Published: 2026-09-09
Sketch-to-layout in ChatGPT Image 2.5, and a real test of edit consistency
Chapters / key moments (click to jump — plays here on the page)
The technique worth copying is the pairing: draw the spatial relationships, then describe what each shape means. Coloured rectangles and circles on a canvas carry position and scale, which is exactly what prose is bad at; the accompanying text says which blob is the desk and which are the cameras. The output is a labelled aerial 3D layout. The other half is a genuine test rather than a demo — the reviewer makes a series of small edits specifically to measure drift, because his standing complaint about the previous model is that a face becomes unrecognisable by the third or fourth generation.
Source video
"ChatGPT Image 2.5 Just Dropped — Here's Everything That's New" by Bart Slodyczka — Watch on YouTube →
Gotchas & Caveats
- This is an image-generation feature rather than agent tooling. It earns a page here for the sketch-plus-description technique and the edit-consistency measurement, both of which generalise to any spatial or layout task you hand a model.
- Availability was still rolling out when recorded; if you do not see the updated images tab, you do not have it yet.
- The consistency finding comes from one sequence of edits on one image. It is a real observation and a small sample.
- Generation timings were measured on the reviewer's account under normal load and will vary.
Key Takeaways
- The sketch feature lives behind the plus icon in the images tab, opening a canvas with shape and freehand tools. Access was still rolling out at the time of recording — the updated images-tab icon is the tell.
- Shapes carry spatial intent, text carries meaning. The working prompt names each shape by colour and says what it represents: white square is floor space, green rectangle is the desk with two monitors, yellow circles are camera positions, purple rectangle is where you sit.
- Generation times, measured: roughly a minute and a half for the first render, about 1:50 for a subsequent edit. Progress climbs at roughly 1% per second, so the bar is a usable estimate.
- Consistency across edits held. Flicking between generations, the reviewer finds essentially everything preserved between the original and the edited version — the only visible difference being lighting. That directly addresses the complained-about failure in GPT Image 2, where a subject degraded by the third or fourth generation.
- It is not perfect and he says so: the generated signage came out sideways, and some details in the render are inventions rather than reproductions of what was drawn.
- Edits are targeted by region. Marking an area and describing the change ("in that red circle, put a really tall plant") is how small changes are requested, rather than re-prompting the whole image.





