Published: 2026-09-09

Sketch-to-layout in ChatGPT Image 2.5, and a real test of edit consistency

Chapters / key moments (click to jump — plays here on the page)

The technique worth copying is the pairing: draw the spatial relationships, then describe what each shape means. Coloured rectangles and circles on a canvas carry position and scale, which is exactly what prose is bad at; the accompanying text says which blob is the desk and which are the cameras. The output is a labelled aerial 3D layout. The other half is a genuine test rather than a demo — the reviewer makes a series of small edits specifically to measure drift, because his standing complaint about the previous model is that a face becomes unrecognisable by the third or fourth generation.

Source video

"ChatGPT Image 2.5 Just Dropped — Here's Everything That's New" by Bart SlodyczkaWatch on YouTube →

Gotchas & Caveats

  • This is an image-generation feature rather than agent tooling. It earns a page here for the sketch-plus-description technique and the edit-consistency measurement, both of which generalise to any spatial or layout task you hand a model.
  • Availability was still rolling out when recorded; if you do not see the updated images tab, you do not have it yet.
  • The consistency finding comes from one sequence of edits on one image. It is a real observation and a small sample.
  • Generation timings were measured on the reviewer's account under normal load and will vary.

Key Takeaways

  • The sketch feature lives behind the plus icon in the images tab, opening a canvas with shape and freehand tools. Access was still rolling out at the time of recording — the updated images-tab icon is the tell.
  • Shapes carry spatial intent, text carries meaning. The working prompt names each shape by colour and says what it represents: white square is floor space, green rectangle is the desk with two monitors, yellow circles are camera positions, purple rectangle is where you sit.
  • Generation times, measured: roughly a minute and a half for the first render, about 1:50 for a subsequent edit. Progress climbs at roughly 1% per second, so the bar is a usable estimate.
  • Consistency across edits held. Flicking between generations, the reviewer finds essentially everything preserved between the original and the edited version — the only visible difference being lighting. That directly addresses the complained-about failure in GPT Image 2, where a subject degraded by the third or fourth generation.
  • It is not perfect and he says so: the generated signage came out sideways, and some details in the render are inventions rather than reproductions of what was drawn.
  • Edits are targeted by region. Marking an area and describing the change ("in that red circle, put a really tall plant") is how small changes are requested, rather than re-prompting the whole image.

Weekly Digest — In Your Inbox

Get the week's top AI agent news, updates, and guides — every Friday.