GPT Image 2
by OpenAI
Near-perfect text in images, and it follows a long brief.
What it is
GPT Image 2 is OpenAI's latest image model, and per its documentation it is "capable of creating extremely detailed images with fine typography". The headline capability is text: "near-perfect text rendering", integrating written language naturally into scenes , handwritten notes, signage, interface labels, posters , with correct spelling and consistent spacing, across Latin and CJK scripts.
The second thing it does unusually well is follow a long instruction. Its documentation notes that instruction-following is significantly improved and that the model preserves composition, lighting choices and fine-grained detail described in long or multi-part prompts. If you write briefs rather than phrases, that is the difference you will notice.
It also supports precise inpainting and outpainting with a mask, so specific regions change while everything else stays untouched.
What it is good at
Near-perfect rendering with correct spelling and consistent spacing, in Latin and CJK scripts. Signage, labels, handwriting, posters.
Composition, lighting and fine detail from long or multi-part prompts are preserved rather than averaged away.
Precise inpainting and outpainting via a mask, so unrelated pixels stay exactly as they were.
A maximum edge of 3840px and flexible dimensions, so the frame fits the job instead of the job fitting the frame.
Where it struggles
- Dimensions must be multiples of 16 on both edges, and total pixels have to land between 655,360 and 8,294,400. The Studio only offers sizes that satisfy that, so you will not run into it, but it is why the size list is what it is.
- The quality setting changes both the reasoning depth and the price, and the difference is not small. It is worth choosing deliberately rather than leaving it.
How to prompt it
Write the brief. This is the model that rewards length and structure: name the composition, the lighting, the materials, and quote any text exactly as it should appear. Where other models need you to be evocative, this one wants you to be specific.
An example that shows what that means
A hand-lettered chalkboard menu leaning against a brick wall in a cafe, the heading "TODAY" at the top in wide capitals, three items beneath it with prices aligned to the right, warm overhead light with a soft falloff to the left, shallow depth of field so the wall behind goes soft, slight chalk dust visible on the ledge.
What we offer here
Straight from the catalogue, so this list never goes stale. The price is the provider’s published rate with our markup on top, and it is on the button before you run anything.
Questions
Is this the best model for text in an image?
It is the strongest claim in the studio: near-perfect rendering with correct spelling and spacing, in Latin and CJK scripts. For layout-heavy work in many languages, Seedream 5 is the other one to try.
Can I edit part of an image and leave the rest?
Yes , precise inpainting and outpainting with a mask. Unrelated pixels stay untouched, which is not true of every editing model.
What does the quality setting change?
Both how deeply the model reasons and what it costs, and the difference is significant. The Studio shows the amount for each setting before you run.