OpenAI Images
OpenAI image generation — gpt-image-2 (flagship, reasoning-based), gpt-image-1 (native multimodal), dall-e-3 (legacy), dall-e-2 (legacy).
Flagship OpenAI image model (Apr 2026) — reasoning-based composition, near-perfect multilingual text rendering (CJK, Arabic, Devanagari), supports up to 3840px on the long edge, accepts up to 16 input images for composite edits.
2 methods
Native GPT-4o image model — t2i + image edit, low/medium/high quality.
2 methods
Legacy DALL·E 3 — text-to-image only, standard or HD.
1 method
Legacy DALL·E 2 — t2i + edit at 256/512/1024 squares.
2 methods
Budget multimodal model for image understanding. NB: inflates image tokens ~33x — gpt-4.1-mini is cheaper and better for vision.
1 method
Cheap, reliable multimodal model — image → text or strict JSON. Recommended OpenAI vision option.
1 method
Lowest-cost OpenAI vision model (reasoning family) — image → text or strict JSON.
1 method
High-accuracy multimodal model — image → text or strict JSON. For accuracy-sensitive jobs.
1 method
Flagship multimodal model — image → text or strict JSON. Highest accuracy, highest cost.
1 method