Ideogram 4 delivers precise prompt adherence via JSON workflow, but counting remains weak
Complex JSON prompts yield stunning accuracy, but the model can't count to three.
Ideogram 4, the latest image generation model from Ideogram, demands a JSON-based prompting approach that trades simplicity for unprecedented control. Users like the Reddit poster found that translating complex prompts into structured JSON—using tools like KJ's ComfyUI nodes and an LLM to generate the basic structure—enables the model to follow detailed instructions to the letter. The example workflow requires 35 steps with Euler/simple sampler, generating images in 90–130 seconds on an RTX 4090. The model can handle NSFW content, though outputs are described as 'Barbie doll' style, and it accurately renders sensitive subjects like Michelangelo's David without unnecessary censorship.
Despite its strengths in composition and detail, Ideogram 4 has a notable weakness: counting. Even with bounding boxes defined in JSON, the model frequently adds extra characters to scenes requiring a specific number of people. The poster selected the best-of-eight images, noting that the model excels at photographic realism and anime styles when prompted correctly. The increased workload—moving bounding boxes and editing JSON directly—is justified for users who need precise, reproducible results, positioning Ideogram 4 as a tool for creators who know exactly what they want, rather than those exploring random outputs.
- Uses JSON-based prompting for fine-grained control; typical generation time 90–130s on RTX 4090
- Excellent at following complex prompts but struggles with counting objects (e.g., extra people)
- Handles NSFW with modesty and accurately depicts culturally sensitive subjects like David statue
Why It Matters
Ideogram 4 raises the bar for precision in AI image generation, but the JSON overhead limits its accessibility to advanced users.