HiDream AI
MIT-licensed image generation foundation model for text-to-image, editing, and subject-driven personalization up to 2048px.
HiDream-O1-Image-Dev-2604 is a focused image model rather than a general reasoning or coding system. Official materials support a 9B-parameter, MIT-licensed model with text-to-image, editing, subject personalization, 28-step inference, and up-to-2048px output. Its 1189 Artificial Analysis Text-to-Image Arena Elo, behind Cosmos3-Super-Text2Image at 1218 and tied with its 4Step variant, supports solid image-generation capability. This makes its multimodal I/O score stronger than Bria 3.2’s 78, while technical capability remains below broad high-end anchors such as GPT-5.5 and Claude Opus 4.8, which have substantially wider validated capabilities. The MIT license, Hugging Face distribution, and public GitHub repository make it materially more accessible for self-hosted experimentation than closed API models. That supports cost effectiveness despite the 9B scale and 28-step workload, but not a high pricing-clarity score: supplied official sources do not establish first-party hosted pricing or service terms. Artificial Analysis reports substantial provider latency variation, from 0.9–2.2 seconds at MachGen to 9.3–206.8 seconds at WaveSpeed, so speed and operational availability depend heavily on the selected endpoint. Coding and general reasoning scores are intentionally low because DeepSWE and LiveCodeBench list no result for this exact model, SWE-bench evidence was not found, and no Terminal-Bench or Aider result is supplied. Public signal is moderate rather than broad: the Arena placement is useful preference evidence, but no adoption metrics are provided. Evidence quality is reasonable for product identity and licensing, but limited for quality, reliability, safety, and deployment comparisons.