Cosmos3-Super-Text2Image
Cosmos3-Super-Text2Image is a 64-billion parameter image generation model developed by NVIDIA as part of its Cosmos 3 family of omnimodal world foundation models. Released in mid-2026, it is designed for high-fidelity text-to-image synthesis and is categorized by its ability to generate sharp, detailed, and photorealistic imagery. The model represents a significant scale-up from previous iterations, specializing the broader Cosmos 3 world model architecture for static visual generation tasks.
Built on a Mixture-of-Transformers (MoT) architecture, the model unifies multiple modalities—including text, image, video, audio, and actions—into a single framework. This unified design allows the model to leverage cross-modal representations, which NVIDIA suggests leads to a deeper understanding of physical properties, motion, and spatial relationships compared to models trained on images alone. It was trained on a massive dataset comprising approximately 20 trillion tokens, including over a billion real-world images and hundreds of millions of video sequences.
The model is often utilized within an agentic prompt-upsampling harness. This system takes simplified user inputs and expands them into highly detailed, structured descriptions that provide the generator with the necessary context for complex scenes. This approach helps maintain high faithfulness to the original intent while maximizing the visual quality and resolution of the output. Upon release, the model achieved top rankings on open-weight image generation leaderboards, such as the Artificial Analysis Text-to-Image benchmark.
In practical use, Cosmos3-Super-Text2Image is served via frameworks like vLLM-Omni or SGLang. Due to its 64B size, the model generally requires high-end workstation or server-grade GPUs (such as those with 80GB+ VRAM) for local inference. It is released under the OpenMDW-1.1 license, facilitating its use across research and commercial applications in physical AI and world simulation.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Cosmos3-Super-Text2Image ranks
Cosmos3-Super-Text2Image is highlighted in the table below. Switch the metric to see how the ordering changes.