HunyuanImage 3.0
HunyuanImage-3.0 is a native multimodal image generation and editing model developed by Tencent. Unlike traditional systems that rely entirely on separate image-only diffusion structures, HunyuanImage-3.0 features a large-scale Mixture-of-Experts (MoE) architecture unified within an autoregressive framework. It totals approximately 80 billion parameters, with around 13 billion parameters active per token during inference, allowing for high computational efficiency relative to its massive scale.
The model series includes dedicated variants such as the base text-to-image generator and the HunyuanImage-3.0-Instruct family, which handles instruction-driven image editing, style transfer, and multi-image fusion. By incorporating world knowledge reasoning and a native chain-of-thought schema, the Instruct variants can parse complex natural language descriptions exceeding one thousand characters, manipulate precise typographic elements or embedded text strings, and preserve unedited regions during modifications.
Official prompt guidelines recommend detailed structural prompts for the base model, while the Instruct checkpoints support automatic prompt expansion and refinement through built-in reasoning layers. Because of its high capacity and MoE layer distribution, the system requires substantial hardware resources and is typically deployed on multi-GPU server environments or specialized quantized setups.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow HunyuanImage 3.0 ranks
HunyuanImage 3.0 is highlighted in the table below. Switch the metric to see how the ordering changes.