MiMo-V2.6-Flash
MiMo-V2.6-Flash is a natively omnimodal sparse Mixture-of-Experts (MoE) reasoning model developed by Xiaomi. Built as a balance between intelligence, efficiency, and cost, the model accepts text, image, video, and audio inputs natively within a single checkpoint while offering a 1 million token context window. It features 309 billion total parameters with 15 billion active parameters per forward pass, designed to execute complex professional workflows, code generation, and multi-agent tasks efficiently.
Architecture and Capabilities
The model utilizes a hybrid attention mechanism alongside a sparse MoE architecture comprising 48 transformer layers and 256 routed experts. To accelerate inference and maintain high throughput, it incorporates specialized modules including a multi-layer speculative decoding drafter (DFlash-style) and support for native FP8 weight configurations. MiMo-V2.6-Flash provides robust tool calling, deep thinking processes, structured outputs, and web search integration, positioning it as a low-cost solution for high-frequency professional tasks.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow MiMo-V2.6-Flash ranks
MiMo-V2.6-Flash is highlighted in the table below. Switch the metric to see how the ordering changes.