Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Language/MiMo-V2.6-Pro
Xiaomi logoXiaomi·Language ModelsOpen weights

MiMo-V2.6-Pro

View rankingsHugging Facemimo.xiaomi.com
Intelligence#22
Context1M
Parameters1.02T (42B active)
ReleasedSep 2026

MiMo-V2.6-Pro is a flagship, natively omni-modal foundation model developed by Xiaomi. Released in September 2026 as part of the MiMo-V2.6 series, the model is designed to process and reason across multiple modalities within a single architecture. It accepts inputs in the form of text, images, video, and audio, enabling it to perform complex cross-modal tasks without the need for external orchestration or specialized adapters.

The model utilizes a Mixture-of-Experts (MoE) architecture with approximately 1.02 trillion total parameters, featuring 42 billion active parameters per token. A central component of its development was the use of scaled reinforcement learning (RL) and recursive self-improvement (RSI) techniques. Xiaomi trained the model using a public reinforcement learning run that focused on verifiable, complex tasks, allowing the model to continuously expand its reasoning capabilities through iterative feedback and exploration.

With a context window of 1 million tokens, MiMo-V2.6-Pro is optimized for long-horizon tasks and professional workflows. Its capabilities extend to multi-agent collaboration, 3D modeling, and advanced software engineering. In practical applications, the model has demonstrated proficiency in materials research—such as designing molecular structures to capture pollutants—and in agentic computer use, where it can interpret user interfaces and execute multi-step operations.

On benchmarks such as the Artificial Analysis Intelligence Index, MiMo-V2.6-Pro achieved a composite score of 46.32, making it a high-performing open-weight model for complex reasoning and agent-based applications. Its architecture includes specialized components like a 681-million-parameter vision encoder and a dedicated audio tokenizer, which provide the foundation for its native multimodal perception and decision-making.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How MiMo-V2.6-Pro ranks

MiMo-V2.6-Pro is highlighted in the table below. Switch the metric to see how the ordering changes.