Ling-3.0-flash
Ling-3.0-flash is a native hybrid reasoning model developed by InclusionAI, the artificial general intelligence initiative of Ant Group. Architected as a Mixture-of-Experts (MoE) system, it comprises 124 billion total parameters with approximately 5.1 billion parameters activated per token. The model is specifically engineered for production-scale agentic workflows, aiming to provide a high-speed execution node that balances reasoning density with low computational overhead.
The model's architecture utilizes a native hybrid-linear attention stack, alternating Kimi Delta Attention (KDA) and Multi-head Latent Attention (MLA) layers at a 5:1 ratio. This configuration includes fine-grained diagonal gating in state updates and a 1/64 sparse MoE activation ratio. These technical optimizations are designed to enhance long-range memory and maintain performance stability during extensive multi-turn interactions and long-horizon agentic tasks.
Ling-3.0-flash supports a native context window of 262,144 tokens, which can be extended up to 1 million tokens. It natively integrates hierarchical caching technologies, such as SGLang HiCache and Mooncake, which utilize physical dual-pools and a cluster-shared L3 cache to minimize redundant recomputation and significantly reduce time-to-first-token (TTFT) in long-input scenarios. The model is particularly optimized for the "planning-execution separation" paradigm, serving as a specialized executor for tasks including coding, search, and deep research.
In practical applications, the model demonstrates refined instruction following and spatial understanding, enabling the construction of physical scene grids and relative positions in programming environments. It was trained using over 10,000 interactive environments to achieve end-to-end closed-loop execution across diverse benchmarks, reportedly matching or exceeding the performance of trillion-parameter predecessors like Ring-2.6-1T in core productivity and tool-driven workflows.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Ling-3.0-flash ranks
Ling-3.0-flash is highlighted in the table below. Switch the metric to see how the ordering changes.