Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Language/Ling-3.0-flash
InclusionAI logoInclusionAI·Language ModelsOpen weights

Ling-3.0-flash

View rankingsHugging Facehuggingface.co
Intelligence#143Coding#114
Context262K
Parameters124B (5.1B active)
ReleasedJul 2026

Ling-3.0-flash is a native hybrid reasoning model developed by InclusionAI, the artificial general intelligence initiative of Ant Group. Architected as a Mixture-of-Experts (MoE) system, it comprises 124 billion total parameters with approximately 5.1 billion parameters activated per token. The model is specifically engineered for production-scale agentic workflows, aiming to provide a high-speed execution node that balances reasoning density with low computational overhead.

The model's architecture utilizes a native hybrid-linear attention stack, alternating Kimi Delta Attention (KDA) and Multi-head Latent Attention (MLA) layers at a 5:1 ratio. This configuration includes fine-grained diagonal gating in state updates and a 1/64 sparse MoE activation ratio. These technical optimizations are designed to enhance long-range memory and maintain performance stability during extensive multi-turn interactions and long-horizon agentic tasks.

Ling-3.0-flash supports a native context window of 262,144 tokens, which can be extended up to 1 million tokens. It natively integrates hierarchical caching technologies, such as SGLang HiCache and Mooncake, which utilize physical dual-pools and a cluster-shared L3 cache to minimize redundant recomputation and significantly reduce time-to-first-token (TTFT) in long-input scenarios. The model is particularly optimized for the "planning-execution separation" paradigm, serving as a specialized executor for tasks including coding, search, and deep research.

In practical applications, the model demonstrates refined instruction following and spatial understanding, enabling the construction of physical scene grids and relative positions in programming environments. It was trained using over 10,000 interactive environments to achieve end-to-end closed-loop execution across diverse benchmarks, reportedly matching or exceeding the performance of trillion-parameter predecessors like Ring-2.6-1T in core productivity and tool-driven workflows.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Ling-3.0-flash ranks

Ling-3.0-flash is highlighted in the table below. Switch the metric to see how the ordering changes.