Ling 3.0 Tiny
Ling 3.0 Tiny is a compact Mixture-of-Experts (MoE) language model developed by InclusionAI. Designed for resource-constrained environments and high-throughput agentic workflows, the model contains approximately 7.9 billion total parameters, with only 1.3 billion parameters activated per token. This sparse activation allows the model to provide reasoning-heavy performance while maintaining the speed and low latency profile typical of smaller models.
Architecture and Design
The model utilizes a native hybrid-linear architecture that integrates Kimi Delta Attention (KDA) with Multi-Head Latent Attention (MLA). This hybrid configuration is engineered to balance high-precision transformer accuracy with linear-scaling efficiency, which is particularly effective for processing long-context inputs. By alternating gated KDA layers with sparse MoE blocks, Ling 3.0 Tiny achieves stable performance in long-horizon tasks while minimizing the overall computational footprint.
Key Capabilities
Ling 3.0 Tiny supports dual execution modes: Thinking mode for complex, multi-step logical reasoning and Instant mode for low-latency, immediate responses. It natively supports function calling, structured outputs, and prompt caching. These features make it suitable for production-grade AI agents that require reliable instruction following and precise tool usage during multi-turn interactions.
The model features a native context window of 262,144 tokens, enabling it to handle extensive documentation, large datasets, or deep code repositories without requiring frequent recomputation of context. This long-context capacity, paired with a 32,768-token maximum output, supports complex agentic workflows such as document intelligence and long-horizon research tasks.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Ling 3.0 Tiny ranks
Ling 3.0 Tiny is highlighted in the table below. Switch the metric to see how the ordering changes.