Ling 3.0 Tiny
Ling 3.0 Tiny is a lightweight language model developed by InclusionAI, the artificial intelligence initiative of Ant Group. As the most compact member of the Ling 3.0 family, it is designed to balance high-level reasoning and agentic performance with a minimal computational footprint. The model is optimized for responsive AI agents, instruction following, and natural multi-turn conversations in resource-constrained environments, including local and edge deployment scenarios.
The model utilizes a sparse Mixture-of-Experts (MoE) architecture with a total of 7.9 billion parameters, of which only 1.3 billion are activated per token. Its internal structure features an efficient hybrid linear attention mechanism that integrates Kimi Delta Attention (KDA) and Multi-Head Latent Attention (MLA) in a 3:1 alternating stack. This configuration uses a sparse MoE feed-forward network with 128 routed experts, where each token activates eight routed experts and one shared expert to maintain long-context modeling capability at a low inference cost.
A central feature of Ling 3.0 Tiny is its dual operational modes: Thinking mode and Instant mode. In Thinking mode, the model performs internal step-by-step reasoning before providing an answer, outputting its chain of thought separately. Instant mode provides direct responses with lower latency. Users can toggle these modes per request via the enable_thinking configuration to optimize for either reasoning depth or response speed.
Equipped with a 262,144-token context window, the model supports native function calling and prompt caching to improve efficiency in complex workflows. It is validated for high-performance execution on consumer-grade hardware, such as Apple Silicon (MacBook and Mac mini) and NVIDIA DGX Spark, enabling sophisticated mathematical reasoning and coding assistance without the need for datacenter-class GPUs.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Ling 3.0 Tiny ranks
Ling 3.0 Tiny is highlighted in the table below. Switch the metric to see how the ordering changes.