Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Language/Ling 3.0 Tiny
InclusionAI logoInclusionAI·Language ModelsOpen weights

Ling 3.0 Tiny

View rankingsHugging Faceinclusion-ai.org
Intelligence#253Coding#259
Context262K
Parameters7.9B
ReleasedAug 2026

Ling 3.0 Tiny is a lightweight language model developed by InclusionAI, the artificial intelligence initiative of Ant Group. As the most compact member of the Ling 3.0 family, it is designed to balance high-level reasoning and agentic performance with a minimal computational footprint. The model is optimized for responsive AI agents, instruction following, and natural multi-turn conversations in resource-constrained environments, including local and edge deployment scenarios.

The model utilizes a sparse Mixture-of-Experts (MoE) architecture with a total of 7.9 billion parameters, of which only 1.3 billion are activated per token. Its internal structure features an efficient hybrid linear attention mechanism that integrates Kimi Delta Attention (KDA) and Multi-Head Latent Attention (MLA) in a 3:1 alternating stack. This configuration uses a sparse MoE feed-forward network with 128 routed experts, where each token activates eight routed experts and one shared expert to maintain long-context modeling capability at a low inference cost.

A central feature of Ling 3.0 Tiny is its dual operational modes: Thinking mode and Instant mode. In Thinking mode, the model performs internal step-by-step reasoning before providing an answer, outputting its chain of thought separately. Instant mode provides direct responses with lower latency. Users can toggle these modes per request via the enable_thinking configuration to optimize for either reasoning depth or response speed.

Equipped with a 262,144-token context window, the model supports native function calling and prompt caching to improve efficiency in complex workflows. It is validated for high-performance execution on consumer-grade hardware, such as Apple Silicon (MacBook and Mac mini) and NVIDIA DGX Spark, enabling sophisticated mathematical reasoning and coding assistance without the need for datacenter-class GPUs.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Ling 3.0 Tiny ranks

Ling 3.0 Tiny is highlighted in the table below. Switch the metric to see how the ordering changes.