Logocrafiq.ai

An AI-powered assets creation platform. Generate, edit & ship content faster.

Explore

  • Home
  • Contact
  • Pricing
  • Blog

Features

  • 2D Assets Generator
  • Text to 3D
  • Video Generator
  • Sound Effects
  • All Features

Rankings

  • Image generation
  • Image upscaling
  • Video generation
  • 3D generation
  • Text generation
  • Music generation
  • Speech generation

© 2026 Crafiq. All rights reserved.

Privacy PolicyTermsImpressum
Models/Language/Ling-3.0-flash-VL
InclusionAI logoInclusionAI·Language ModelsOpen weights

Ling-3.0-flash-VL

View rankingsHugging Faceinclusion-ai.org
Intelligence#126Coding#92
Context262K
Parameters124B
ReleasedSep 2026

Ling-3.0-flash-VL is a multimodal large language model (MLLM) developed by inclusionAI, the AI research laboratory of Ant Group. Released in September 2026, it is the first vision-integrated model within the Ling-3.0-flash family. The model is designed to natively process and reason across text, image, and video inputs, supporting applications such as document intelligence, medical report interpretation, and agentic workflows that require visual feedback for GUI navigation.

The model utilizes a sparse Mixture-of-Experts (MoE) architecture with a total of 124 billion parameters, of which 5.5 billion are activated per token during inference. Its 42-layer hybrid backbone, identified as BailingMoeV3VL, alternates between Knowledge-Dense Attention (KDA) and Gated Multi-head Latent Attention (MLA) layers at a 5:1 ratio. This architectural choice is optimized for inference efficiency and long-context stability, enabling the model to support a context window of up to 1 million tokens.

Visual processing is powered by a ViT visual encoder and a two-layer MLP projector that aligns visual features with the model's textual representations. To handle the temporal dimension of video data, Ling-3.0-flash-VL employs VideoRoPE (Rotary Positional Embedding), which allows the model to encode both spatial positions and the temporal order of frames. This facilitates advanced capabilities such as event localization, temporal reasoning, and visual change detection over time. The weights for Ling-3.0-flash-VL are released under the MIT license, supporting both open-source research and commercial integration.

Create with Crafiq

Generate images, 3D models, video and audio in one studio.

Explore the studio

How Ling-3.0-flash-VL ranks

Ling-3.0-flash-VL is highlighted in the table below. Switch the metric to see how the ordering changes.