Qwen3.8 2.4T A95B
Qwen3.8-2.4T-A95B is a massive Mixture-of-Experts (MoE) language model developed by Alibaba's Qwen team. As the open-weights counterpart to the Qwen3.8-Max flagship, it is the first model in the "Max" performance tier to be released for public use. It is optimized for high-level reasoning, complex programming tasks, and long-horizon agentic operations.\n\nThe model architecture comprises 2.4 trillion total parameters, with approximately 95 billion parameters activated during each forward pass. It utilizes a 92-layer structure organized in a repeating 23-block pattern that alternates between Gated DeltaNet linear attention and traditional Gated Attention. This system incorporates 512 experts, routing tokens to 10 specific experts plus one shared expert per step to maintain a balance between frontier-level capability and inference efficiency.\n\nA key feature of the model is its support for configurable reasoning depth through the reasoning_effort parameter, allowing users to modulate the complexity of the model's internal thinking process. Additionally, the preserve_thinking flag enables the retention of reasoning context across conversation turns, which is particularly useful for multi-step engineering and research agents.\n\nWhile derived from the multimodal Qwen3.8-Max, this open-weight release is a text-only model and does not include native vision or multimodal inputs. It natively supports a context window of 262,144 tokens, though it is architecturally capable of extension up to approximately 1.01 million tokens. The model weights are distributed in the Safetensors format and optimized for deployment using high-throughput inference engines like vLLM and SGLang.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Qwen3.8 2.4T A95B ranks
Qwen3.8 2.4T A95B is highlighted in the table below. Switch the metric to see how the ordering changes.