K2 Horizon MoVA 36B A4B
K2 Horizon MoVA 36B A4B is an open-weight large language model developed by the Institute of Foundation Models (IFM) at MBZUAI. The model is built on an advanced Mixture-of-Experts (MoE) architecture. The "A4B" designation signifies its efficient parameter utilization: while the model contains a total of 36 billion parameters, only approximately 4 billion parameters are active per token during inference, balancing high-capacity representation with computational efficiency.
The model's distinguishing feature is the MoVA (Mixture-of-Value Attention) mechanism, which introduces sparsity directly into the multi-head attention layers rather than relying solely on feed-forward MoE routing. By dynamically routing token representations across specialized value projections, MoVA expands model capacity and expressiveness without the computational and memory-bandwidth overhead of dense attention layers.
Architecture and Training
K2 Horizon MoVA 36B A4B was pre-trained on an extensive corpus of approximately 20 trillion tokens, with a heavy emphasis on step-by-step reasoning trajectories, synthetic logic data, and programming code. The model natively supports an expansive context window of 512,000 tokens (512K), enabling long-horizon agentic task execution and full repository-level code comprehension in a single prompt.
Released under the permissive Apache 2.0 license, K2 Horizon MoVA 36B A4B demonstrates competitive performance across challenging reasoning and agentic benchmarks, including Terminal-Bench, tau-bench, GPQA Diamond, and Humanity's Last Exam (HLE). Its architecture delivers frontier-level reasoning and autonomous tool-use capabilities with the serving footprint and latency profile of a compact 4B active-parameter model.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow K2 Horizon MoVA 36B A4B ranks
K2 Horizon MoVA 36B A4B is highlighted in the table below. Switch the metric to see how the ordering changes.