Motif 3
Motif 3 is a large-scale, decoder-only language model developed by the South Korean AI company Motif Technologies. Released as part of Korea's sovereign foundation model project (known as "Dokpamo"), it is notable for being built entirely from scratch using a proprietary architecture rather than building upon existing open-source frameworks. The model is designed to provide high intelligence and reasoning capabilities while maintaining competitive inference efficiency compared to other frontier models.
The architecture of Motif 3 utilizes a Mixture-of-Experts (MoE) design with a total of 314.8 billion parameters. It employs a fine-grained sparsity mechanism featuring 384 routed experts per layer, with only 8 experts (approximately 13.2 billion parameters) activated per token. This configuration allows the model to maintain a massive knowledge base and expert capacity while constraining the computational cost required for individual inference steps.
Technical Innovations
Motif 3 introduces several novel architectural components aimed at optimizing attention and stability:
- Grouped Differential Latent Attention (GDLA): A mechanism that integrates the compressed Key-Value (KV) representations of Multi-head Latent Attention with Grouped Differential Attention to reduce memory overhead and improve signal-to-noise ratios.
- Expert-Specific PolyNorm: A specialized activation function that uses learned polynomial coefficients constrained by a sigmoid parameterization, promoting higher specialization within individual experts.
- Modified Manifold-Constrained Hyper-Connections (mHC): A multi-stream convex mixing structure that replaces standard residual additions to enhance optimization stability during large-scale training.
Training and Performance
The model was pretrained on a diverse 12.5 trillion token corpus spanning STEM, code, mathematics, and multilingual content. It supports a native context window of 256,000 tokens (262,144), handled through window-aware context parallelism. Motif 3 also includes a built-in Multi-Token Prediction (MTP) head to facilitate self-speculative decoding. Benchmarks indicate strong performance in long-horizon agentic tool use, mathematical reasoning, and instruction following, particularly in contexts requiring high fidelity and calibrated abstention from hallucinations.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Motif 3 ranks
Motif 3 is highlighted in the table below. Switch the metric to see how the ordering changes.