Mercury 2.5
Mercury 2.5 is a production-grade diffusion language model (dLLM) developed by Inception AI. Unlike conventional autoregressive models that generate tokens sequentially, Mercury 2.5 utilizes a diffusion-based framework to produce and refine blocks of tokens in parallel. This architectural distinction enables exceptionally high throughput, exceeding 1,000 tokens per second on standard hardware, which significantly reduces latency and operational costs for real-time applications.
Positioned as a high-performance reasoning model tailored for latency-sensitive workflows, Mercury 2.5 supports advanced features such as parallel tool calling and structured JSON outputs via explicit schema alignment. It features an extended 260K token context window, making it suitable for complex agentic loops, code generation, and enterprise search tasks where both speed and reasoning depth are critical.
Architecture and Capabilities
The model builds upon Inception's proprietary diffusion paradigm, starting with a rough draft of the output and iteratively refining tokens simultaneously. This design aims to mitigate the traditional tradeoff where deeper reasoning forces slower sequential generation speeds. Mercury 2.5 is accessible exclusively via managed APIs and enterprise hosting platforms rather than open weights, serving as a specialized component for high-frequency middle-tier agent tasks.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Mercury 2.5 ranks
Mercury 2.5 is highlighted in the table below. Switch the metric to see how the ordering changes.