Celeris-1
Celeris-1 is a flagship language model developed by Celeris Labs, an AI research organization focused on ultra-low latency inference. Introduced in July 2026, the model is built on a novel diffusion-based architecture rather than the traditional autoregressive approach common in large language models. This design allows the model to generate text in parallel rather than token-by-token, significantly reducing response times for interactive applications.
Architecture and Performance
The model's diffusion decoding process differs from the standard of sequential token generation. Instead of predicting one word at a time based on previous context, Celeris-1 generates and refines entire sequences simultaneously. This architectural shift enables p50 response latencies of approximately 158 milliseconds and a throughput exceeding 1,600 tokens per second. In benchmark testing, the model achieved a 75.9% score on MMLU-Pro, positioning its reasoning capabilities near frontier-class systems.
Capabilities and Use Cases
Celeris-1 is specifically optimized for high-frequency, latency-sensitive agentic workflows. Key applications include structured data extraction, classification, query rewriting, and real-time voice orchestration. The model maintains a context window of 8,192 tokens, making it suitable for short-to-medium length tasks where immediate response delivery is a priority over long-horizon document analysis. Its interface follows standard chat-completions structures, allowing for integration into existing tool-use and automation pipelines.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Celeris-1 ranks
Celeris-1 is highlighted in the table below. Switch the metric to see how the ordering changes.