TechNewsReel
Live

Inception Labs Debuts Mercury 2.5 Diffusion LLM Hitting 1,107 Tokens Per Second

The new dLLM architecture eliminates latency bottlenecks for real-time AI agents and voice applications.

TechNewsReel Newsroom · September 8, 2026

Inception Labs has released Mercury 2.5, a diffusion-based language model (dLLM) designed to break the latency barriers of traditional AI. By refining tokens in parallel rather than generating them sequentially, the model achieves speeds that could fundamentally alter the responsiveness of real-time AI applications.

According to Inception Labs, Mercury 2.5 reaches a performance speed of 1,107 tokens per second when running on NVIDIA GPUs. The model features a context window of 260,000 tokens and is positioned as a high-speed alternative to traditional autoregressive models. Standard pricing is set at $0.20 per million input tokens and $0.75 per million output tokens, though the company is offering an 80% discount during the launch period. Inception Labs claims the model provides a significant intelligence boost over its predecessor, Mercury 2, with quality levels comparable to Gemini 3.5 Flash-Lite and Claude Haiku 4.5.

The Shift to Diffusion

Most modern large language models rely on autoregressive generation, meaning they predict the next token one by one. This process creates a linear delay that can hinder the user experience in high-frequency environments. Inception Labs is pivoting toward a diffusion-based architecture to solve these bottlenecks. This approach allows the model to refine multiple tokens simultaneously, drastically reducing the time between a user's prompt and the model's complete response.

This architectural shift is specifically targeted at the next generation of AI agents. For high-frequency coding sub-agents or complex RAG (Retrieval-Augmented Generation) pipelines, where milliseconds of delay can disrupt a workflow, the ability to generate text at over 1,000 tokens per second transforms the interaction from a turn-based exchange into a near-instantaneous stream.

Industry Implications

The release of Mercury 2.5 suggests a narrowing of the traditional trade-off between speed and intelligence. Shruti Koparkar, Senior Manager of Product at NVIDIA, noted that the combination of sustained speeds, low costs, and increased intelligence demonstrates how quickly new architectures can mature into production-ready systems on the NVIDIA platform.

For developers, the primary value lies in the reduction of P99 response times. Oliver Silverstein, Co-founder and CEO of OpenCall, stated that after switching to Mercury, their P99 response time dropped from several minutes to just one second, while their P50 dropped from 0.4 seconds to under 0.2 seconds. This level of performance is critical for voice agents, where any perceptible lag can make a conversation feel unnatural.

What to Watch

While the raw speed and pricing are verified, the industry will be watching to see how dLLMs handle complex, multi-step reasoning compared to their autoregressive counterparts. As Inception Labs pushes the boundaries of throughput, the next benchmark for the company will be maintaining this speed across increasingly complex agentic tasks and larger-scale deployments.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.