Nari Labs Optimizes Qwen3 Speech Models to Lead Coval Benchmarks
New serving engines for open-source TTS and ASR models outperform official endpoints in latency and accuracy.
Nari Labs has announced that its serving engines for Qwen3-TTS and Qwen3-ASR now lead the Coval Voice AI benchmarks in quality and latency. The company claims its optimizations place these open-source models on the Pareto frontier for quality, latency, and cost across both speech-to-text and text-to-speech categories.
According to data from Nari Labs, the Qwen3-ASR Fast implementation is ranked first in Time-to-Final-Segment (TTFS) with a p50 of 44 ms and a Word Error Rate (WER) of 3.6%. In the text-to-speech category, Qwen3-TTS Fast ranks first in WER at 3.8% and second in Time-to-First-Audio (TTFA) with a p50 of 63 ms. These performance gains are particularly evident when compared to official sources; Nari Labs notes that the official Qwen3 TTS Flash Realtime endpoint exhibits a significantly higher WER of 8.8% and a median TTFA of 692 ms.
The Inference Problem
This push for efficiency follows Nari Labs' previous work on Dia, an open-source model designed for natural dialogue. The company is now focusing on what it describes as the "inference problem," arguing that the current market dominance of closed-source speech models is driven by serving efficiency rather than superior model architecture. By optimizing how the Qwen3 series is served, Nari Labs aims to prove that high-performance open-source speech models are viable for production-grade voice agents.
Market Implications
Low latency, specifically in TTFA and TTFS, is critical for the perceived responsiveness of AI voice agents. When open-source models can match or exceed the speed and accuracy of closed-source leaders, it lowers the barrier for developers to build sovereign and cost-effective voice interfaces. Nari Labs is positioning its offerings as a high-value alternative, pricing Qwen3-TTS Fast at $10 per 1 million characters—tied for the cheapest on Coval's pricing directory. Similarly, Qwen3-ASR Fast is priced at $0.12 per hour, tying for the second lowest price among models with known public rates.
What's Next
As Nari Labs pushes these optimizations, the industry will be watching to see if other open-source serving frameworks can replicate these gains. While the Coval benchmarks provide a snapshot of current performance, the long-term viability of these models will depend on their stability at scale in real-world production environments. For now, the Nari Labs team maintains that their implementation sits on the quality-latency Pareto Frontier, challenging the necessity of closed-source ecosystems for high-speed voice AI.