Gradium AI Debuts TTS Model with 81% Accuracy at Lightning Speed
Voice agents finally sound reliable on the phrases that cause the most trouble—order numbers, callback digits, email addresses. Gradium AI has rolled out a new default text-to-speech model that delivers an 81.0% human-rated pass rate on a 500-sentence hard-case set spanning five languages, surpassing Cartesia Sonic 3.6 (75.1%) and ElevenLabs v3 Conversational (65.4%). At the same time, it achieves a 216 ms median time to first audio on Coval’s benchmark, 170 ms faster than its predecessor and with a tight 30 ms interquartile spread.
Built for the edge cases that break other systems
Gradium’s evaluation set focuses on ten strict criteria across English, German, French, Spanish and Portuguese, including spelling, acronyms, alphanumeric tokens, dates, numbers of all sizes, and emails. Each sentence must be pronounced flawlessly by independent native speakers; a single missed digit or letter fails the entire test. The company open-sourced the dataset on Hugging Face under CC BY 4.0, inviting external scrutiny and continuous improvement.
No migration, no downtime
Existing customers simply continue using the same voice IDs and endpoints; the new model became the default across Gradium’s API and Studio on August 31, 2026. New teams can integrate via the Python SDK and connect to the WebSocket TTS endpoint without altering their workflows. Gradium is also offering one million credits to users who submit complete failure reports on its Discord, reinforcing a feedback loop that could further sharpen accuracy.
Why it matters
For any business running voice agents, the difference between 65% and 81% accuracy on critical phrases can translate directly into fewer callbacks and higher first-call resolution. Meanwhile, halving the time-to-first-audio from roughly 386 ms to 216 ms reduces perceived latency for callers, a tangible improvement in user experience. By combining strict, human-validated benchmarks with open data and seamless deployment, Gradium is nudging the industry toward models that are both trustworthy and fast enough for real-time interactions.
Source: MarkTechPost. AI-assisted editorial synthesis — TechnoExpress.

