We're launching our realtime TTS endpoints for Qwen3-TTS 1.7B today.
> "Fast" endpoint: client-side p90 time-to-first-audio (TTFA) at 50ms.
> "Standard" endpoint: dirt cheap at $5 per 1M characters, still faster than industry avg at 200 ms TTFA.
For comparison, 11Labs V3 is $100 for 1M and Cartesia Sonic 3.6 is $50.
We are working hard on further lowering the price of our endpoints, as well as adding more voices and support for input streaming.
We believe open-source will win - not just in LLMs but also in multimodal :)
During public beta, all endpoints are free. Try it out and let me know how it goes!
toebee•51m ago
We're launching our realtime TTS endpoints for Qwen3-TTS 1.7B today. > "Fast" endpoint: client-side p90 time-to-first-audio (TTFA) at 50ms. > "Standard" endpoint: dirt cheap at $5 per 1M characters, still faster than industry avg at 200 ms TTFA.
For comparison, 11Labs V3 is $100 for 1M and Cartesia Sonic 3.6 is $50.
TTFA is measured using our OSS benchmarking tool from us-west-1 and us-east-1 (https://github.com/nari-labs/benchmarks.git).
Quality: link to TTS samples (https://narilabs.com/blog/introducing-nari-qwen3-tts/). Based on preliminary benchmarks by a 3rd-party (Speko AI), we scored above leading TTS providers such as Grok and Rime (https://narilabs.com/blog/qwen3-tts-speko-benchmark/)!
We are working hard on further lowering the price of our endpoints, as well as adding more voices and support for input streaming.
We believe open-source will win - not just in LLMs but also in multimodal :) During public beta, all endpoints are free. Try it out and let me know how it goes!