Qwen3-ASR 1.7B is an excellent OSS Speech-to-Text model. But it is slow out of the box. We built an inference engine around it, and now it is the fastest realtime transcription model (as of Sep 2026) at 40 ms time-to-final-segment (TTFS). And the word error rate (WER) sits top 3 on Coval’s independent benchmarks. OSS will win, not just in LLMs but also in multimodal, one step at a time.
toebee•53m ago