3 ms·
Routing makes even more sense for voice than text, but the constraint space is trickier: latency budget (round-trip vs streaming), WER on accented speech, and T
by modgate 2mo ago
Routing makes even more sense for voice than text, but the constraint space is trickier: latency budget (round-trip vs streaming), WER on accented speech, and TTS naturalness all trade off non-linearly. In our voice pipeline, DeepSeek-V4-Pro plus a small dedicated STT beat an end-to-end frontier voice model on cost-per-minute by ~10x while staying inside a 300ms added-latency budget - purely because we could mix and match components. The hard part is honest benchmarking: WER numbers are only comparable within the same eval set, so a "find the optimal combo" service lives or dies by its methodology. Do you expose per-component benchmarks (STT WER, TTS MOS) separately, or only the combined scores?