3 ms·
OpenAI released an app last month with press-to-talk that transcribes voice-to-text, but there's no in-built text-to-speech (use a screen reader I guess?) I th
by QuantumG 3y ago
OpenAI released an app last month with press-to-talk that transcribes voice-to-text, but there's no in-built text-to-speech (use a screen reader I guess?)
I think this a dumb architecture. The Whisper model has been released and runs well on an 8GB consumer GPU. Train a new head on it to produce speech until it exceeds the other voice-to-voice models, and then fine-tune it to banter instead of translate. Is that possible? Sure, but it's a pretty small model, so you wouldn't expect large LLM performance.