3 ms·
Since voice-to-text has gotten so good I've used it a lot more and also noticed how distracting and confusing it can be. Using Apple's dictation has a similar f
by mavsman 3y ago
Since voice-to-text has gotten so good I've used it a lot more and also noticed how distracting and confusing it can be. Using Apple's dictation has a similar feel to this where you're constantly seeing something that's changing on the screen. It's kind of irritating and I don't really know what the solution is.
One suggestion I have here is to have at least two different sections of the UI. One part would be the actual document and the other would be the scratchpad. It seems like much of what you say would not actually make it into the document (edits, corrections, etc) so those would only be shown in the scratchpad. Once the editor has processed the text from the scratchpad then it can go into the document how it's supposed to. Having text immediately show up in the document as it's dictated is weird.
Your big challenge right now is just that STT is still relatively slow for this usecase. Time will be on your side in that regard as I'm sure you know.
Good luck! Voice is the future of a lot of the interactions we have with computers.
- codercowmoo 3y agoDistil-whisper is incredibly fast. Realtime on a 3060 Ti, and I used it to transcribe an 11 hour audiobook in 9 minutes.
- peddling-brink 3y agoYou know, those audiobooks already have transcriptions. Often written by the original author! I kid. Your comment made me think of a shower thought I had recently where I wished my audiobook had subtitles.
- jazzyjackson 3y agoIt really is a little absurd IMO that the text of the book is sold separately from the audio.
- chrisaiv 3y agoBook publishing industry is different from audio recording industry.
- robbomacrae 3y agoNot trying to hijack this. Great demo! But STT can be very much real-time now. Try SoundHound's transcription service available through the Houndify platform [0] (we really don't market this well enough). It's lightning fast and it's half of what powers the Dynamic Interaction demos that we've been putting out. I actually made a demo just like this aqua voice internally (unfortunately didn't get prioritized) but there is really no lag. However it will always be the case where the model will want to "revisit" transcribed words based on what comes next. So if you want the best accuracy you do want to wait a sec or two for the transcription to settle down a bit. [0]: https://www.houndify.com https://www.houndify.com