9 ms·
I am not sure if "Word Error Rate" captures what has always been wrong with transcription. My biggest complaint is that it inserts sentence breaks in random pla
by jeffbee 1mo ago
I am not sure if "Word Error Rate" captures what has always been wrong with transcription. My biggest complaint is that it inserts sentence breaks in random places, then fails to evaluate the result, even though it is obviously wrong. Then I have to go fix it which can be harder than having just typed it myself, due to the difficulty of positioning the Android cursor, the fact that it automatically capitalizes if you delete a capital letter, etc. And much of the time I fail to notice the errors until later.
- verdverm 1mo agohave another model do a pass to clean it up, saw a demo of local STT where someone did this, can fix a lot of things, especially with gotchas for the STT model in a clean-transcript.md
- jeffbee 1mo agoI think the model can even evaluate itself. If it looks afterward at an output like "do you. Want to get lunch?" in the absence of affirmative evidence that the user wanted it that way, it should be able to see that it goofed.
- verdverm 1mo agoIt's typical to use a special STT model (audio in only), which will not be able to clean up afterwards. If you are using an LLM for the STT part, you're leaving stuff on the table
- coder543 1mo agoI haven't tried it, but this looked promising for that exact task: https://huggingface.co/superwhisper/s1-mini https://huggingface.co/superwhisper/s1-mini