5 ms·
I've been looking for a tool that can passively transcribe audio files to make them easier to search - this looks like it could almost solve that use case - may
by AlphaWeaver 3y ago
I've been looking for a tool that can passively transcribe audio files to make them easier to search - this looks like it could almost solve that use case - maybe with a scripted tagger.
- calciphus 3y agoCheck out Whisper, it is surprisingly good and several of the output formats include time codes, works with multiple languages. https://github.com/openai/whisper https://github.com/openai/whisper
- yorwba 3y agoWhisper is now outclassed by Facebook's MMS https://github.com/facebookresearch/fairseq/tree/main/examples/mms https://github.com/facebookresearch/fairseq/tree/main/exampl... though the integration work to make it a turnkey solution for a wider audience hasn't happened yet. (E.g. https://github.com/ggerganov/whisper.cpp/issues/950 https://github.com/ggerganov/whisper.cpp/issues/950 is still open.)
- angrais 3y agoHave you tested MMS with real-world data? It's perhaps outclassed on evaluation metrics but on real-world data it is not as good as whisper.
- ck_one 3y agoHave you tried MMS on real world data or is it just assumption?
- angrais 3y agoYes, of course. Real world data being: one on one interviews (no background noise), small groups of people chatting (lots of background noise), and specific audio recordings ( with varying British regional accents. In all three instances whisper produced a more accurate transcription. This is for personal use. The license of MMS is also restrictive so cannot he used for commercial uses while whisper can. Another key consideration when wondering what to choose. On the other hand, one can train MMS (so using own custom dataset) so for some projects it may be more suitable.
- sandreas 3y agoThat is funny. For audio books I'm currently working on an `epub` command for `tone` which will be able to extract text from `epub` files, e.g.: tone epub --format="markdown" --extract-sentences --one-file-per-chapter output-path/ As a result, you can use https://github.com/readbeyond/aeneas https://github.com/readbeyond/aeneas with the generated text / markdown files to create a json mapping file looking like this: { "fragments": [ { "begin": "0.000", "children": [], "end": "7.920", "id": "f000001", "language": "eng", "lines": [ "This is the first sentence of the audio book." ] } } Since aeneas is a bit inaccurate, I'm also working on an improvement with silence detection for these mapping files. If you are looking for something that is "ready to use", you could check out https://github.com/r4victor/syncabook https://github.com/r4victor/syncabook or the according library https://github.com/r4victor/afaligner https://github.com/r4victor/afaligner If you have audio files, that are NOT audio books, the epub approach will not help you and the other comments are more helpful.
- thangalin 3y agoSee my comment showing how to transcribe using Whisper: https://news.ycombinator.com/item?id=35366778 https://news.ycombinator.com/item?id=35366778