4 ms·
I recently swapped out the AI model for voice transcription on revoldiv.com and replaced it with Whisper. The results have been truly impressive - even the smal
by userhacker 4y ago
I recently swapped out the AI model for voice transcription on revoldiv.com and replaced it with Whisper. The results have been truly impressive - even the smaller models outperform and generalize better than any other options on the market. If you want to give it a try, our model is capable of faster transcription by utilizing multiple GPUs and some other enhancements, and it is all free
- thundergolfer 4y agoYour app is very similar to our demo app! https://modal-labs-whisper-pod-transcriber-fastapi-app.modal.run/#/ https://modal-labs-whisper-pod-transcriber-fastapi-app.modal... How come you don't support audio files longer than 1hr? Is it because of $$ cost? The above demo app gets faster transcription by chunking audio and parellelizing over dozens of CPUs, so you can transcribe a about 1hr of audio for $0.10.
- userhacker 4y ago> https://modal-labs-whisper-pod-transcriber-fastapi-app.modal https://modal-labs-whisper-pod-transcriber-fastapi-app.modal... Interesting, which model are you using? We use the medium model which is the sweet spot between time/performance ratio. We also chunk, We try to detect words and silences to do better chunking at word boundaries but if you do more chunking and you don't get the word boundaries right it seems like whisper loses some context and the accuracy suffers. We will soon support longer hours. We just want to make sure the wait time for transcription doesn't suffer for most users. But great demo, reach out to me if you want to collaborate
- thundergolfer 4y agoWe’re using base. The code is open source at modal-labs/model-examples repo if you want to see anything we’re doing
- 71a54xd 4y agoI'd love to try it out, who should I ping to run this locally on my GPU server?
- userhacker 4y agoIt's not a model you can run on your own server but a free service on revoldiv.com. You can expect 40 to 50 second wait time to transcribe an hour long video/audio. We combine whisper with our model to get word level timestamps, paragraph separation and sound detections like laughter, music etc... We recently added very basic podcast search and transcription.
- getcrunk 4y agoIm getting an unknown error, errro code 551