8 ms·
OTranscribe: A free and open tool for transcribing audio interviews
- nullbar 2y agoMaybe it isn't perfectly clear, but OTranscribe isn't an automatic speech-to-text tool, but instead, a UI for assisting in manual transcribing. So no AI here, folks.
- space_oddity 2y agoYep, it's designed to assist with manual transcription
- kimoz 2y agoAnyone knows a free tool for generating subtitles for movies and series videos ?
- doug_life 2y agohttps://github.com/McCloudS/subgen https://github.com/McCloudS/subgen worked very well for me. I had a TV series where somehow the last few seasons timestamps did not match up with subtitle files I could find online. I used subgen and it worked surprisingly well.
- BrunoJo 2y agoYou can try https://www.transcripo.com/ https://www.transcripo.com/ for free
- drtgh 2y agoSubtitleEdit is one of the most complete and has many online tutorials from users. Make sure they are recent tutorials because they will probably mention how to use the automated generation tools/plugins that wasn't available years ago. https://github.com/SubtitleEdit/subtitleedit https://github.com/SubtitleEdit/subtitleedit
- jagermo 2y agofantastic tool; I used it a lot to transcribe interviews during plane travels where there was no internet, and I needed to fill the time. Really useful to have if you do a lot of interviews
- dotancohen 2y agoFrom the homepage: > A free web app to take the pain out of transcribing recorded interviews How did you use a web app on the plane with no internet?
- Havoc 2y agoIt’s MIT licensed so presumably self hosted
- grandfunction 2y agoRan the server on his or her laptop... You don't need the internet to use a web browser
- tampueroc 2y agoThe web app saves an offline copy for use the first time you open it. https://otranscribe.com/help/#can_i_use_otranscribe_offline https://otranscribe.com/help/#can_i_use_otranscribe_offline
- jagermo 2y agoit works offline if you preload the website :)
- TrojanHookworm 2y agoUse this a lot. It's nice and simple and has exactly the tools you need (playback speed control, easy pause/play) and nothing more. Greatly prefer it over automatic transcription tools give you 40 pages of 'umm's and 'ahhhh's to filter through and edit.
- stavros 2y agoCan you not give the transcript to an LLM to remove the umms and ahhs?
- BiteCode_dev 2y agoPeople not used to AI have blind spots that prevent them from seing evident use case like this. I'm always surprised at the amazed look of my friends when they see me concretely use the tool. They just didn't picture it until they saw it in action.
- stavros 2y agoIt's not even people not used to AI, I developed a tool that uses AI to do something, and then kind of couldn't be bothered to fix some of the output manually. It only occurred to me days later that I can ask the AI to fix it.
- phoronixrly 2y agohttps://github.com/oTranscribe/oTranscribe https://github.com/oTranscribe/oTranscribe
- cube2222 2y agoI needed to do this this week (transcribe an interview with multiple speakers) and used https://github.com/MahmoudAshraf97/whisper-diarization https://github.com/MahmoudAshraf97/whisper-diarization Worked excellent. It generates both a file that just contains a line per uninterrupted speaker speech prefixed with the speaker number, as well as a file with timestamps which I believe would be used as subtitles.
- RamblingCTO 2y agoI had better success with whisperx, as whisper-dia does sometimes have weird issues I couldn't resolve: https://github.com/m-bain/whisperX https://github.com/m-bain/whisperX
- cube2222 2y agoiirc whisper-diarization uses whisperx under the hood. I’ll be honest, I haven’t dived much into this as I just needed something transcribed quickly, but when I was looking at WhisperX I couldn’t find a CLI that would just out of the box give me a text file with a line per speaker statement (not per word).
- RamblingCTO 2y agoI use it like this: whisperx $file int8 --min_speakers 3 --max_speakers 3 --language de --hf_token $token --diarize
- stavros 2y ago> iirc whisper-diarization uses whisperx under the hood. It seems like it does: https://github.com/MahmoudAshraf97/whisper-diarization/blob/main/requirements.txt https://github.com/MahmoudAshraf97/whisper-diarization/blob/...
- adipasquale 2y agoI have had very good results using Spectropic [1], a hosted Whisper Diarization API service as a platform. I found it cheap and way easier and faster than setting up and using whisper-diarization on my M1. Audiogest [2] is a web service built upon Spectropic, I have not yet used it. disclaimer : I am not affiliated in any way, just a happy customer! I had some nice mail exchanges after bug reports with the (I believe solo-)developer behind these tools. --- [1] https://spectropic.ai/ https://spectropic.ai/ [2] https://audiogest.app/ https://audiogest.app/
- choya-love 2y agoAny new language support in the future? Fingers crossed for japanese
- fabianmg 2y agoAm I missing something?. For what I checked it supports every language, as is yourself the one transcribing by hand. This is just an UI to watch the video or audio while you're typing it.
- comradesmith 2y agohttps://tactiq.io https://tactiq.io is made for meetings, but also does uploaded transcripts and supports Japanese!
- ilt 2y agoI currently use Aiko’s free iOS app which does offline transcription using OpenAI’s Whisper model. It has been working pretty well for me so far. It can export in formats like SRT, TXT, CSV, JSON and text with timestamps too. https://sindresorhus.com/aiko https://sindresorhus.com/aiko
- deleted 2y ago[deleted]
- left-struck 2y ago[dead]
- BetterWhisper 2y agoIf you are looking for something automatic that also allows you to interact with your transcripts chatgpt style then I would recommend https://www.videototextai.com/ https://www.videototextai.com/
- Terretta 2y agoThat cookies box though... Dark pattern (accept lots + accept all, fake drag affordance, covering a quarter of the page) for cookies doesn't bode well for privacy protections around the transcripts.
- BetterWhisper 2y agoYou are allowed to delete any transcription you make and with that we do not keep any copy of the transcripts :) . The cookie banner is there to comply with the EU laws.
- teddyh 2y agoSee also TranscriberAG: <https://transag.sourceforge.net/ https://transag.sourceforge.net/>
- dmitrykan 2y agoI'm working on the tool, that includes AI. My original target is to test it on my https://www.youtube.com/c/VectorPodcast https://www.youtube.com/c/VectorPodcast by offering something that Lex Fridman does for his episodes. Current features: 1. Download from YT 2. Transcribe using Vosk (output has time codes included) 3. Speaker diarization using pyannote - this isn't perfect and needs a bit more ironing out. What needs to be done: 4. Store the transcription in a search engine (can include vectors) 5. Implement a webapp If anyone here is interested to join forces, let me know.
- ciaran00 2y agoTalio.ai allows you to do this with chatGPT style chat with the transcript plus numerous other features https://talio.ai https://talio.ai
- jrochkind1 2y agoKinda surprised to not have AI integration. You do still need to proof and QA even AI results, if you want a publication quality result, and do things like attribute who is speaking when (at least Whisper can't do that), and correct "unusual" last names and things. So I feel like people using AI still need good tools for the correcting/finishing/proofing too, that would be similar to the tools for non-assisted transcription.
- MattieTK 2y agoThis was written a really long time ago by a former WSJ Graphics reporter (Elliot Bentley) who is now at Datawrapper. It is now operated by Muckrock and hasn't seen changes made to it in a while. That's why it doesn't have any of these integrations, the technology just didn't exist.
- jrochkind1 2y agoAha, good to know! That's actually important context, that this is not a recent release, and doesn't necessarily have a lot of ongoing development.
- bcherny 2y agoLooks cool! Unclear from the docs, but does it support non-English languages? How about mixed-language interviews?
- avodonosov 2y agoYes! Any language you understand is supported!
- avodonosov 2y agoI made a similar tool for making tables of contents for youtube videos: https://youtoc.by/ https://youtoc.by/ Not developing it actively after I created tables of contents for the several videos I needed, years ago. If I ever need it again, I will probably work on mobile UI (aka responsive)
- tkgally 2y agoI was curious how good a transcription I could get from what may be the best multimoldal LLM currently, Gemini-1.5-Pro-Experiment-0801, so I had it transcribe five minutes of an interview between Ezra Klein and Nancy Pelosi from earlier today. The results are here: https://www.gally.net/temp/20240809geminitranscription/index.html https://www.gally.net/temp/20240809geminitranscription/index... Aside from some minor punctuation and capitalization issues, Gemini’s transcription looks nearly perfect to me. There were only one or two words that I think it misheard. If I had transcribed the audio myself, I would have made more mistakes than that. One passage struck me in particular: And then he comes up with "weird," which becomes viral and the rest, and here he is. How did Gemini know to put “weird” in quotation marks, to indicate—correctly—that the speaker was referring to Walz’s use of the word as a word? According to Politico, Walz first used the word in that context in the media on July 23. https://www.politico.com/news/2024/07/26/trump-vance-weird-00171470 https://www.politico.com/news/2024/07/26/trump-vance-weird-0...
- moritzwarhier 2y agoMaybe two factors helped achieve the impressive result with the quotation marks: - auditory cues - the sentence would be gramatically incorrect and make no sense without them Just guessing out of the blue. But I think it's likely that LLMs (and other speech recognition systems) need to exploit sentence context to recognize individual words and punctuation, and this is an example were it went well. Human listening is similar in a way, we can recognize words even when spoken very mumbly or fast, if we have context. So we always hear phrased rather than words.
- sebastiennight 2y agoIt's very likely that the model is capable of picking up on the verbal cues surrounding quotes. Do you have the audio or video file? I'd like to run it through our AI video editor and see how it punctuates the transcript.
- tkgally 2y agoThe mp3 file that I gave to Gemini (a five-minute excerpt from the audio podcast) is linked in the source code of the page. Here is the full URL: https://www.gally.net/temp/20240809geminitranscription/interview.mp3 https://www.gally.net/temp/20240809geminitranscription/inter... The full interview including video is on the New York Times website, though you might need a subscription to view it: https://www.nytimes.com/2024/08/09/opinion/ezra-klein-podcast-nancy-pelosi.html https://www.nytimes.com/2024/08/09/opinion/ezra-klein-podcas... The NYT’s closed captions do not put “weird” in quotation marks; they also divide sentences weirdly and have other mistakes. But they get some things better than Gemini, such as capitalizing “House” when it means the House of Representatives. I haven’t compared the audio-only podcast version and the video version carefully; it’s possible that parts of the audio were edited or re-recorded for one or the other. Let us know how your AI video editor does!
- deleted 2y ago[deleted]
- kgdiem 2y agoI started making an open source macOS app to do this with whisper and potentially pyannote. It is functional but a bit slow. I think using whisper directly instead of swift bindings will help a lot. Really interested in adding diarisation but having a lot of trouble converting Pyannote to CoreML. Pyannote runs so slowly with torch on CPU. Haven’t gotten around putting my latest work for that on GitHub yet. Happy to accept contributions — Some priorities right now: * Fixing signing for local builds * Replace swift whisper with whisper cpp * Allowing users to provide their own models https://github.com/Stack-Studio-Digital-Collective/Auditif https://github.com/Stack-Studio-Digital-Collective/Auditif
- vunderba 2y agoYeah definitely switch over to using ggerganov's whisper implementation, I use it in a little home brewed python app on my M1 for handling speech transcripts. The base EN model chews through minutes of audio in seconds, it's insanely fast.
- accidbuddy 2y agoAnyone knows one with transcription and translate in real time? Nowadays, I use libretranslate/libretranslate and pluja/whishper to do this, but not at real time.
- Bayko 2y agoAh this brings back memories. When I was in college with limited money, I used to pirate movies and most of them didn't have subtitles and I used to daydream of writing a VLC plug-in which would real time generate subtitles. But I had better things to do like play video games...
- space_oddity 2y agoMany of us have had those ambitious tech ideas...
- leiferik 2y agoYou're always welcome to try my service TurboScribe https://turboscribe.ai/ https://turboscribe.ai/ if you need a transcript of an audio/video file. It's 100% free up to 3 files per day (30 minutes per file) and the paid plan is unlimited and transcribes files up to 10 hours long each. It also supports speaker recognition, common export formats (TXT, DOCX, PDF, SRT, CSV), as well as some AI tools for working with your transcript.
- rsingel 2y agoThis looks great. Did you have an API or plan to release one?
- leiferik 2y agoThanks! Nothing to announce on the API front right now, but appreciate you asking :)
- windthrown 2y agoThanks! I've had good results with Turboscribe (paid plan) and appreciate having this as a service. I typically use it for 2-3 hour long video recordings with a number of speakers and appreciate the editing tools to clean things up before export.
- justinclift 2y agoFrom their FAQ: Does oTranscribe automatically convert audio into text? Sorry! It doesn’t. oTranscribe makes the manual task of transcribing audio a lot less painful. But you still have to do the transcription.
- btown 2y agoAre there any open-source or paid apps/shareware/freeware that can: - Transcribe word-by-word in real time as audio is recorded - Work entirely locally - Use relatively recent open-source local models? I've been using otter.ai for real-time meeting transcriptions - letting me multitask and instantly catch up if I'm asked a question by skimming the most recent few seconds worth of the transcript - but it's far from perfect and occasionally their real-time service has significant transcription delays, not to mention it requires internet connectivity. Most of the Whisper-based apps out there, though, as well as (when I last checked) the whisper.cpp demo code, require an entire recording to be ingested at once. There are others that rely on e.g. Apple's dictation frameworks, which is a bit dated in capability at the moment. Anything folks are using out there?
- uohzxela 2y agoI have built my own local-first solution to transcribe entirely locally in real time word by word, driven by a different need (I'm hard of hearing). It's my daily driver for transcribing meetings, interviews, etc. Because of its local-first capability, I do not have to worry about privacy concerns when transcribing meetings at work as all data stays on my machine. It's about as fast as Otter.ai although there's definitely room for improvements in terms of UX and speed. Caveat is that it only works on MacBooks with Apple silicon. Happy to chat over email (see my HN profile).
- WaitWaitWha 2y agoI have some staff with combined hearing and visual needs. Have you researched the one-, two- all-party consent requirements? Asking because I hope to identify transcription as "non-recording".
- CyberDildonics 2y agoWhat did your own research turn up?
- btown 2y agoCalifornia has an exception for hearing aids and other similar devices, but it’s unclear if transcription aids count, or if this has been tested in court. https://codes.findlaw.com/ca/penal-code/pen-sect-632/ https://codes.findlaw.com/ca/penal-code/pen-sect-632/ (Not a lawyer, this is not legal advice.)
- ulrischa 2y agoPretty amazing what a webapp an do. I whished there were more lile them and not all these native apps
- matejmecka 2y agoJust pitching in a transcription tool that lets you transcribe video and audio files using Whisper and WASM in your browser, and get a .txt, .srt, .vtt file. Maybe in the future support for Whisper Turbo? https://video2srt.ccextractor.org/ https://video2srt.ccextractor.org/ Disclaimer: Working on this project.
- neves 2y agoDoes anybody tested it with Brazilian Portuguese? It is a hard problem, since we have too many accents.
- dmd 2y agoI don't understand what the issue is. You don't know how to type the different diacritical marks? Or the textbox isn't accepting them? (Which seems like it would be a browser issue, not an issue with the site.)
- freeduhm 2y ago[dead]
- freeduhm 2y ago[dead]
- space_oddity 2y agooTranscribe is a free option for transcription but in many cases it's just too simple
- bilater 2y agoIf you just want quick transcriptions of YouTube video this works pretty well https://www.you-tldr.com/ https://www.you-tldr.com/
- ldenoue 2y agoYou can also try Scribe (free chrome extension and iOS app) https://www.appblit.com/scribe https://www.appblit.com/scribe