17 ms·
MacWhisper: Transcribe audio files on your Mac
- mkmk 3y agoI’ve gotten confused between the different whispers. How is this different from the openai api endpoint?
- miki123211 3y agoIt runs locally, using Whisper.cpp[1], a Whisper implementation optimized to run on CPU, especially Apple Silicon. Whisper itself is open source, and so is that implementation, the OpenAI endpoint is merely a convenience to those who don't wish to host a Whisper server themselves, deal with batching, renting GPUs etc. If you're making a commercial service based on Whisper, the API might be worth it for the convenience, but if you're running it personally and have a good enough machine (an M1 MacBook Air will do), running it locally is usually better. [1] https://github.com/ggerganov/whisper.cpp https://github.com/ggerganov/whisper.cpp
- smoldesu 3y agoFWIW, I will add that most laptops made in the past 10 years are fast enough for real-time transcription. Unless you're trying to transcribe in bulk, running it locally will usually be the best option.
- miki123211 3y agoThis depends on the model being used, if you're doing anything which isn't English, you pretty much need large, and that needs considerable resources.
- ajhai 3y agoShameless plug: recently launched LLMStack (https://github.com/trypromptly/LLMStack https://github.com/trypromptly/LLMStack) and I have some custom pipelines built as apps on LLMStack that I use to transcribe and translate. Granted my use cases are not high volume or frequent but being able to take output from Whisper and pipe it to other models has been very powerful for me. It is also amazing how good the quality of Whisper is when handling non English audio. We added LocalAI (https://localai.io https://localai.io) support to LLMStack in the last release. Will try to use whisper.cpp and see how that compares for my use cases.
- Flimm 3y agoIf you're looking for an alternative that runs on Linux, I just recently discovered Speech Note. It does speech to text, text to speech, and machine translation, all offline, with a GUI: https://flathub.org/apps/net.mkiol.SpeechNote https://flathub.org/apps/net.mkiol.SpeechNote https://github.com/mkiol/dsnote https://github.com/mkiol/dsnote
- circularfoyers 3y agoI must have missed this on my usual peruse of new apps on Flathub. Thanks for making it, will look forward to trying it out.
- bilater 3y agoIf you'd rather use a web app with minimal cost upfront check out PlainScribe :) https://www.plainscribe.com/ https://www.plainscribe.com/
- uger 3y agoGreat tool, but I can't wait until it can do real-time live transcribing.
- nafizh 3y agoThe main problem I have faced with the whisper model (large) is when there is silence or a sizable gap without audio, it hallucinates and just puts out some random gibberish repeatedly until the transcription ends. How does this app handle this?
- _tom_ 3y agoI've run into that, many times. Would be nice to have a fix.
- userhacker 3y agoIf you want a quick and free web transcription and editor tool, We've built https://revoldiv.com/ https://revoldiv.com/ with speaker detection and timestamps. Takes less than a minute to transcribe 1 hour long video/audio
- beardedwizard 3y agoYes but the point of this project is that it doesn't require you to share sensitive data with third parties.
- userhacker 3y agoGood point but the problem with local hosting is that if you want to use the larger models it will take a long time to transcribe a file. We use multiple gpus and we do speaker detection, sound detection and it is has a rich audio editor.
- beardedwizard 3y agoTotally agree, having built a similar app I know speaker diarization is a killer feature that's hard to get. My problem is I'll never share these recordings ;).
- deegles 3y agoIs gumroad a good platform for selling software like this? How is licensing handled?
- beardedwizard 3y agoSeems shady to me to charge for running larger free models you don't provide on hardware your users provide. You are charging for openAi features not yours.
- firloop 3y agoThey're not charging for the model, they're charging for the UI.
- beardedwizard 3y agoThe UI is free, the premium features are the model. Read the website.
- michelb 3y agoDid you even look at the app?
- beardedwizard 3y agoYes now let me help you read the paid feature: Supports Tiny (English Only), Tiny, Base, Small, Medium and Large models Translate audio file into another language through Whisper (use the Medium or Large models, the results will not be perfect and I'm working on more advanced ways to do this)
- satvikpendem 3y agoWhile whisper.cpp is faster than faster-whisper on macOS due to Apple's Neural Engine [0], if you have a GPU on Windows or Linux, faster-whisper [1] is a lot faster than OpenAI's reference Whisper implementation as well as whisper.cpp, with the CLI being wscribe or whisper-ctranslate2 as faster-whisper is only a Python library. It's pretty good. [0] https://github.com/guillaumekln/faster-whisper/discussions/368 https://github.com/guillaumekln/faster-whisper/discussions/3... [1] https://github.com/guillaumekln/faster-whisper https://github.com/guillaumekln/faster-whisper
- bonney_io 3y agoAny insight on how Whisper works on older Intel Macs? I have a 2012 Mac mini with 16GB of RAM doing nothing; if I could use it to (slowly) transcribe media in the background, this becomes a must-buy.
- nchudleigh 3y agoNot well with Intel Macs unfortunately.
- jbverschoor 3y agoSo this is not Whisper Transcription 4 from the appstore?
- ycstohley 3y agoSeriously great program. Licensing model just fine. I use this all the time, so do my collegues at other companies. The developer Jordi has a great speech online about product development.
- googlryas 3y agoOut of curiosity, does anyone know what the state of the art for transcription is? Is there a possibility it will soon be "better than a person carefully listening and manually transcribing"? I ask because I asked a friend to record a (for fun) lecture I couldn't attend, and unfortunately the speech audio levels are quite low, and I'm trying to figure out how to extract as much info as possible so I can hear it. If I could add context to the transcriber like "This is about the Bronze Age collapse and uses terminology commonly used in discussions on that topic", it might be even more useful.
- userhacker 3y agoTry to upload it on https://revoldiv.com/ https://revoldiv.com/ we pre-process the file to make it a little Intelligible and you can supply your context when uploading.
- neocodesoftware 3y agohttps://github.com/chidiwilliams/buzz https://github.com/chidiwilliams/buzz Brew install buzz Its great
- 8f2ab37a-ed6c 3y agoLove the idea behind this. High quality transcription + the data not leaving your device is excellent. Any chance there's an iOS version of this coming down the pike? It would be great to have a voice-based note taking app that you can use when you are driving or walking and you don't want to type into your phone, but you just want to save that thought you just had somewhere by quickly dictating it, and having it accessible as text later.
- saYu 3y ago[dead]
- ZoomerCretin 3y agoA few weeks ago I found myself wanting a speech to text transcriber that directly captures my computer's audio output (I.e. not mic input, not am audio file), but I could not find one. The best alternative I found was to have my computer direct audio output to a virtual audio input device, but I could not do this on my desktop because I do not have a sound card. I found software that did this, but it did not allow me to listen to the audio output while it was redirected to a virtual audio input. Has anyone else tried to do something similar? How did you achieve it?
- spectre3d 3y agoAudio Hijack[1] will let you route any audio to multiple virtual or actual outputs while adding the ability to listen to any part of the signal chain. Hope that solves it for you, it’s saved my sanity a number of times! [1] https://rogueamoeba.com/audiohijack/ https://rogueamoeba.com/audiohijack/
- ukuina 3y agoMany such apps exist. I use Hello Transcribe from the App Store, $7 across all iDevices, with CoreML optimization.
- not_the_fda 3y agoIs this just a front end to OpanAI's whisper? https://github.com/openai/whisper https://github.com/openai/whisper
- idorosen 3y agoWhy? Just use whisper directly. The model and code is available and I think there’s even a homebrew formula...
- michelb 3y agoI just want to drag and drop my files and be done.
- nickthegreek 3y agoWhy? Just use MacWhisper and have a great interface with a bunch of options.
- speedgoose 3y agoI have both installed. I use macwhisper because the GUI is convenient.
- deleted 3y ago[deleted]
- Etheryte 3y agoWhy use a web browser? Just use curl directly. The code is available and I think there's even a homebrew formula...
- idorosen 3y agoI deserved that curl snark. :o) Web browsers are mostly free and don't try to upsell you to a Pro paid version. The MacWhisper author deserves to be compensated for their work, so I'm not objecting to the existence of a paid version. This feels like yet another relatively low value freemium/upsell wrapper in the Mac shareware ecosystem to me. I'm probably wrong and there's a real population that benefits from this work, clearly some folks perceive it as useful enough to pay for it and I'm just not in that audience to see it. I think part of what rubs me the wrong way about this is that it feels to me like commercial freeloading due to the thinness of the commercialized wrapper around a free/open core in this case (whisper model + code); it feels ethically questionable unless the author contributes back some portion of the proceeds to research in some way -- I didn't see evidence of that. I'm probably being naive here, happy to have a less snarky discussion about it though.
- shawnc 3y agoBeen using it for a couple months, and Jordi keeps improving on it at a steady clip. It's great!!
- pgt 3y agoAnyone have a cached page? Seems to hugged to death.
- deleted 3y ago[deleted]
- patrick91 3y agoI really like this app, I wish there was a way to play a video while editing the subtitles though!
- paulmd 3y agoWhisper is cool. Back in college I wanted to do some projects with speech-to-text and text-to-speech as an interface like 10-12 years ago, but at that point the only option was google APIs that charged by the word or second. On top of that, constantly sending data to google would have chewed a ton of battery compared to the "activation word" style solutions ("ok google/siri") that can be done on-device. The power for on-device processing was obviously going to come down over time, while wireless is much more governed by the laws of physics, and connectivity power budgets haven't gone down nearly as much over time. I am pretty sure there is a fundamental asymptotic limit for this, governed by Shannon entropy limit/channel width and power output. In the presence of a noise floor of X, for a bandwidth of Y, you simply cannot use less than Z total power for moving a given amount of data. BTLE is really the first game-changer (especially if you are hooking into a broad network of receivers like apple does with airtags) but even then you are not really breaking this rule - you are just transmitting less often, and sending less data. It's just a different spot on the curve that happens to be useful for IOT. If you are, say, doing a keyboard over BTLE where the duty cycle is higher, the power will be too. Applications that need "100% duty cycle"/"interactive" (reachable at any time with minimal latency") still have not improved very much. In hindsight I guess the answer would have been writing a mobile app that ties into google/siri keywords and actions, and letting the phone be the UI and only transmit BT/BTLE to the device. But BTLE hadn't hit the scene back then (or at least not nearly to the extent it has now) and I was less experienced/less aware of that solution sapce.
- kulesh 3y agosuperwhisper.com is also cool
- nchudleigh 3y agoThanks!
- ukuina 3y ago$165!
- davidf18 3y ago[dead]
- _rs 3y agoI've used this for a few months to transcribe interviews and it works pretty well. The UI for dealing with multiple speakers is a bit cumbersome, and there are occasional crashes, but overall definitely a great app and worth the money
- miki123211 3y agoThis basically does the same thing but free: https://apps.apple.com/us/app/aiko/id1672085276 https://apps.apple.com/us/app/aiko/id1672085276
- Nezteb 3y agoThat's awesome that the dev released Aiko for free! Not a deal breaker, but it was last updated 3 months ago and lacks a few QoL features of MacWhisper. Jordi is frequently pushing updates to MacWhisper: https://nitter.net/jordibruin/status/1692133387299864638 https://nitter.net/jordibruin/status/1692133387299864638
- michelb 3y agoHmm that one has a lot less features.
- washadjeffmad 3y agoNothing that isn't scriptable with existing projects, tbh
- agentdrtran 3y agoDoes anyone know of an easy to use whisper fork with speaker attestation?
- masukomi 3y agoHere's a multi-platform open source app that does the same thing but uses vosk instead of whisper. https://github.com/bugbakery/audapolis https://github.com/bugbakery/audapolis
- MaxikCZ 3y agoWould be nice if it allowed importing mkv files, in the end its just a container..
- deleted 3y ago[deleted]
- dagaci 3y agoThe OpenAi CLI does that, follow the instructions https://github.com/openai/whisper https://github.com/openai/whisper
- tornato7 3y agoI have a Python script on my mac that detects when I press-and-hold the right option key, and records audio while it's pressed. On release, it transcribes it with whispercpp and pastes it. Makes it very easy to record quick voice notes. Here it is: https://github.com/corlinp/whisperer/tree/whisper.cpp https://github.com/corlinp/whisperer/tree/whisper.cpp I was working on a native version in the form of a taskbar app with customizable prompt and all. However I quickly realized that the behaviors I want the app to do require a bunch of accessibility permissions that would block it from the app store and require more setup steps. Would anybody still find that useful?
- rpastuszak 3y ago> However I quickly realized that the behaviors I want the app to do require a bunch of accessibility permissions Which behaviours specifically? Personally, I wouldn't worry too much about the App Store. I'm distributing Enso (http://enso.sonnet.io http://enso.sonnet.io) via gumroad.com, and people download/pay for it. I think it's easier than using the App Store Connect route anyway. Here's a good intro: https://rambo.codes/posts/2021-01-08-distributing-mac-apps-outside-the-app-store https://rambo.codes/posts/2021-01-08-distributing-mac-apps-o...
- tornato7 3y agoDetecting an alt-key push even when it's not an active window, and editing the selected text field are both accessibility permissions. Thanks for the info about your app. It looks great!
- alin23 3y agoEditing the field definitely needs the permissions, but detecting Alt-key holding should not. You can do that using something like: var reactOnOptionKeyHeld: DispatchWorkItem? { didSet { oldValue?.cancel() } } NSEvent.addGlobalMonitorForEvents(matching: .flagsChanged) { (event) in guard event.modifierFlags == [.option] else { reactOnOptionKeyHeld = nil return } reactOnOptionKeyHeld = DispatchWorkItem { // start recording } // schedule to run if held for at least 1 second DispatchQueue.main.asyncAfter(deadline: .now() + 1, execute: reactOnOptionKeyHeld!) } I see you're using Python with pynput though, which is creating a full key listener so I guess that is why you need the permissions.
- holdodd 3y agohttps://github.com/MahmoudAshraf97/whisper-diarization https://github.com/MahmoudAshraf97/whisper-diarization This project has been alright for transcribing audio with speaker diarization. A big finicky. The OpenAI model is better than other paid products(Descript, Riverside) so I’m looking forward to trying MacWhisper.
- zitterbewegung 3y agoThere is a great library that has support not only with OpenAIs whisper but many others that also work offline. https://github.com/Uberi/speech_recognition https://github.com/Uberi/speech_recognition
- simonw 3y agoI've been using MacWhisper for a few months, it's fantastic. Sometimes I'll send a mp3 or mp4 video through it and use the resulting transcript directly. Other times I'll run a second step through https://claude.ai/ https://claude.ai/ (because of its 100,000 token context) to clean it up. My prompt for that at the moment is: > Reformat this transcript into paragraphs and sentences, fix the capitalization and make very light edits such as removing ums That's often not necessary with Whisper output. It's great for if you extract captions directly from YouTube though - I wrote more about that here: https://simonwillison.net/2023/Aug/6/annotated-presentations/ https://simonwillison.net/2023/Aug/6/annotated-presentations...
- rpastuszak 3y agoThis is so good! I studied English, then moved to linguistics, then lived in the UK for almost a decade and due to my accent none of the TTS tools are close to the approach you just mentioned (whisper + LLM). Thanks Simon!
- mosselman 3y agoI didn’t know whisper could differentiate voices for the per speaker transcription. Is that new? Is it also available in the command line whisper builds?
- wahnfrieden 3y agoIt can’t