10 ms·
Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
I built this because I wanted to see how far I could get with a voice-to-text app that used 100% local models so no data left my computer. I've been using a ton for coding and emails. Experimenting with using it as a voice interface for my other agents too. 100% open-source MIT license, would love feedback, PRs, and ideas on where to take it.
- charlietran 6mo agoThank you for sharing, I appreciate the emphasis on local speed and privacy. As a current user of Hex (https://github.com/kitlangton/Hex https://github.com/kitlangton/Hex), which has similar goals, what are your thoughts on how they compare?
- ipsum2 6mo agoParakeet is significantly more accurate and faster than Whisper if it supports your language.
- yeutterg 6mo agoAre you running Parakeet with VoiceInk[0]? [0]: https://github.com/beingpax/VoiceInk https://github.com/beingpax/VoiceInk
- treetalker 6mo agoI have been using Parakeet with MacWhisper's hold-to-talk on a MacBook Neo and it's been awesome.
- rahimnathwani 6mo agoRight, and if you're on MacOS you can use it for free with Hex: https://github.com/kitlangton/Hex https://github.com/kitlangton/Hex
- lloyd-christmas 6mo agoOr write your own custom one with the library that backs it: https://github.com/FluidInference/FluidAudio https://github.com/FluidInference/FluidAudio I did that so that I could record my own inputs and finetune parakeet to make it accurate enough to skip post-processing.
- rahimnathwani 6mo agoThere's a fork of FluidAudio that supports the recent Cohere model: https://github.com/altic-dev/FluidAudio/tree/B/cohere-coreml-asr https://github.com/altic-dev/FluidAudio/tree/B/cohere-coreml... It's used by this dictation app: https://github.com/altic-dev/FluidVoice/ https://github.com/altic-dev/FluidVoice/
- obrajesse 6mo agoAnd indeed, Ghost Pepper supports parakeet v3
- totetsu 6mo agoParakeet supports japanese now, but I cant find a version ported to apple silicone yet.
- goodroot 6mo agoNice one! For Linux folks, I developed https://github.com/goodroot/hyprwhspr https://github.com/goodroot/hyprwhspr. On Linux, there's access to the latest Cohere Transcribe model and it works very, very well. Requires a GPU though. Larger local models generally shouldn't require a subordinate model for clean up. Have you compared WhisperKit to faster-whisper or similar? You might be able to run turbov3 successfully and negate the need for cleanup. Incidentally, waiting for Apple to blow this all up with native STT any day now. :)
- deleted 6mo ago[deleted]
- hephaes7us 6mo agoThanks for sharing! I was literally getting ready to build, essentially, this. Now it looks like I don't have to! Have you ever considered using a foot-pedal for PTT? Apple incidentally already has native STT, but for some reason they just don't use a decent model yet.
- goodroot 6mo agoThey do, and they even have that nice microphone F5 key for it, and an ideal OS level API making the input experience >perfect<. Apparently they do have a better model, they just haven't exposed it in their own OS yet! https://developer.apple.com/documentation/speech/bringing-advanced-speech-to-text-capabilities-to-your-app https://developer.apple.com/documentation/speech/bringing-ad... Wonder what's the hold up... For footpedal: Yes, conceptually it’s just another evdev-trigger source, assuming the pedal exposes usable key/button events. Otherwise we’d bridge it into the existing external control interface. Either way, hooks are there. :)
- jiehong 6mo agoThe only issue with Apple models is that they do not detect languages automatically, nor switch if you do between sentences. Parakeet does both just fine.
- 6mo ago
- konaraddi 6mo agoThat’s awesome! Do you know how it compares to Handy? Handy is open source and local only too. It’s been around a while and what I’ve been using. https://github.com/cjpais/handy https://github.com/cjpais/handy
- youniverse 6mo agoI love and have been using handy for a while too, what we need is this for mobile apps I don't think there's any free apps and native dictation is not always fully local and not as good.
- swaptr 6mo agoHandy is awesome! I used it for quite a while before Claude Code added voice support. Solid software, very good linux and mac integration. Shoutout to Parakeet models as well, extremely fast and solid models for their relatively modest memory requirements.
- stavros 6mo agoHandy is fantastic.
- vunderba 6mo agoI’d also be interested to know what the impetus was for developing ghost-pepper, which looks relatively recent, given that Handy exists and has been pretty well received. Extra bonus is that Handy lets add an automatic LLM post-processor. This is very handy for the Parakeet V3 model, which can sometimes have issues where it repeats words or makes recognition errors for example, duplicating the recognition of a single word a dozen dozen dozen dozen dozen dozen dozen dozen times.
- rob 6mo agoYep. Using Handy with Parakeet v3 + a custom coding-tailored prompt to post-process on my 2019 Intel Mac and it's been working great. Once in a while it will only output a literal space instead of the actual translation, but if I go into the 'history' page the translation is there for me to copy and paste manually. Maybe some pasting bug.
- mathis 6mo agoIf you don't feel like downloading a large model, you can also use `yap dictate`. Yap leverages the built-in models exposed though Speech.framework on macOS 26 (Tahoe). Project repo: https://github.com/finnvoor/yap https://github.com/finnvoor/yap
- hyperhello 6mo agoFeature request or beg: let me play a speech video and transcribe it for me.
- MattHart88 6mo agoI like this idea and it should work -- whatever microphone you have on should be able to hear the speaker. LMK if not (e.g., are you wearing headphones? if so, the mic can't hear the speaker)
- aristech 6mo agoGreat job. How about the supported languages? System languages gets recognised?
- MattHart88 6mo agoThanks! We currently have 2 multi-lingual options available: - Whisper small (multilingual) (~466 MB, supports many languages) - Parakeet v3 (25 languages) (~1.4 GB, supports 25 languages via FluidAudio)
- lostathome 6mo ago[dead]
- guzik 6mo agoSadly the app doesn't work. There is no popup asking for microphone permission. EDIT: I see there is an open issue for that on github
- ttul 6mo agoAnd many people are mailing in Codex and Claude Code generated PRs - myself included. Fingers crossed, I suppose.
- MattHart88 6mo agoThanks to everyone who submitted PRs! The fix is merged, new version is up.
- parhamn 6mo agoI see a lot of whisper stuff out there. Are these the same old OpenAI whispers or have they been updated heavily? I've been using parakeet v3 which is fantastic (and tiny). Confused why we're still seeing whisper out there, there's been a lot of development.
- zackify 6mo agosame, even have kokoro for speech back to text for home assistant and parakeet on mac os through voice ink. Also vibe coded a way to use parakeet from the same parakeet piper server on my grapheneos phone https://zach.codes/p/vibe-coding-a-wispr-clone-in-20-minutes https://zach.codes/p/vibe-coding-a-wispr-clone-in-20-minutes
- daemonologist 6mo agoWhisper is still old reliable - I find that it's less prone to hallucinations than newer models, easier to run (on AMD GPU, via whisper.cpp), and only ~2x slower than parakeet. I even bothered to "port" Parakeet to Nemo-less pytorch to run it on my GPU, and still went back to Whisper after a couple of days.
- goodroot 6mo agoWhisper is very good in many languages. It's also in many flavours, from tiny to turbo, and so can fit many system profiles. That's what makes it unique and hard to replace.
- 71bw 6mo agoI'm also wondering whether or not it would be beneficiary for my workload to switch over to Parakeet. Problem is, I'm using a lot of lingo - and in Polish, as well! - so it's not exactly the best case and whisper (v3), so far, works.
- gegtik 6mo agohow does this compare to macos built in siri TTS, in quality and in privacy?
- realityfactchex 6mo agoExactly my question. I double-tap the control button and macOS does native, local TTS dictation pretty well. (Similar to Keyboard > Enable Dictation setting on iOS.) The macOS built-in TTS (dictation) seems better than all the 3rd party, local apps I tried in the past that people raved about. I have tried several. Is this better somehow? If the 3rd party apps did streaming with typing in place and corrections within a reasonable window when they understand things better given more context, that would be cool. Theoretically, a custom model or UX could be "better" than what comes free built into macOS (more accurate or customizable). But when I contacted the developer of my favorite one they said that would be pretty hard to implement due to having to go back and make corrections in the active field, etc. I assume streaming STT in these utilities for Mac will get better at some point, but I haven't seen it yet (been waiting). It seems these tools generally are not streaming, e.g. they want you to finish speaking first before showing you anything. Which doesn't work for me when I'm dictating. I want to see what I've been saying lately, to jog my memory about what I've just said and help guide the next thing I'm about to say. I certainly don't want to split my attention by manually toggling the control (whether PTT or not) periodically to indicate "ok, you can render what I just said now". I guess "hold-to-talk" tools are for delivering discrete, fully formed messages, not for longer, running dictation. AFAICT, TFA is focused on hold-to-talk as the differentiator, over double-tap to begin speaking and double-tap to end speaking?
- realityfactchex 6mo agos/TTS/STT/
- Supercompressor 6mo agoI've been looking for the opposite - wanting to dump text and it be read to me, coherently. Anyone have good recommendations?
- realityfactchex 6mo agoSure, Chatterbox TTS Server is rather high quality: https://github.com/devnen/Chatterbox-TTS-Server https://github.com/devnen/Chatterbox-TTS-Server You could hook it up to some workflow over the local API depending on how you want to dump the text, but the web UI is good too. The Show HN by the author was at: https://news.ycombinator.com/item?id=44145564 https://news.ycombinator.com/item?id=44145564
- Supercompressor 6mo agoAppreciated - thank you.
- ericmcer 6mo agoI see quite a few of these, the killer feature to me will be one that fine tunes the model based on your own voice. E.G. if your name is `Donold` (pronounced like Donald) there is not a transcription model in existence that will transcribe your name correctly. That means forget inputting your name or email ever, it will never output it correctly. Combine that with any subtleties of speech you have, or industry jargon you frequently use and you will have a much more useful tool. We have a ton of options for "predict the most common word that matches this audio data" but I haven't found any "predict MY most common word" setups.
- MattHart88 6mo agoI've found the "corrections" feature works well for most of the jargon and misspelling use cases. Can you give it a try and let me know edge cases?
- sorenjan 6mo agoWhisper supports a prompt, you can put your "Donold" there. https://developers.openai.com/cookbook/examples/whisper_prompting_guide https://developers.openai.com/cookbook/examples/whisper_prom...
- bonkler59 6mo agoMy experience is that Aqua voice does a good job of this with custom dictionary and replacements.
- __mharrison__ 6mo agoCool, I've been doing a lot of "coding" (and other typing tasks) recently by tapping a button on my Stream Deck. It starts recording me until I tap it again. At which point, it transcribes the recording and plops it into the paste buffer. The button next to it pastes when I press it. If I press it again, it hits the enter command. You can get a lot done with two buttons.
- coldfoundry 6mo agoThis is exactly what I am building right now, Stream Deck with two buttons too (push to talk and enter)! It's a sweet little pet project, and has been a blast to build so far. Excited to finally add it to my workflow once its working well.
- purplehat_ 6mo agoHi Matt, there's lots of speech-to-text programs out there with varying levels of quality. 100% local is admirable but it's always a tradeoff and users have to decide for themselves what's worth it. Would you consider making available a video showing someone using the app?
- semiquaver 6mo agoSlop
- douglaswlance 6mo agodoes it input the text as soon as it hears it? or does it wait until the end?
- romeroej 6mo agoalways mac. when windows? why can you just make things multios
- naikrovek 6mo agoBecause like all other modern Macs, the GPU in my Mac uses the same API as the GPU in your Mac. Also, on a Mac with 32GB of RAM, 24GB of that (75%) is available to the GPU, and that makes the models run much faster. On my 64GB MacBook Pro, 48GB is available to the GPU. Have you priced an nvidia GPU with 48GB of RAM? It’s simply cheaper to do this on Macs. Macs are just better for getting started with this kind of thing.
- patja 6mo agoFair enough for GPU-intensive stuff like running Qwen locally. But do you really need a GPU for decent local TTS? I run parakeet just on CPU.
- patja 6mo agoI've been using Chirp which uses parakeet on Windows. Learned about it here: https://news.ycombinator.com/item?id=45930659 https://news.ycombinator.com/item?id=45930659 Works great for me!
- cootsnuck 6mo agoHandy has Windows support. https://handy.computer/ https://handy.computer/
- SquareWheel 6mo agoWindows has a native (cloud-based) dictation software built-in[1], so there's likely less demand for it. Nonetheless, there are still a handful of community options available to choose from. [1] https://support.microsoft.com/en-us/windows/use-voice-typing-to-talk-instead-of-type-on-your-pc-fec94565-c4bd-329d-e59a-af033fa5689f https://support.microsoft.com/en-us/windows/use-voice-typing...
- primaprashant 6mo agoSpeech-to-text has become integral part of my dev flow especially for dictating detailed prompts to LLMs and coding agents. I have collected the best open-source voice typing tools categorized by platform in this awesome-style GitHub repo. Hope you all find this useful! https://github.com/primaprashant/awesome-voice-typing https://github.com/primaprashant/awesome-voice-typing
- RobertTheNerd 6mo ago[dead]
- ArlenBales 6mo agoCan you explain how exactly dictation is used for development? I type about 120 WPM so typing is always going to be way faster for me than talking. Aside for accessibility, is dictation development for slower typers or is it more so you can relax on a couch while vibe coding? If this comes off as condescension it's not intended, I am genuinely out of the loop here.
- KerrickStaley 6mo agoI think most people can speak faster than 120 WPM. For example this site says I speak at 343 WPM https://www.typingmaster.com/speech-speed-test/ https://www.typingmaster.com/speech-speed-test/, and I self-measure 222 WPM on dense technical text.
- thakoppno 6mo agoMicro machines guy could be vibe coding at an absurd rate.
- mememememememo 6mo agoMy LLM types at 2k WPM. So I ise that to talk to my LLMs
- Zizizizz 6mo agoMost English speakers speak faster than 120 wpm so that's probably why people, especially those who can't type at speeds like you can, prefer it.
- rcarmo 6mo agoNot sure why I should use this instead of the baked-in OS dictation features (which I use almost daily--just double-tap the world key, and you're there). What's the advantage?
- qq66 6mo agoI haven't used this one but WisprFlow is vastly better than the built-in functionality on MacOS. Apple is way behind even startups, even for fundamental AI functionality like transcribing speech
- ibero 6mo agoWisprFlow has a lot of good recommendations behind it but the fact they used Delve for SOC2 compliance gives me major pause.
- janalsncm 6mo agoThe fact that a company could slurp up all of your data and then use Delve for their SOC2 is a great reason to use local models.
- deleted 6mo ago[deleted]
- jonwinstanley 6mo agoI use the baked in Apple transcription and haven't had any issues. But what I do is usually pretty simple. What makes the others vastly better?
- MattDamonSpace 6mo agoI’ve rarely had macOS TTS produce a sentence I didn’t have to edit Whisper models I barely bother checking anymore
- qq66 6mo agoI'm speaking for >1 minute and including bulleted lists, etc. WisprFlow gets all of the bulleted lists formatted correctly, and I'm not saying things like "Bullet 1" -- just speaking as I'd speak to a person.
- atlgator 6mo agoThis thread is a support group for people who have each independently built the same macOS speech-to-text app.
- brcmthrowaway 6mo agoOh to be 20-something and do a bunch of free work for your portfolio again
- obrajesse 6mo agoI'll have you know that I'm Matt's top contributor to Ghost Pepper and I'm nearly fifty But I did it because I wanted it to work exactly the way I wanted it. Also, for kicks, I (codex) ported it to Linux. But because my Linux laptop isn't as fast, I've had to use a few tricks to make it fast. https://github.com/obra/pepper-x https://github.com/obra/pepper-x
- dotancohen 6mo agoI'll look at this, thank you. I haven't yet gotten around to vibe coding my own itch yet so maybe your scratching will do.
- deleted 6mo ago[deleted]
- tpowell 6mo agoI cobbled my own together one night before I came across the thoughtfully-built KeyVox and got to talking shop with its creator. Our cups runneth over. https://github.com/macmixing/keyvox/ https://github.com/macmixing/keyvox/
- lxe 6mo agohahaha I’m glad I’m just a procedurally generated NPC I built one for cross platform — using parakeet mlx or faster whisper. :)
- karimf 6mo ago
- Ecko123 6mo ago[dead]
- tito 6mo agoThis is great. I'm typing this message now using Ghost Pepper. What benefits have you seen from the OCR screen sharing step?
- janalsncm 6mo agoI think the jab at the bottom of the readme is referring to whispr flow? https://wisprflow.ai/new-funding https://wisprflow.ai/new-funding
- dakila5 6mo agoMacWhisper is also a good one
- arkensaw 6mo agoThis is great, and I'm not knocking it, but every time I see these apps it reminds me of my phone. My 2021 Google Pixel 6, when offline, can transcribe speech to text, and also corrects things contextually. it can make a mistake, and as I continue to speak, it will go back and correct something earlier in the sentence. What tech does Google have shoved in there that predates Whisper and Qwen by five years? And why do we now need a 1Gb of transformers to do it on a more powerful platform?
- com2kid 6mo agoMicrosoft OneNote had this back in 2007 or so, granted the speech to text model wasn't nearly as advanced as they are now. I was actually on the OneNote team when they were transitioning to an online only transcription model because there was no one left to maintain the on device legacy system. It wasn't any sort of planned technical direction, just a lack of anyone wanting to maintain the old system.
- deleted 6mo ago[deleted]
- rudhdb773b 6mo agoI remember trying out some voice-to-text around 2002 that I believe was included with Windows XP.. or maybe Office? You had to go through some training exercises to tune it to your voice, but then it worked fairly well for transcription or even interacting with applications.
- silon42 6mo agoOS/2 had it built in in 1996.
- adamsmark 6mo agoThe accuracy is much lower though. I've switched away from Gboard to Futo on Android and exclusively use MacWhisper on MacOS instead of the default Apple transcription model.
- deleted 6mo ago[deleted]
- raybb 6mo agoWould also like to know how it compares to https://github.com/openwhispr/openwhispr https://github.com/openwhispr/openwhispr I like that openwhisper lets me do on device and set a remote provider.
- atlasagentsuite 6mo ago[flagged]
- thatxliner 6mo agowhy isn't the cleanup done on the transcription (as opposed to screen record)
- pmarreck 6mo agoHow does this compare with Superwhisper, which is otherwise excellent but not cheap?
- cupcake-unicorn 6mo agohttps://handy.computer/ https://handy.computer/ already exists?
- smcleod 6mo agoYeah props to Handy, really nice tool.
- forbiddenvoid 6mo agoMore than one solution can exist for the same problem.
- semiquaver 6mo agoI have a few qualms with this app: 1. For a Linux user, you can already build such a system yourself quite trivially by getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem. From Windows or Mac, this FTP account could be accessed through built-in software. 2. It doesn't actually replace a USB drive. Most people I know e-mail files to themselves or host them somewhere online to be able to perform presentations, but they still carry a USB drive in case there are connectivity problems. This does not solve the connectivity issue. 3. It does not seem very "viral" or income-generating. I know this is premature at this point, but without charging users for the service, is it reasonable to expect to make money off of this?
- Graziano_M 6mo agoI got that reference!
- MegagramEnjoyer 6mo agowhy does it need to generate money?
- morelikeborelax 6mo agoThis is the reply that was posted when Dropbox was first shown off on HN. It's a joke :)
- fiatpandas 6mo agoThe clean up prompt needs adjusting. If your transcription is first person and in the voice of talking to an AI assistant, it really wants to “answer” you, completing ignoring its instructions. I fiddled with the prompt but couldn’t figure out how to make it not want to act like an AI assistant.
- vaulpann 6mo agovery cool - huge open source drop!
- zephyrwhimsy 6mo ago[flagged]
- kushalpandya 6mo agoSpeecg-to-text is basically AI version of Todo app that we used to build every week when new frontend framework would release.
- deleted 6mo ago[deleted]
- maxmorrish 6mo agolove seeing more local-first tools like this. feels like theres been a real shift since the codebeautify breach last year, people are actually thinking about where there data goes now. nice work on keeping it all on device
- aesopturtle 6mo ago[dead]
- boudra 6mo agoInteresting, I'm surprised you went with Whisper, I found Parakeet (v2) to be a lot more accurate and faster, but maybe it's just my accent. I implemented fully local hands free coding with Parakeet and Kokoro: https://github.com/getpaseo/paseo https://github.com/getpaseo/paseo
- ghm2199 6mo agoI've been using handy since a month and its awesome. I mainly use it with coding agents or when I don't want to type into text boxes. How is this different? Part of the reason handy is awesome is because it uses some of the same rust infra for integrating with the model, so that actually makes it possible to use the code as a library in android or iOS. I have an android app that runs on a local model on the phone too using this.
- imazio 6mo agois this the support group for people building speech-to-text apps? I built https://yakki.ai https://yakki.ai No regrets so far! XP
- pdyc 6mo agointeresting, i wanted something like this but i am on linux so i modified whisper example to run on cli. Its quite basic, uses ctrl+alt+s to start/stop, when you stop it copies text to clipboard that's it. Now its my daily driver https://github.com/newbeelearn/whisper.cpp https://github.com/newbeelearn/whisper.cpp
- snickell 6mo agoCan somebody help me understand how they use these, I feel like I'm missing something or I'm bad at something? I only spent 10 minutes with Handy, and a similar amount of time with SuperWhisper, so pretty ignorant. I tried it both with composing this comment, and in a programming session with Codex. I was slightly frustrated to not be hands free, instead of typing, my hands were having to press and release a talk button (option-space in handy, right-command in superwhisper), but then I couldn't submit, so I still had to click enter with Codex. Additionally, for composing this message, I'm using the keyboard a ton because there's no way I can find to correct text I've typed. Do other people get really reliable and don't need backspace anymore? Or.... what text do you not care enough to edit? Notes maybe? My point of comparison is using Dragon like 15 years ago. TBH, while the recognition is better (much better) on handy/superwhisper, everything else felt MUCH worse. With dragon, you are (were?) totally hands free, you see text as you say it, and you could edit text really easily vocally when it made a mistake (which it did a fair bit, admittedly). And you could press enter and pretty functionally navigate w/o a keyboard too. Its weird to see all these apps, and they all have the same limitations?
- sorkhabi 6mo agoWell done
- jannniii 6mo agoOh dear, why does it not use apfel for cleanup? No model download necessary…
- Sukhbat 6mo ago[flagged]
- eddie-wang 6mo ago[dead]
- miki123211 6mo agoWhat do you actually use for STT, particularly if you prize performance over privacy and are comfortable using your own API keys? I was on WhisperFlow for a while until the trial ran out, and I'm really tempted to subscribe. I don't think I can go back to a local solution after that, the performance difference is insane.
- k9294 6mo agoTry ottex.ai - it has an OpenRouter like gateway with most STT models on the market (Gemini, OpenAI, Groq, Deepgram, Mistral, AssemblyAI, Soniox), so you can try them all and choose what works best for you. My favorites are Gemini 3 Flash and Mistral Voxtral Transcribe 2. Gemini when I need special formatting and clean-up, and Voxtral when I need fast input (mostly when working with AI).
- ianmurrays 6mo agoI had Claude make this hammerspoon config + daemon that does pretty much the same, in case anyone is interested. https://github.com/ianmurrays/hammerspoon/blob/main/stt.lua https://github.com/ianmurrays/hammerspoon/blob/main/stt.lua
- jiusanzhou 6mo ago[dead]
- jwr 6mo agoI currently use MacWhisper and it is quite good, but it's great to see an alternative, especially as I've been looking to use more recent models! I hope there will be a way to plug in other models: I currently work mostly with Whisper Large. Parakeet is slightly worse for non-English languages. But there are better recent developments.
- bambushu 6mo ago[flagged]
- joshuahart 6mo ago[dead]
- zephyrwhimsy 6mo ago[flagged]
- marktolson 6mo agoI got it to transcribe this: "Create tests and ensure all tests pass" and instead of transcribing exactly what I said it outputs nonsense around "I am a large language model and I cannot create and execute tests". Other than that issue I like it.
- therealdeal2020 6mo agobtw I know at least a dozen doctors that still pay for software like this. I think doctors are THE profession that likes to use speech-to-text all day every day
- edoardobambini- 6mo ago[dead]
- ezVoodoo 6mo agoHi, nice project! Quick question, when I speak Chinese language, why it output English as translated output? I was using the multilingual (small) model. Do I need to use the Parakeet model to have Chinese output? Thx.
- nidnogg 6mo agoI really like the project and am eager to try and fit this into some of my workflows. However, this bothered me a bit: "All models run locally, no private data leaves your computer. And it's spicy to offer something for free that other apps have raised $80M to build." I’d straight up drop the comparison to big AI labs. This isn’t rebellious or subversive, it’s downstream of a ton of already-funded work. Calling it “spicy” is a bit misframed.
- nidnogg 6mo agoThis got me thinking that the smaller these local first LLMs get - the more they're gonna looking the next bread and butter of app dev. Reminds me how Electron gained a lot of traction for making it easy to package prettier apps. At the measly cost of gigabytes of RAM, give or take.
- rpdaiml 6mo ago[dead]
- acjacobson 6mo agoNice app! Feedback since you asked: The most obvious must-have feature IMO is to paste automatically. Don't require me to hit a shortcut (or at least make it configurable) The next most critical thing I think is speed and in my tests it's just a little bit slower than other solutions. That matters a lot when it comes to these tools. The third thing, more of a nice to have is controlling formatting. By this I mean - say a few sentences, then "new line" and the model interprets "new line" as formatting, not as literal text.
- philbitt 6mo ago[dead]
- zhichuanxun 6mo ago[dead]
- ssz7820 6mo ago[dead]
- cisco801 6mo ago[dead]
- wujiahua 6mo ago[dead]
- mft_ 6mo agoDoes it show your spoken words on the screen live (i.e. streaming) or does it wait until you’ve finished speaking? I find it very helpful to see my words live - for some reason it helps my simple brain structure what I’m saying, and I’m much more fluent as a result. I went on a mission a few weeks ago and tried every freely available MacOS STT app I could find (and there are lots of them) - but none I tried had this feature and was otherwise satisfactory. (I vibe-coded a PoC which could do this, so it’s definitely possible.)
- deleted 6mo ago[deleted]
- deleted 6mo ago[deleted]
- leeeeep101 6mo agoi also did this dictatorflow lee101/voicetype i open sourced it also. nice. might be good reference :)
- kingofbits 6mo agoNicely done! Ive been abusing chatgpt's overlay window for this, until now
- aimemobe 6mo ago[flagged]