7 ms·
Handy – Free open-source speech-to-text app written in Rust
- Leftium 1y agoRead the creator's description in the original Show HN: https://hw.leftium.com/#/item/44302416 https://hw.leftium.com/#/item/44302416
- roscas 1y agoIt downloads the model at first execution and also checks versions in github. That is ok for what is brings. Nice program. Very "handy".
- sipjca 1y agoIf you prefer a more stripped down version: the original releases (0.1.0 and 0.1.1) shipped with Whisper tiny included and no auto-update feature
- majorchord 1y agoTypeScript 53.9% Rust 44.9% FYI
- typpilol 1y agoLmao. At least it's typescript and not JavaScript!
- shakabrah 1y agoWho’s gonna tell him?
- loloquwowndueo 1y agoDon’t you dare!
- nicce 1y agoYeah. Rust compiles to machine code.
- typpilol 1y agoI thought it was a clever joke
- yoavm 1y agoThe README is very clear about it: Frontend: React + TypeScript with Tailwind CSS for the settings UI Backend: Rust for system integration, audio processing, and ML inference
- areeba_iqbal 1y agoThat's great, nice to see more and more projects of Machine learning being written in rust
- dcre 1y agoIt’s not really a machine learning project. It’s an application that calls existing models.
- amelius 1y agoRepo says: CPU-optimized speech recognition with Parakeet models
- dcre 1y agoI understand that it uses ML models. My point is that it is an end-user application making use of such models. It is recording audio, passing it to the model, and pasting in the resulting text to the focused input. The fact that the middle step happens to involve an ML model is not really intrinsic to anything the app does. If there was a good speech to text program that did not use ML, the app could use that instead and not really be any different.
- sipjca 1y agoTo be fair on the other side there is a fair lack of specific ML inference libraries in Rust, and this project is pushing some of that forward with Parakeet at the very least. The Rust library `transcribe-rs` came from it and hopefully will support more models in the future. While certainly it's not an ML project in the sense of I am not training models, the inference stack is just as important. The fact is the application does do inference using ONNX and Whisper.cpp.
- ajsnigrutin 1y agoMore than half the code is typescript.
- ranger_danger 1y agoAnyone know of the opposite? A really easy-to-use text-to-speech program that is cross-platform?
- jszymborski 1y agoI've used Speech Note, which works well for STT and TTS.
- geor9e 1y agoI've tried a lot of them, and the best I found so far is Edge browsers built in microsoft (natural) voices, which I call via javascript or the browsers read aloud function.
- yoavm 1y agoCheckout https://github.com/rany2/edge-tts https://github.com/rany2/edge-tts , which exposes it as a Python library and a CLI tool.
- ompogUe 1y agoBeen having fun with this one https://addons.mozilla.org/en-CA/firefox/addon/read-aloud/ https://addons.mozilla.org/en-CA/firefox/addon/read-aloud/ Read Aloud allows you to select from a variety of text-to-speech voices, including those provided natively by the browser, as well as by text-to-speech cloud service providers such as Google Wavenet, Amazon Polly, IBM Watson, and Microsoft. Some of the cloud-based voices may require additional in-app purchase to enable. ... the shortcut keys ALT-P, ALT-O, ALT-Comma, and ALT-Period can be used to Play/Pause, Stop, Rewind, and Forward, respectively.
- derekja 1y agoI’ve been enjoying Kokoro Amazing what it can do with only 82M parameters https://www.kokorotts.io/ https://www.kokorotts.io/
- sipjca 1y agoCurious your use case, I now have quite a lot of experience with releasing desktop apps, and I have done some accessibility work as well, and may be curious to put together a TTS toolkit as well into a desktop app (or even Handy)
- hu3 1y agoVery cool. Uses whiper small uder the hood. https://github.com/openai/whisper https://github.com/openai/whisper
- geor9e 1y agonvidia parakeet v3 was the default out of the box and it works surprisingly well it offers all the different sizes of openai models too
- jonahx 1y agoHow good will this local model be compared to, say, your iphone builtin STT?
- dcre 1y agoIt’s way better. iPhone’s is awful. On macOS, interestingly, the built in dictation seems a bit better than on iOS, but still not as good as Whisper and Parakeet. Worth noting I have never used Whisper Small, only large and turbo. Another comment says Parakeet is the default now, though, despite what the site says.
- sipjca 1y agoAuthor here! The default recommendation is Parakeet (mainly because it runs fast on a lot more hardware), but definitely think people should experiment with different models and see what is best for them. Personally I found Whisper Medium to be far better than Turbo and Large for my speech, and Parakeet is about on par with Medium, but each have their own quirks. I'll update the site soon!
- dcre 1y agoThat's really interesting about medium being better than large. I never bothered trying the smaller models since the big ones were fast enough.
- sipjca 1y agoBenchmarks definitely say otherwise, but my anecdotal experience says medium is the best for this application with my voice and microphone
- rgbrgb 1y agothis is a great landing page. I downloaded. great onboarding too, using it now. Very handy, thanks!
- b_e_n_t_o_n 1y agoWhy does the title specify the language used when it's not even mentioned on the home page?
- nicce 1y agoMarketing. Honestly, might not be good here since it is not library and not completely written in Rust.
- ajsnigrutin 1y agoMarketing for what exactly? I mean... why would I want this app instead of some other app? Just because it's written in the language of the week? If it said "20% faster than xyz" it would be a much better marketing than saying it's written in rust, even though more than half the code is typescript.
- quicklime 1y agoThe title also mentions that it’s open source, so it could be marketing for potential contributors.
- sipjca 1y agoIt's primarily this. I'm a novice Rust developer and really would like to improve the code quality across the board, and some of this comes to attracting the right kind of developers to help. Maybe "Rust" in the title helps, maybe it doesn't. Clearly HN doesn't like it and that's okay. I stated my need for help on the about page as well > This is my first Rust project, and it shows. There are bugs, rough edges, and architectural decisions that could be better. I’m documenting the known issues openly because I want everyone to understand what they are getting into, and encourage improvement in the project.
- nicce 1y ago> Maybe "Rust" in the title helps, maybe it doesn't. Clearly HN doesn't like it and that's okay HN definitely likes it, when it is used in the correct context. Using Rust in the title is a soft promise for better reliability and quality for the software than on average. But it starts to get controversial when Rust is not purely the controlling part of the software anymore. So people start to complain because it can be misleading marketing which is based on the promise that Rust can offer.
- perfmode 1y agohow’s it differ from macos dictation?
- dcre 1y agoI find state of the art speech to text models like Whisper and Nvidia Parakeet are a lot better than macOS dictation. I use them through MacWhisper, but this is basically the same.
- geor9e 1y agoJust compare them side by side. On one side, the dictation tech baked into you OS, the other side transformer models like Whisper Large or Parakeet. Mumble from across the room from the mic. The difference is staggering.
- m13rar 1y agoAwesome . I was looking to build this on my own. Will look at the code and consider contributing cheers.
- sipjca 1y agoHey author of Handy here! Would absolutely love any help, please let me know if there's any way I can make contributing easier!
- deleted 1y ago[deleted]
- efskap 1y agoCool, you just might've saved me some carpal tunnel in the long run xD. I guess there's no way for the AppImage to use GPU compute, right? Not that it matters much because parakeet is fast enough on CPU anyway.
- Leftium 1y agoI think the Whisper models will all use GPU. Only the Parakeet model is limited to CPU. (I'm unfamiliar with AppImage. Was the model included in the app image, or was there a download after selecting the model?)
- mamonoleechi 1y agonot sure this might help, but when you launch the .appimage in a terminal, it shows you the command to extract the files it contains (to speed the loading) ; this might help you find the files you're searching for, maybe :)
- sipjca 1y agoWhisper uses Vulkan and Metal acceleration with whisper.cpp Parakeet is currently CPU only
- daakus 1y agoShameless plug: A brutally minimalist Linux only, whisper.cpp only app: https://github.com/daaku/whispy https://github.com/daaku/whispy I wanted speech-to-text in arbitrary applications on my Linux laptop, and I realized that loading the model was one of the slowest parts. So a daemon process, which triggers recording on/off using SIGUSR2, records using `pw-record` and passes the data to a loaded whisper model, which finally types the text using `ydotool` turned out to be a relatively simple application to build. ~200 lines in Go, or ~150 in Rust (check history for Rust version).
- efskap 1y agoI'm very curious about the rewrite. Was Rust slowing you down too much?
- daakus 1y agoJust for fun. I like both languages. I thought Rust would be better fit on account of interop with whisper.cpp, but turns out the use of cgo was straight forward in this case. I like that the Go version has minimal 3rd party dependencies compared to the Rust version.
- DoctorOW 1y agoWhy Linux only? Isn't Go and Whisper.cpp cross platform?
- daakus 1y agoIt relies on `pw-record` for recording audio and `ydotool` for triggering keyboard input. These are Linux specific. I don't know about Windows, but on my Mac I have a not-yet-public Swift + whisper + CoreAudio + Accessibility based solution that provides similar functionality.
- atoav 1y agoThat was my guess. Crossplatform Audio input isn't exactly as trivial as using pipewire.
- vladstudio 1y ago+1, happy user and a humble contributor.
- sipjca 1y agoYou're awesome Vlad!
- primaprashant 1y agobuilt something similar for terminal lovers. It's a CLI tool built in Python called hns [1] and uses faster-whisper for completely local speech-to-text. It automatically copies the transcription to the clipboard as well as writes to stdout so you seamlessly paste the transcription in any other application or pipe/redirect it to other programs/files. [1]: https://github.com/primaprashant/hns https://github.com/primaprashant/hns
- oulipo2 1y agoNice! There's also the VoiceInk open-source project https://github.com/Beingpax/VoiceInk/ https://github.com/Beingpax/VoiceInk/
- atmanactive 1y agoMacOS only.
- precompute 1y agoThis is local, but I've found that external inference is fast enough, as long as you're okay with the possible lack of privacy. My PC isn't beefy enough to really run whisper locally without impacting my workflow, so I use Groq via a shell script. It records until I tell it to stop, then it either copies it to the clipboard or writes it into the last position the cursor was in.
- sipjca 1y agoWhat computer are you using? You really should give Parakeet a try, I find it runs in a few hundred milliseconds even on a Skylake i5 from 10 years ago.
- deleted 1y ago[deleted]
- kwar13 1y agoNicely done! Seeing that it uses a port of Whisper, here's my shameless plug for a gnome extension I made using Whisper: https://extensions.gnome.org/extension/8238/gnome-speech2text/ https://extensions.gnome.org/extension/8238/gnome-speech2tex...
- amelius 1y agoHow handy is this for coding? ;)
- skeptrune 1y agoAmazing! I have been desperately wanting this. Livecaptions doesn't seem to be maintained super well.
- mzimbres 1y ago[flagged]
- amelius 1y agoHow can I call this library from C++?
- thelittleone 1y agoIs it able to isolate the speaker from background noises / voices?
- sipjca 1y agoRight now there is fairly minimal processing done to the audio. There is a VAD filter to reduce the non-speech areas. But there is no noise-reduction as such. The audio pipeline could support it though, so if you know any good real time noise reduction filters let me know. Would love to improve the SNR into the models
- mathverse 1y agoEven being in Tauri this application just by doing these things takes around 120MB on my M3 Max. It's truly astonishing how modern desktop apps are essentially doing nothing and yet consume so much resources. - it sets icon on the menubar - it display a window where I can choose which model to use That's it. 120MB FOR doing nothing.
- 3oil3 1y agoI feel the same astonishment! Our computers surely are today faster and stronger and smaller than yesterdays', but did this really translate in something tangible for a user? I feel that besides boot-up, thanks to SSDs rather than gigaHertz, it's not any faster. It's like, all this extra power is used to the maximum, for good and bad reasons, but not focused on making 'it' faster. I get a bit puzzled to why my mac could freeze half a second when I 'cmd+a' in some 1000+ files-full folder. Why doesn't Excel appear instantly, and why is it 2.29GB now when Excel 98 for Mac was.. 154.31MB? Why is a LAN transfer between two computers still as slow as 1999, 10ishMB/s, when both can simultaneously download at > 100MB/s? I'm not starting with GB-memory-hoarding tabs, when you think about it, it's managed well as a whole, holding 700+ tabs without complaining. And what about logs? This is a new branch of philosophy, open Console and witness the era of hyperreal siloxal, where computational potential expands asymptotically while user experience flatlines into philosophical absurdity?
- dreamcompiler 1y agoIt me takes longer to install a large Mac program from the .dmg than it takes to download it in the first place. My internet connection is fairly slow and my disk is an SSD. The only hypothesis that makes sense to me is that MacOS is still riddled with O[n] or even O[n^2] algorithms that have never been improved and this incompetence has been made less visible by ever-faster hardware. A piece of evidence supporting this hypothesis: rsync (a program written by people who know their craft) on MacOS does essentially the same job as Time Machine, but the former is orders of magnitude faster than the latter.
- sipjca 1y agoA lot of the bloat comes from dependencies like ONNX or whisper.cpp to accelerate running the model itself While the UI is doing “nothing” most of the bloat is not from the UI
- NaomiLehman 1y agojust a heads up. There are many more accurate and faster models than Whisper nowadays. https://huggingface.co/spaces/hf-audio/open_asr_leaderboard https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
- sipjca 1y agoIt also uses one of the fastest and most accurate on the ASR leaderboard, Parakeet.
- queenss90 1y ago[flagged]
- sidhusmart 1y agoI love this tool. Been using this for the past 2 weeks and it works great. Struggles a bit in noisy settings but it's weird talking to your computer in a coffee shop anyways :P