9 ms·
Transcribe.cpp
- SamPentz 2mo agoIs there a way to add speaker identification easily?
- semiquaver 2mo agoThat would be Diarize.cpp, not Transcribe.cpp.
- sipjca 2mo agoHey! It’s actually in progress right now, probably will come this week :)
- arikrahman 3mo agoExcellent work, paired with the 500kb TTS model headlining today I can see the full stack coming together.
- zuzululu 3mo agosaw the demo its impressive but the audio was robotic
- aarvin_roshin 3mo agoSpot on: > I think as we look forward to the future, more inference will start happening locally for one reason or the other. This brings the distribution story front and center. In order to have more applications running inference locally, we need to make running inference easier. This makes these projects so much more trustworthy and easier to approach: > Were any of the words here written using AI? Nope. They came from my mouth or my fingers.
- boplicity 3mo ago>This makes these projects so much more trustworthy and easier to approach: >> Were any of the words here written using AI? Nope. They came from my mouth or my fingers. I have to push back on this a bit, as I believe (quite strongly) that we're shaped by the tools we use; text-to-speech LLMs are still LLMs, and generally their mistakes are shaped by the expectations inherent in their training. This, in turn, shapes the words that appear on the screen. For those who regularly use them, you then learn which word sequences are likely to be accurately transcribed, and this definitively becomes part of your thinking process. Over time, the LLM becomes tangled into your thinking; the use of AI, even in this way, very much can and often does shape the resulting words.
- nullsanity 2mo ago[dead]
- eventualcomp 2mo agoIsn't this like saying "my words are not really my own when I speak to my family, because I know my father is a non-native English speaker and hard of hearing so I try to use words which are well enunciated and are few in syllable count"?
- iezepov 2mo agoYou can take it one step further! As Tyutchev wrote, "A thought once uttered is a lie." [1] Speech is a projection of a thought, and a lossy one. So no matter who is the listener, the speaking/writing does affect the thinking. Though comment on LLM transcribing is spot on. 1. https://www.poetryloverspage.com/poets/tyutchev/silentium/literary-analysis https://www.poetryloverspage.com/poets/tyutchev/silentium/li...
- deleted 2mo ago[deleted]
- ChadNauseam 2mo agoParakeet, probably the most popular model for handy-style transcription, doesn't include a "text-to-speech LLMs" or any other form of LLM
- yjftsjthsd-h 3mo agoSo it's mostly intended to be a better replacement for whisper? Mostly? With better support for more models and maybe acceleration backends?
- sipjca 2mo agoMore or less yes, for whisper.cpp, just trying to make local transcription more accessible to anyone building an app, etc
- sbinnee 3mo agoI saw that metal is almost x10 faster than vulkan? Why so much gap?
- sipjca 2mo agoIt very much depends on the hardware! An M4 max is being compared against a Ryzen 4750U with an integrated GPU! The M4 max has probably 10x the compute and memory bandwidth hahaha
- zuzululu 3mo agowould love to see a demo handy is fantastic although its still behind the frontier models
- therealpygon 3mo agoPretty sure I saw Handy using it; if you have the latest version, you’re probably already demoing it.
- loufe 3mo agoauthor of the blogpost is the maintainer of Handy, so almost guaranteed!
- zuzululu 2mo agoI installed it but I don't think I see the streaming transcriptions. I do think the transcription is a bit faster. I am using the latest version.
- qntmfred 2mo agoYou have to change the model to one that supports streaming. The latest parakeet does. I've been using it the last week or so. It's good stuff :)
- therealpygon 2mo agoI tried nemotron but the zero punctuation was a no go. I also used unified but didn’t see streaming. That said, I had to disable direct input because my shortcut for handy includes a ctrl-, so direct causes menu and shortcut activations at lightning speed, so I am only getting output when recording stops. Still, you can pry Handy from my cold dead hands-y either way. So useful.
- sipjca 2mo agoYep the latest version has support! Virtually all of the SOTA open models are supported by Handy including the streaming ones like Nemotron Streaming Parakeet Unified Voxtral Mini Realtime If something you want is not supported, open an issue on transcribe.cpp!
- ghm2199 2mo agoCongrats on shipping this. I love handy on my Mac, my phone for STT in situations where it’s not possible/poor performance of the native Model for STT(e.g apple’s thing is not upto scruff, like mistranslating words corresponding to a domain). Noob question: How do you think about funding from a foundation(i have no clue if you need it or not, I do hope you have a way to get paid one way or another because handy is amazing) for maintenance of this? if you did or were going to get paid by asking for maintaining such a project what might be the kind of organizations you would look for to get supported and how would you do it?
- sipjca 2mo agoThanks! What an excellent question, I’m not sure I have a good answer. I kind of became an open source maintainer by accident as Handy became popular Certainly I am very lucky that quite a few people donate to Handy, and also some people and organizations who sponsor the work I do To be honest I just love contributing to open source and wish to continue to do so. So anyone who supports this is good to me. Organizations which believe in OSS and push it forward are typically most aligned with me Of course you can always email me (contact@handy.computer) and we can discuss in more detail
- sneak 2mo agoOS-native dictation on iOS requires uploading your address book to Apple on every request, even if you don’t use iCloud. I unfortunately have to leave it disabled for this reason.
- bengotow 2mo agoThis is an incredible contribution to the community and it's just... one guy? I kept reading expecting a Series A funding announcement at the bottom. It's a nice reminder: You can use AI to slop cannon at maximum speed, or you can use it to scale your ambitions and build something more rigorous and lasting than ever before. I'd build Transcribe.cpp into the apps I maintain, but I feel like this functionality should (generally) be integrated into the OS or "everywhere" via an app like Handy.
- sipjca 2mo agoHey, yep author and maintainer here! Certainly sponsors help and the wonderful community who donates to Handy as well! Mozilla AI was very helpful in getting this work off the ground. It was a pipe dream for me to build for Handy and they helped to sponsor me so I could make time to take this project seriously and get a v0.1.0 release out the door I agree this should be everywhere and I hope to distribute libtranscribe some day properly so it is more a system library! It will take time to stabilize but I think we can get there
- espetro 2mo agoHey, for all of us builders out there, this is some quality work you're putting out there. How do you handle to talk about it, distribute it, just basically making it known to people, and still have time to focus on delivering actually useful, extensible quality work?
- shade 2mo agoNice - I'm definitely going to take a look at this. I've built my own cross-platform (Mac/Win/Linux) live captioning app on top of Nemotron, and it works well but dealing with ONNX is kind of annoying. With this having Rust support (I built it on Rust/Tauri) it should be a pretty solid candidate; I'll have to see if I can find a Silero VAD implementation that doesn't depend on ONNX, or maybe I'll see if the clankers can migrate it for me.
- nohup2 2mo agoHave you published your app? I would love to take a look and test it
- shade 2mo agoI have not formally published it, but it's open source: https://github.com/edmistond/larmindon https://github.com/edmistond/larmindon and https://github.com/edmistond/larmindon-core https://github.com/edmistond/larmindon-core - you'll want to clone them into the same root directory. Right now it only supports languages supported by parakeet-rs and Nemotron (so... English only as far as I'm aware) and you'll need the ONNX version of Nemotron: https://huggingface.co/altunenes/parakeet-rs/tree/main/nemotron-speech-streaming-en-0.6b https://huggingface.co/altunenes/parakeet-rs/tree/main/nemot... The first run experience isn't great, you'll need to download all the files from the model, start the app, and then go to settings and configure the model directory. It runs well on Mac and Windows; I haven't tested it on Linux in a couple of months since my Linux install is out of commission currently.
- aomix 2mo agoWhat good timing to spot this. I've been reading more and more people talk about bringing TTS into their prompting toolkit and wanted to give that a try. The idea of rambling brain dump into a doc -> edit pass -> send to the robot loop sounds appealing.
- ukuina 2mo agoWhat's the easiest way to add speaker separation to this?
- sipjca 2mo agoHey! It’s actually in progress right now, probably will come this week :)
- simonw 2mo agoAwesome! I found the in-progress diarization PR here: https://github.com/handy-computer/transcribe.cpp/pull/85 https://github.com/handy-computer/transcribe.cpp/pull/85 Looks like it's using IBM's Granite-Speech-4.1-2B-Plus https://huggingface.co/ibm-granite/granite-speech-4.1-2b-plus https://huggingface.co/ibm-granite/granite-speech-4.1-2b-plu... and/or MOSS-Transcribe-Diarize https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize
- sipjca 2mo agoYep, but I am in the process of also porting NVIDIAs Sortformer for multi speaker diarization as well :) I’m not sure how many specific models will be supported as the library is more focused on transcription specifically. But the models which support diarization natively must be supported I think. And parakeet multitalker was the primary driving force for this change
- oezi 2mo agoHow close do you aim for when it comes to drop-in vs whisper.cpp? Are timestamps per word and character something aimed for? How about multi-lingual transcription or hallucination suppression? The github page doesn't seem to go into depth on these orthogonal topics. May have missed it.
- sipjca 2mo agoEventually I would like to be more fully drop in compatible, right now some feature support is a bit sparse. And whisper has so much work done to it over the years so it’s hard to support every possible thing. Right now it’s a more bog standard implementation than anything special. Right now stabilizing the core header is probably among the primary goal, but if people want to contribute model specific things im happy to review test and pull in. Whisper is a good case for this as there is a header extension already so it’s easier
- copypirate 2mo agoExcellent work CJ
- sipjca 2mo agoThanks :)
- wolvoleo 2mo agoLooks interesting, I'll give it a try. Though I'm really happy with faster-whisper on a GPU.
- diimdeep 2mo agoCongrats on delivering good value to the people. I have used transcribe.cpp a few weeks ago to do near realtime offline stt on a 10 year old phone, writing simple adhoc app for my use case, it's crazy what is happening right now.
- sipjca 2mo agoHa amazing, love to hear it
- simonw 2mo ago> Maintainer supported bindings in 4 Languages Nice. Here's the Python one: https://github.com/handy-computer/transcribe.cpp/tree/main/bindings/python https://github.com/handy-computer/transcribe.cpp/tree/main/b... - looks like it's not yet available as a binary wheel on PyPI with the dependency included (the library on PyPI right now uses ctypes to call a separately installed library) but that's planned for a future release.
- sipjca 2mo agoYes, I’ve put a PR up on pypi for extra storage for CUDA but it has not been accepted yet afaik If there’s any issues or improvements on the bindings I would love help to make the DX the best it can be
- lxe 2mo agoWhat's the best local TTS model right now? I'm running parakeet on a mac which transcribes all my uh's and aahs. I'm running whisper on linux/cuda and I by far prefer that one over parakeet.
- jv22222 2mo ago> parakeet on a mac which transcribes all my uh's and aahs You should be able to fix this by playing with the mic speech floor. It happens when to much ambient stuff slurps in. It's actually gaslighting you, you don't say that many ums and ahs ;)
- lxe 2mo agoThat is definitely not the case. It transcribes everything very accurately, especially all the "ahs" and "uhms".
- sipjca 2mo agoParakeet unified for me no longer does this and it’s also a streaming transcription model! But the answer largely depends on you, the languages you speak, and personal preference. Whisper is still excellent and supported in transcribe.cpp Cohere Transcribe is also excellent, but many of the new models are as well
- TomGarden 2mo agoI run the same, if you want try a simple filter post transcription to remove them, and while you're at it add some simple word replacements like 'cloud MD' to 'CLAUDE.MD'
- aghilmort 2mo agoaudio.cpp @ github repo has good collection of TTS models
- zaptheimpaler 2mo agoAmazing, i've been looking for something like this and ended up doing transcription + diarization on a local server for now. Are you looking for contributions? Have you tried this one for diarization - https://huggingface.co/pyannote/speaker-diarization-community-1 https://huggingface.co/pyannote/speaker-diarization-communit... - it performed much better than Sortformer for me.
- sipjca 2mo agoContributions are always welcome! There’s a WIP diarization PR rn, and after it’s merged would love to have support if it fits well into the interface. And if not would love to figure out a good interface for it
- rcarmo 2mo agoYeah, diarization is the real feature these days. STT needs uniformization, but quality of diarization is what is setting personal solutions apart in this field.
- sipjca 2mo agoFor sure, it was not initially a target because I didn’t need it for Handy but I do understand the importance in the broader context
- ilaksh 2mo agoI don't suppose this works in the browser?
- sipjca 2mo agoOut of the box no probably not, but if people are interested there’s probably ways forward
- _medihack_ 2mo agoAbsolutely interested
- abdullahkhalids 2mo agoFor anyone looking to build on top of this. I have tried a few different STT systems, and they accurately capture what I am saying. Unfortunately, they don't support the reasonable workflow I want to open an office document, for example, and start talking. And I want the software to continuously type what I am saying at the cursor with minimal latency. The continuous part is crucial. Many software will paste whatever I said after I have stopped recording, but that is not useful.
- mijoharas 2mo agoAgreed. It's something I've found annoying about a few systems. It looks like the rust bindings have streaming examples so hopefully there is a nice solution here.
- sipjca 2mo agoYou can fairly easily modify [Handy](https://handy.computer https://handy.computer) to do this if you want I’m planning on having it as a first class feature of the app too just too many other issues to work on first
- mft_ 2mo agoI’m really glad to hear this! A while ago, I auditioned about 10 different STT apps on my Mac, with this realtime/streaming transcription as a goal. I failed to find that feature in an app I was happy with, but settled on Handy as the best option otherwise. So if Handy adds this, it will be perfect!
- rolisz 2mo agoCan you give some pointers around this? I'd gladly help with a PR for this, but if you have anything docs/ideas around this it would be helpful.
- sipjca 2mo agoI’m on a train right now but off the top of my head the audio pipeline may have to be modified slightly to emit partial text segments as they come in from the transcription engine. And then calling the appropriate paste method the user has in their settings. It may be easier than expected in some way since we already emit events for the live overlay, so it could be as small as a function call, but I don’t know the code path well enough from memory and what complexities it has. Probably with the Tauri context and a bit of other mess we have as this bit of code has gone through a lot of pain
- kzyxx11 2mo agoExcellent work
- dostick 2mo agoDoes this support filtering of “umm”,”err”, “ugh”, or that is nit yet possible with open source models?
- sipjca 2mo agoNot in the library itself, it’s pure inference. Some models have this trained out of them anyhow. Otherwise this is a post processing task which is not really inference
- primaprashant 2mo agoHandy is an amazing cross-platform app for dictation from the author. There are other awesome open-source dictation tools as well like native macOS ones. You do not need SaaS subscription in this day and age for transcription. I maintain this list of all the best open-source ones in this awesome-style GitHub repo. People looking for open-source dictation tools, hope you find something that works for you here: https://github.com/primaprashant/awesome-voice-typing https://github.com/primaprashant/awesome-voice-typing
- rubidium 2mo agoAny that also support translation? How much harder/ easier of a problem is local translation compared to transcription?
- primaprashant 2mo agoIf you're talking about translated text, then that should be super easy. Most of these dictation tool support post-processing with LLM to remove filler words, fix punctuation, etc. I'd imagine you can change the system prompt for the post-processing step to do the translation instead, and you'd get translated text.
- rubidium 2mo agoYea I’m looking local hosted transcription and translation with diarization of 2 (or more ) speakers. This is to speed up collaborative technical work between two teams who speak different languages where want all local processing (assume no cloud access).
- markisus 2mo agoThe post makes it seem like ONNX is CPU only. I've used ONNX runtime to run models on Nvidia GPUs. The runtime can even dispatch to TensorRT. I'm not sure what the performance is on Apple hardware so maybe that was the motivation for moving away from ONNX.
- sipjca 2mo agoTensorRT and CUDA is effectively the same speed as CPU for the speech to text models I was testing via ONNX at a huge binary bloat penalty. WGPU is hard to ship and also equivalent speed or slower. This may not be the case for LLM or other models but the runtimes did not seem well supported for what I needed to do. ONNX is incredibly well optimized for CPU, best in class even, but the other execution providers at least for STT seemed lacking. I did this investigation before creating transcribe.cpp it would have been much more convenient and save me literal months of work. Happy to share the repo and binaries produced as well, but it was mostly throw away work to profile how to ship accelerated ONNX in Handy.
- scoriiu 2mo ago[flagged]
- 0xnyn 2mo agohandy has been invaluable in my workflow, and having a fast, local, c++-based transcription library with first-party ts bindings is incredibleee tysm for shipping this, keep up the great work OP
- jerieljan 2mo agoNice. I did transcriptions on a casual project before that went through something like this. Transcribing videos or audio files with Whisper? Very common. But having to swap it out with Qwen3 or a different family of ASR models? Oops, not as straightforward. For Qwen for example you gotta deal with the forced aligner or it won't be good as subtitles, and then gotta deal with some requirements and considerations if you want to make use of MLX on a Mac or something. Will definitely check this out since it sounds like it eases through the pain of dealing with these.
- bazzingadev 2mo agoHey, thanks for this.
- ctas 2mo agoI'm using Handy on macOS and love it. Unfortunately, hotkeys still doesn't seem to work on Wayland, which make it unusable.
- sipjca 2mo agoYeah I’m working on it, Linux is a big pain point especially Wayland Once things are more or less ironed out on MacOS and Windows a lot of attention will be turned towards Linux I know a lot of Linux PRs are open it just takes me so long to get around and test them. And often multiple different implementations trying to fix similar issues which is a lot of overhead sometimes
- ctas 2mo agoReally appreciate your work. Is there any way people can help? From your last sentence, it sounds like another PR isn't it and the opposite might be needed. But would love to contribute with testing if helpful. I'm regularly jumping between XFCE, KDE, GNOME, Niri, etc..
- sipjca 2mo agoTesters by far as the most needed thing, I do maintain a list of per platform people who help to test so if you drop a GitHub username (or email me) I will add you to the list and ping for help Basically the biggest blocker is me being the sole maintainer and reviewer at the moment and it just ends up taking a lot of time for the scale of the project. Which is why it moves slow and features typically are much slower than someone can vibe code. I know each added feature inevitably has bugs so I try to be careful with them. But also Linux has historically been a minefield, fixing something for someone breaks for someone else so yeah testers really needed. Or anyone with deeper Linux DE knowledge than I have. I’m much more accustomed to server based Linux distros
- boomskats 2mo agoI have a personal fork of hyprvoice[0] which I use almost everywhere now (w/ the big cohere-transcribe running on a local vLLM instance). It does a similar thing, but that's not why I'm mentioning it; I think it's worth looking at because it's a clean reference for the few elegant ways you can implement text injection in modern Linux (wayland). It supports ydotool[1], wtype[2] and "clipboard fallback with clipboard restore". The first two you can probably think of as AHK equivalents - they wire in at the input layer and inject keystrokes when injecting text. wtype is wayland-only and a bit less invasive, ydotool supports non-wayland also apparently, but I haven't tried it. Neither approach provides 'instant text' - you have to watch the text get typed out, and you don't touch your keyboard while it's happening; the clipboard implementation is fallback for a reason as it's the least reliable. The first two work 'well enough' though, and are fairly tunable. The other thing hyprvoice does in probably the most linux-friendly and universal way is the 'hotkey handling'. The server creates a socket in /tmp that the cli can then ping when the user triggers the start/stop/cancel, and they do this by binding whatever their DE's keyboard shortcut mapping mechanism is to trigger `hyprvoice toggle` as a background shell command. This works extremely well and is much cheaper than you'd intuitively think coming from Windows. This way you don't have to interface with DE-specific global keyboard listeners etc, but leave that to the WM (that's not to say that your installer couldn't prompt the user to configure the keyboard shortcut for them with their detected WM, you just wouldn't do it in the software itself). I haven't actually looked at your project in too much depth yet as I have a solution for this already, so apologies if none of the above is news to you. Hope it helps though - happy to poke around and contribute something if the gap's still there. [0]: https://github.com/leonardotrapani/hyprvoice https://github.com/leonardotrapani/hyprvoice [1]: https://github.com/ReimuNotMoe/ydotool https://github.com/ReimuNotMoe/ydotool [2]: https://github.com/atx/wtype https://github.com/atx/wtype
- JeremyHerrman 2mo agoAnother happy user of Handy here! After seeing so many *subscription based* transcription apps all wrapping *open source models*, finding Handy was a real delight and I'm happy to see the author keep on building!
- luciana1u 2mo ago[flagged]
- hackrmn 2mo agoIs transcription a form of _inference_ though? I mean I see the word being thrown around and I understand what it means (or at least I think I do) in context of LLMs doing the thing that they do -- intelligently predict the next token, but do speech-to-text models do that?
- l-albertovich 2mo agoThanks CJ, you've put some pretty cool things out there!
- rmunn 2mo agoLooks very cool. One thing I have been looking for, which this doesn't seem to cover (at least I didn't see any mention of IPA in the model documentation), is a way to transcribe unknown languages phonetically, using the International Phonetic Alphabet to spell them (sound-based spelling rather than meaning-based spelling). I know several linguists doing research on minority languages (fewer than 10,000 speakers in some cases), which are small enough that they will never have enough effort made towards training language-specific models in that language. Are there models I'm not aware of that are trained for this task? Taking audio in an unknown language, and rather than identifying the language, just transcribing the sounds to IPA? That would not be useful to most people, but it would be a Godsend to many, many linguists working with minority languages around the world.
- shenberg 2mo agoWe take a lot of shortcuts when speaking, it's actually much harder to transcribe phonemes than to transcribe words, even when aware of the language being spoken. Some models have been trained for the task (e.g. look at https://huggingface.co/spaces/KoelLabs/IPA-Transcription-EN https://huggingface.co/spaces/KoelLabs/IPA-Transcription-EN ), but the error rate is really high.
- rmunn 2mo agoThere are, broadly, two kinds of audio recordings that linguists want to transcribe. One is native speakers telling traditional stories, where they're speaking naturally and taking the natural shortcuts (such as "wanna" and "gonna" in English). The other is native speakers reading words (or short example sentences) very carefully and distinctly, so that the linguist can listen to the recording over and over to learn how to pronounce the word right. In those recording, they'll say "want to" and "going to" rather than "wanna" and "gonna". Thanks for the pointer; I'll check out that model and see if it handles the "slowly and carefully" type of recording better than the "natural speaking" type. (And depending on what kinds of errors the model makes, even the recordings where it makes errors can prove useful: for example, a linguist studying regional variations in speech would want the model to produce the IPA for "gonna" rather than "going to").
- 2mo ago
- kelvinjps10 2mo agoIs there something but for transcribing what you watch like videos and not your microphone? Samsung has this in my phone and it's useful for language learning. (Thought is not that accurate)
- JesseHowell 2mo agoReally cool that every model is actually tested for accuracy instead of just claiming it works, I think alot of 'we support everything' tools skip that step. How are you checking accuracy for models that don't have an obvious "official" version to compare against?
- sipjca 2mo agoEvery model with open weights has some code which can be used to inference it. So we download the published weights and run against inference library they suggest, be it transformers, Nemo, etc
- kmfrk 2mo agoWell this almost seems to be to good to be true. :) I assume this is going to make maintaining SubtitleEdit a lot easier from now on, too: https://github.com/SubtitleEdit/subtitleedit/ https://github.com/SubtitleEdit/subtitleedit/. Anyone know a good Windows app that's just a window that transcribes - and translates - whatever goes through your output device, and not the microphone like most apps do?
- solarkraft 2mo agoOh, I like this! I’ve been looking into locally hosting a transcription API server and came away feeling pretty close to the problem statement. The things most frequently lacking were streaming support (which I’m so glad this has!) and the support for special words to boost during recognition (which I guess there’s some hope they might add???).
- embedding-shape 2mo ago> I’ve been looking into locally hosting a transcription API server I've been hosting my own since whisper.cpp appeared on the scene, thrown up on a server with a 3090ti. Even if there is better/faster stuff out today, it just keeps on working without any issues, the weights are tiny and it's faster than I could need. This is basically what you need to get this working today: MODEL="/home/user/projects/ggml-org/whisper.cpp/models/ggml-large-v3-turbo.bin" WHISPER_SERVER_BIN="/home/user/projects/ggml-org/whisper.cpp/build/bin/whisper-server" "$WHISPER_SERVER_BIN" --model "$MODEL" --language en --host 127.0.0.1 --port 7812 Very simple stuff, throw it on some local homelab server and now you have a local transcription API :) Might need to play around with some of the inference parameters, but once you've locked them in, seems to work really well.
- solarkraft 2mo agoDoes it support streaming? I find that this is the #1 thing missing from almost all implementations.
- sipjca 2mo agoword boosting will probably come on a much longer time horizon, but streaming is here! I'm really hoping someone either contributes a good server example to the codebase (and is willing to help with issues) or use transcribe.cpp or the bindings to create a robust server in another language :) would be happy to link it from the main project directly as well
- jech 2mo ago[dead]
- tangsoupgallery 2mo ago[flagged]
- maverickaayush 2mo ago[dead]
- paweladamczuk 2mo agoI've been using this one for a week for local transcription, working pretty well so far
- hermes_scanner 2mo ago[flagged]
- tlamponi 2mo agoHas anybody experience with using this with strong dialects, like e.g. bavarian-family (German) based ones? Or other languages one too, as I'd figure basic behavior and approaches to improve detection of such is often similar in principle for dialect style variants of a language. I mean, I naturally should try myself, and plan to do so, but slightly lower on my free time priority list and I figured someone else might have explored this already.
- terhechte 2mo agoI'm using this in one of my side projects, Emyn ( https://github.com/terhechte/Emyn https://github.com/terhechte/Emyn ) a macOS virtual camera app for composing camera video, app windows, backgrounds, effects, notes, and captions into a polished live presentation feed. It works very well, the integration is much easier than before, users have model choice. So happy that this exists!
- sipjca 2mo agoWow, it's amazing to hear that even though I released this so recently people are already using it properly! Thanks! Please let me know any issues you run into
- nojvek 2mo agoI use Handy everyday. It’s a great project. Thank you for making offline ASR work great on modern machine.
- sorenjan 2mo agoWhy not include transcribe-cli in the release archives to make it easier to use for people that can't compile it themselves? I downloaded the Cuda version but it's only the dll files, I don't really want to have to deal with Cuda SDK, I doubt most people want to.
- sipjca 2mo agoRight now I intend to maintain this as a library. The examples are just that, examples for programmers/agents. If someone in the community wants to step up to maintaining release binaries I will gladly have that support, it's just impossible to do as a sole maintainer
- sorenjan 2mo agoThat's up to you of course, but is it that much more work to compile the cli binary at the same time as you compile the libraries? How am I supposed to actually use the Cuda binaries available in the releases section, through a separately downloaded Python wheel?
- sipjca 2mo agoIt's not that much extra work to compile, the extra work comes from the maintenance and feature requests. By not shipping the binary directly I am defending my time until other contributors want to step up and maintain things. I am one person with limited time and I don't want to spend all of it in front of a computer Yes you use them through a wheel. If you have specific questions on packaging and how to use things lets move it over to the discussions/issues in the repo itself so it can be more broadly accessible to more people and we can make the packaging of the library as useful as possible
- leumon 2mo agoThank you! I found this to work much better then the old transcribe-rs lib. I updated my Offline Voice Input App to also use the new library and it's much faster now: https://github.com/notune/android_transcribe_app https://github.com/notune/android_transcribe_app
- apitman 2mo agoJust wanted to say I started using Handy last week and I love it. It might single handedly cure my RSI. Well, hopefully double handedly.
- fenix1851 2mo agothanks man! Been searching for solution like yours :)
- alabhyajindal 2mo agoCongrats! I just tried Handy again which now uses transcribe.cpp and it works brilliantly. Love the streaming output from Parakeet Unified EN 0.6B. I remember using Handy about a year ago and it's amazing to see the improvements!
- vardalab 2mo agoOne thing I find that's missing a lot or at least I haven't come across other than commercial offerings like AquaVoice is a decent injected technical vocabulary so that the initial transcript requires minimum cleanup afterwards. Because I mostly use these tools to essentially ramble at the command line with coding agents. So there's a lot of technical terms that don't translate well. Like OpenBao comes out as open bowel sometimes,lol. That necessitates significant cleanup prompt or background text available to the cleanup llm, usually in the form of screenshot or something that gets converted to text but that in turn requires good hw for speed to be almost imperceptible. For example m5 max turns cleanup into a noticeable delay while 5090 is decent. Only way I have found that's relatively easy to inject technical vocab is to use whisper, but limited, I think to about 220 or so tokens. Whisper has sort of like a priming prompt where one can put in a bunch of technical words and it will try to recognize those. But again, that's limited to small number tokens. And that limits one use a relatively slow, by today's standards, whisper.cpp. I benchmarked it across a bunch of different hardware that I have available, and Whisper gives decent performance as far as speed goes only on a pretty top-end GPU, such as a 5090 or 4070, like for example on Strix Halo, it's still relatively slow for longer transcriptions because I prefer just a stream of consciousness ramblings for minutes and then that being transcribed and cleaned up versus short sentences. So in that scenario something like 5090 really is good because the cleanup prompt runs fast using usually Qwen 3.6 MOE model. Whisper on 4070 itself is about 0.7 seconds for two or three minute transcription. So the total wait time for a three-minute transcription is roughly a second, or a little bit more than a second, so totally acceptable. But it does take decent hardware, and it grows to be double that on if running totally local. Well, in my case, it's all local, but it's my own hardware all over the place, but truly running on laptop, it's much faster using Parakeet, but then the cleanup is the bottleneck. Anyway, it's just my experience messing around with this for the last year. I did start using AquaVoice, but their speed was exceptional, and tech vocab was exceptional, but they would have some annoying delays occasionally, and I didn't like paying the money and sending sensitive topics and screenshots into the cloud, and I had hardware, so my local solution is basically almost as good as commercial one. But I think they train their own model. So what I'm doing is I collect all the samples of my transcriptions, and I am slowly building my own data set that hopefully at some point when I get energy I will find some way to fine tune something.
- ktosobcy 2mo agoYet another happy user of Handy - one of the best applications out there, kudos! <3
- larnon 2mo agoDoes the whisper model allow for entering context(which improves the accuracy greatly) as it does on Whisper.cpp?
- sipjca 2mo agoyes
- PalmPilotProMax 2mo agoAny plans of including the VAD step into this? When whisper.cpp added Silero it really smoothed out the UX.
- winterscott 2mo agoThe numerical validation and WER testing are what stand out to me here. A lot of local ASR projects claim broad model support, but it is often difficult to know whether the converted models still match their reference implementations. Having one embeddable engine across Vulkan, Metal and CUDA, along with maintained language bindings, addresses a real distribution problem. How stable do you expect the C API and model format to be after v0.1? In particular, could an application eventually switch between different model families without needing model-specific preprocessing code?
- aghilmort 2mo agofully local transcribe.cpp in your browser at https://mic.sloop.ai/ https://mic.sloop.ai/ github repo soon, finalizing ggml / emscripten magic <> simd, pthreads, etc. parakeet for now, adding models, webgpu, wasi server, & electron for native