5 ms·
Speech Recognition and TTS in less than 500kb
- sgt 3mo agoGreat work!
- sgt 3mo agoQuick link to the video where he demos it: https://www.youtube.com/watch?v=kMliOFYBiz4 https://www.youtube.com/watch?v=kMliOFYBiz4
- 8bitsrule 3mo agoThanks for that ... impressive!
- JSR_FDED 3mo agoAmazing that this works. As an aside, and I appreciate this is just a demo, if the use case is to get a device to join a WiFi network - would a single or double line lcd with 3 buttons not be cheaper than 520KB?
- pxx 3mo agothe target rp2350 is a sub-$1 chip. a 16x2 LCD module is over $1. but more importantly, you might have this much ram sitting around unused on whatever you're building anyway.
- andai 3mo agoIs that Microsoft Sam? :) (Also, I know it's besides the point but this might be the most painful way to connect to Wifi physically possible. "Make normal everyday tasks slow, tedious and painful" is a bit of an odd choice for a product demo.) Say, speaking of Sam, what were the memory requirements for SAM (Software Automatic Mouth) on C64. I guess they were not more than 64K? Although, the bulk here is probably for the speech recognition, not the TTS. (And this one does sound a little nicer :) Browser demo of a reversed SAM: https://discordier.github.io/sam/index.html https://discordier.github.io/sam/index.html
- raphlinus 3mo agoThere are two Sams here, the Microsoft one and the C64 one. I don't believe there's any connection between the two other than the name. According to [1], the weight of a modern runnable version is around 39k. The ratio of how good it sounds compared to how much computing power it uses is ridiculous. The C64 has ballpark 3 orders of magnitude less CPU throughput as an RP2350, and the codebase uses an impressive array of tricks to do actual formant synthesis (barely) and a pretty refined form of Elovitz text to phoneme conversion. One of my favorite tricks is its up and down bouncy pitch, which is not random, but based on the opposite contour as the first formant. It's simplistic but enough to make it not sound like a robotic monotone. I've been playing around with this some myself and SAM is an inspiration, along with other landmark systems like MITalk (predecessor to DECtalk), SP0256, and other. I believe it's possible to use modern techniques to get pretty good sounding speech in, say, 64k and 10% of the throughput of a RP2350. It's really cool to see projects like OP, especially under permissive license. [1]: https://simulationcorner.net/index.php?page=sam https://simulationcorner.net/index.php?page=sam
- 0xnyn 3mo agongl, it looks incredible
- zarmin 3mo agoThank you for this. I love your work on Curb Your Enthusiasm.
- toilet 3mo agoWhat work?
- stavros 3mo agoProbably a joke that the author looks like someone on the show? I'm puzzled as well.
- shermantanktop 3mo agoPossibly Cousin Andy? https://curb-your-enthusiasm.fandom.com/wiki/Andy_David https://curb-your-enthusiasm.fandom.com/wiki/Andy_David Played by the great Richard Kind, who my wife swears she saw on the Highline in NYC.
- clayhacks 3mo agoI made a little python wrapper around it to serve an HTTP endpoint that’s OpenAI/elevenlabs compatible https://github.com/clayrosenthal/bootlegger https://github.com/clayrosenthal/bootlegger
- senkora 3mo agoWow, it seems like this might beat out flite for very-low-memory TTS? I ended up abandoning a project of mine because I couldn't get high enough quality or low enough memory usage out of flite, so I'm very excited to try this out. Flite for comparison: https://github.com/festvox/flite https://github.com/festvox/flite
- jedberg 3mo agoDo you have any accuracy benchmarks? I’ve worked in this space. TTS in a small footprint isn’t the hard part —- it’s doing it accurately that’s hard. Although for the use cases OP is targeting, lower accuracy may be good enough!
- amelius 3mo ago> I’ve worked in this space. TTS in a small footprint isn’t the hard part —- it’s doing it accurately that’s hard. This actually holds for everything in AI.
- jedberg 3mo agoVery true!
- kamranjon 3mo agoIf you look at this chart here it seems the tiny model has a WER of ~12%… not sure about the micro model: https://github.com/moonshine-ai/moonshine#when-should-you-choose-moonshine-over-whisper https://github.com/moonshine-ai/moonshine#when-should-you-ch...
- yorwba 3mo agoThat's the error rate for STT, not TTS. TTS is generally easier than STT because you only need to produce one valid pronunciation and don't need to handle variation within and between individuals.
- stfurkan 3mo agoIt looks great, thank you! I'll see if I can use it for my in browser AI assistant project's ( https://aidekin.com https://aidekin.com ) voice part. It's currently using Nemotron-3.5-ASR and supertonic-3 but overall it requires 1.2gb download.
- orliesaurus 3mo agoI installed the command line version using uv uv init uv add moonshine-voice uv run moonshine-voice mic --language en super nice to be able to run it to test it like this good job on a clear readme.md tbh
- pwgawron 3mo ago`uvx moonshine-voice mic --language en` That is even simpler.
- t0mpr1c3 3mo agoVery cool. I've done TTS on a 32K Arduino but it was pretty croaky. https://youtu.be/ErGDboTpwM0 https://youtu.be/ErGDboTpwM0
- smcameron 3mo agoFor TTS I wonder how this compares to nanotts[1] with the en-GB voice, which is sort of unreasonably good. [1] https://github.com/gmn/nanotts https://github.com/gmn/nanotts
- dwa3592 3mo agothis is good to see. i also trained a stt under 500kb for sub dollar chips. it had about 20 words that it could understand(like start, stop, left, right, go, up etc) and then the spell mode where you could say the word spell and then say the individual english alphabets and close with spell. it was super fun to work on. these tend to be extremely unstable though, like confusion between p and t (at least for my accent). will have to try this one now.
- NooneAtAll3 3mo agoI remember someone training smart kettle to use its speaker as microphone
- laidoffamazon 3mo agoIIRC the Alexa enabled voice remotes also used a similarly small model though perhaps not this small
- schoen 3mo agoCould you get people to use the NATO phonetic alphabet for the spelling part? I suppose a challenge is that many people don't know the whole thing, even if they're aware it exists.
- jodrellblank 3mo agoNATO phonetic is to be understandable over a noisy radio channel, if you want just distinct sounds then Talon Voice users settled on shorter ones easier to use all the time: air a bat b cap c drum d each e fine f gust g harp h sit i jury j crunch k look l made m near n odd o pit p quench q red r sun s trap t urge u vest v whale w plex x yank y zip z
- close2 3mo agoInteresting, that some words don't start with the letter they represent.
- deleted 3mo ago[deleted]
- userbinator 3mo agoThis looks like an extreme point for AI-based TTS, as formant/tract modeling synths tend to be more accurate if you want TTS in a tiny amount of compute, but sound distinctly robotic. TTS (neural diphone synth @ 16 kHz) ~1.8 MiB voice pack This is in the realm of Microsoft Sam.
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- HarHarVeryFunny 3mo agoPresumably it's not, but the TTS voice in the video sounds to me more like formant synthesis than diphone - it reminds me of my DECtalk. The project credits does mention espeak (which is formant based) as well as various other TTS projects, although it sounds like they are only using the pronunciation part of espeak, not the voice synthesis. https://github.com/moonshine-ai/moonshine#acknowledgements https://github.com/moonshine-ai/moonshine#acknowledgements
- Joel_Mckay 3mo agoIt certainly sounds similar, but seems more nimble with phonetic pronunciation in the demo. Having it run on a pico would be pretty impressive =3 http://cmuflite.org/ http://cmuflite.org/ https://github.com/festvox/flite https://github.com/festvox/flite
- HarHarVeryFunny 3mo ago> Having it run on a pico would be pretty impressive Yes, although relative to the DECTalk DTC01, a Pi Pico is a beast ! Pico : dual core ARM @ 133 MHz, 2MB flash, 264K RAM DECTalk: 68000 @ 10 MHz + TMS 32010 @ 20 MHz (5 MIPS), 256K ROM, 64K RAM
- jjcm 3mo agoThe voice activity detection alone here is compelling - very useful for doing things like highlighting a speaker who's transmitting in realtime. At that rate the impact on perf will be so minimal that you could easily run it in the browser across devices.
- irfan_99 3mo agovery nice I love it
- irfan_99 3mo agoIs the dataset open
- gitgud 3mo agoSo at that tiny 500kb size I imagine it could be compiled to web assembly, and run entirely in the browser right? Couldn’t find a link, is that hard to do?
- hahahaa 3mo ago500k memory but not sure about disk.
- scoriiu 3mo agoShould be very doable. I ship a small CNN in a browser extension via onnxruntime-web and the model weights were never the bottleneck, the runtime was. The wasm backend adds a few MB of runtime before your first inference, so a 500kb model with a lean hand-rolled wasm build would actually beat most "tiny" browser ML deployments in total download. One gotcha if anyone wants this in a Chrome extension: MV3 requires 'wasm-unsafe-eval' in the CSP for any wasm at all, which surprised me the first time a build that worked fine as a web page died silently as an extension.
- salamo 3mo agoYeah, I also found that for ultra low footprint models ORT is a big portion of the total payload, because it contains logic for general ONNX graph operations. In my case I found that ORT alone was 3.4MB over the wire, so I swapped it out for a tiny wasm that was 850x smaller and only contained the operations I needed: https://blog.lukesalamone.com/posts/creating-tiny-semantic-search/ https://blog.lukesalamone.com/posts/creating-tiny-semantic-s...
- scoriiu 3mo agodid you skip simd just because the model's tiny? naive conv perf is honestly the only reason i haven't done exactly this for the cnn
- salamo 3mo agoYeah, the model is small enough that inference is already basically instant for my usecase (only 6 transformer layers for the blog search).
- walrus01 3mo agoGiven the tiny size of this, I wonder about possible future integration with esphome compatible hardware https://esphome.io/ https://esphome.io/
- KennyBlanken 3mo agoI suppose, but for home automation, esps are best for getting the audio to something more powerful. If this lets a raspberry pi do voice recognition really fast, that alone is worth it.
- nserrino 3mo agoVoice is one of the most latency-sensitive modalities in AI. Moonshine is doing awesome stuff
- 1vuio0pswjnm7 3mo agoCompare https://github.com/ggml-org/whisper.cpp/ https://github.com/ggml-org/whisper.cpp/
- nutanc 3mo agoThis is awesome. I am trying to build a full scale ASR system within 20-25MB. Now that we have Claude code to run experiments, I have started running some experiments. Promising results so far. First realization is that you can capture the nuances of speech in just 3300 embedding vectors(786d). This sequence can be decoded with a small CTC system to get text. Next experiments are on reducing the 768 dimension space into a 64D space. Thats also show some promising results. Hooking up my system so that the agent blogs the results everyday[1]. So my research "claw" setup does the experiments and posts results which I check in the morning and adjust the experiment direction as needed. Its not fully automated yet, but almost there. [1] https://blog.trulm.com/posts/speech-as-independent-parts/ https://blog.trulm.com/posts/speech-as-independent-parts/
- lunixbochs 3mo agoI think Google's Conformer paper is SOTA at the <30M model size, where I think they put an incredible amount of flops into a 10M param model to reach around 2% lsc clean (the whole model and RNN decoder were trained domain specific to librispeech here). I think my small Talon models are next, around 3% lsc clean at ~28M (greedy CTC decoding, no external encoder, no LM, not trained in a domain specific way). I reached around 6.5% at 10M. I've been working on some new baselines I want to release soon as public artifacts. This article is inspiring me to try pushing the param size down a bit. I suspect we can do large vocabulary end to end in the <5M range.
- almogo 3mo agoStt/tts systems always seem to me so promising, but I pretty much never use voice to interface with a computer. Sometimes instead of typing on my phone, I use a voice dictation. I would be keen to use voice to control Claude code, but I've always felt that the way I speak is different from the way I write good prompts. Fishing for anecdotes here, does anyone have any good tts/stt experiences?
- k9294 3mo agoI'm founder of ottex.ai, I use stt pretty much all the time when work with AI and quite often for communications to draft emails and chat messages. I started ottex half a year ago after I tested gemini 2.5 flash native audio support. I was blown away by the quality of transcripts and decided to built an app to use it myself. Currently the default model in the app is Gemini 3 flash, but you can connect to 9 providers and God knows how many models to play with. I would suggest you to try this models for ai prompting: - Gemini 3 / 3.5 flash - Soniox rtt v5 - Mistral transcribe v2 - assembly 3.5 pro
- DrSiemer 3mo agoOne of my side projects is a tool that lets you control your entire system with STT. It's built on Whisper and supports hot swapping custom profiles, so you can add easy commands for any software. I intend to use it to work on low stakes vibe coding projects while I'm doing other stuff. Todays LLMs are a lot better at interpreting rambling dictation with mid-message corrections. There are a few paid programs out there that do the same, but they made my vibe slop sense tingle and are not aimed at development.
- arend321 3mo agoI do a ton of coding (codex) with a tts/stt wrapper. During walks, cycling, in the car. Not every task is suited to this style of interaction, but many are. Long form codex replies are condensed, code blocks are suppressed all in the name of making it work for tts feedback. So it works best on well defined projects with guardrails, where you know the agent can perform well.
- fragmede 3mo ago
- Ecko123 3mo ago[dead]
- deleted 3mo ago[deleted]
- stanko 3mo agoThis is really impressive. If I get time, I would like to try compiling it to WASM. This would allow me to swap my robot poet’s native browser voice synthesis for it. Not sure if it is worth it, but it will be fun to play around with. Edit: typo [0] https://muffinman.io/bard/ https://muffinman.io/bard/
- agnishom 3mo agoWill it be able to understand my English with an Indian accent?
- jkwang 3mo ago[flagged]
- Kyuren 3mo agoThis makes me want to have a server room with 5 of these around my house and control everything that house in LabRats
- hermes_scanner 3mo ago[flagged]
- yako21000 3mo agowow now that's a really tiny tts model. is there a comparison to https://github.com/kyutai-labs/pocket-tts https://github.com/kyutai-labs/pocket-tts ?
- 1saadcodes 3mo agoI've a local dictation workflow for coding, and one thing I've learned is that transcription accuracy is only half or even less than half the problem now. The other half is latency. Once the delay gets low enough that you stop noticing it, voice input starts feeling much more natural. It'll be interesting to see where this lands compared to Whisper-based setups for continuous dictation