7 ms·
Neural audio codecs: how to get audio into LLMs
- mondainx 1y agoThanks for sharing this well written post that I will share with my team; we just recently started using audio/voice in our AI suite and the details herein will be helpful and informative.
- amelius 1y ago> Many LLMs have voice interfaces, but they usually work by transcribing your speech, generating the answer in text, and using text-to-speech to read the response out loud. That’s perfectly fine in many cases (...), but it’s a wrapper, not real speech understanding. But I can say the same about tokenization. LLMs first convert groups of characters to tokens, then use that to generate tokens, and then convert the tokens back to characters. That's not real understanding! If LLMs are so smart, we should be able to skip the tokenization step.
- Workaccount2 1y agoNothing is real understanding because we have no benchmark for understanding because we don't mechanistically know what understanding is. The best we have is people "vibe knowing" a benchmark that they made up on the spot.
- vvolhejn 1y agoThere's a great blog post from Sander Dieleman about exactly this - why do we need a two step pipeline, in particular for images and audio? https://sander.ai/2025/04/15/latents.html https://sander.ai/2025/04/15/latents.html For text, there are a few papers that train the tokenization and language model end-to-end, see: https://arxiv.org/abs/2305.07185 https://arxiv.org/abs/2305.07185
- trollbridge 1y agoAn ongoing question I have is why effort wasn't put into tokenising speech (instead of transcribed words) and then making an LLM out of that. There are huge amounts of speech available to train on.
- MichealCodes 1y agoI don't think we've had the transformer moment for audio training yet, but yes, in theory audio-first models will be much more capable.
- trollbridge 1y agoParticularly interesting would be transformations between tokenised audio and tokenised text. I recall someone telling me once up to 90% of communication can be non-verbal, so when an LLM sticks to just text, it's only getting 10% of the data.
- benob 1y agoAudio tokenization consumes at least 4x tokens versus text. So there is an efficiency problem to start with. Then is there enough audio data to train a LLM from scratch?
- trollbridge 1y agoStart an MVNO that offers cheaper phone plans and and train on all those phone calls. There are big libraries of old speeches. Simply capture all all current radio/tv transmissions and train on that (we've already established copyright doesn't apply to LLM training, right?)
- miki123211 1y ago> Start an MVNO that offers cheaper phone plans and and train on all those phone calls. q: What is 2+2? A: The warranty for your car has expired...
- 542354234235 1y agoDon't we have tens of thousands of hours (hundreds of thousands?) of closed captioned tv shows and movies? How many hours of news broadcasts with transcripts do we have? Maybe I just don't understand what is needed, but it seems like we have a lot of data to work with.
- bkitano19 1y agoAwesome post!
- krackers 1y agoIndeed, the title undersells it and I'm glad I didn't skip over it, the article is basically an information-dense but approachable summary of audio generation.
- robviren 1y agoThis has got to be one of the most visually pleasing explanations I have seen of these concepts. Congrats! I attempted some similar VQ-VAE work instead trying to tokenize rendered text. I was curious if I could make a visual llm working on 10 pt rendered font, but I also tried using PDF sources. The basic idea was to do what more advanced diffusion image models can do where they generate images of text. Make a specific image text diffusion model to do completions. Further I wondered if I could embed things like document type and language so you could have a latent representation of text more abstracted than current dictionary tokenizers. Learned a lot and thought it was all beautifully displayed in this post.
- crazygringo 1y agoThis is fascinating. Obviously working directly with audio is vastly more complex than with text. But it is very exciting to see how part of making LLMs work natively with speech, is finding a codec that is maximally efficient at encoding speech. I even have to wonder if, at some point, we ultimately create a popular voice codec usable with LLMs based not on the Fourier transform or similar, but rather on some kind of set of physical parameters describing vocal cord shape, tongue position, throat/chest/mouth shape, etc. I can imagine such a model being arrived at statistically (determining the necessary number of parameters), and then almost becoming "hard-coded" as a standard since human anatomy doesn't change much there, beyond certain ranges. I think it's called formant speech encoding, and it would be interesting if LLMs wind up massively advancing that field. Since I think historically it's had to do more with speech synthesis than audio compression.
- quinndupont 1y agoThere’s a long history of attempts at artificial speech that take this approach, recreating mouth parts and vibrating air. They are all pretty silly, like this work, which fails to understand how writing isn’t just a derivative of speech.
- crazygringo 1y ago> They are all pretty silly, Huh? How? > like this work which fails to understand how writing isn’t just a derivative of speech. The whole point of the article is that writing isn't just a derivative of speech. It's in the introduction.
- duped 1y agoIn speech coding/synthesis this called a "source-filter" model (decompose speech production into a sound generator in the vocal folds and filter in the vocal tract, and parameterize them) and it's actually older than Tukey and Cooley's rediscovery of the FFT.
- vvolhejn 1y agoAuthor here, thanks for the kind words! I think such a physics-based codec is unlikely to happen: in general, machine learning is always moving from handcrafted domain-specific assumptions to leaving as much as possible to the model. The more assumptions you bake in, the smaller the space of sounds you can model, so the quality is capped. Basically, modern ML is just about putting the right data into transformers. That being said, having a more constrained model can also lead to some really cool stuff. The DDSP paper learns how to control a synthesizer to mimic instruments: https://arxiv.org/abs/2001.04643 https://arxiv.org/abs/2001.04643 You could probably do something similar for a speech model. The result would not sound as good but you could get away with much fewer parameters, because much of the modelling work is done by the assumptions you put in. Compare also KokoroTTS, a tiny TTS that's so tiny because it uses a handcrafted system to turn text into phonemes, and then just synthesizes from those phonemes: https://huggingface.co/spaces/hexgrad/Kokoro-TTS https://huggingface.co/spaces/hexgrad/Kokoro-TTS
- bob1029 1y agoWhy not normal audio codecs? How are JPEG and MP3 (i.e., DCT/MDCT) not a reasonable way to go about tokenizing spatial and time domain signals for these kinds of models? Each MP3 frame is entirely self-contained and can completely reconstruct a few tens of milliseconds of original audio. It does not require other frames to do this. I think this is the most important element. At 128kbps CBR, each MP3 frame is ~418 bytes and covers ~26 milliseconds of time. This is a reduction of 10-11x over the raw PCM waveform. MP3 is also designed to eliminate the information that humans don't seem to care about. I don't know if it's possible to use 400 byte tokens in a transformer model, but I would be very compelled to try.
- PaulDavisThe1st 1y agoThe approach in TFA encodes into a 32 dimensional space. I suspect this is significantly more dimensions than any psycho-acoustic compression algorithm uses. Also, throwing away information that our hearing systems can't process very well is not particularly useful if your goal is speech (or more generally, audio) synthesis from scratch.
- bob1029 1y ago> throwing away information that our hearing systems can't process very well is not particularly useful if your goal is speech (or more generally, audio) synthesis from scratch. I'm not sure I follow. If there is a set of tokens that the average human cannot perceive, why wouldn't we want to eliminate them from the search space? Who is the target audience for this model?
- CaptainOfCoit 1y agoMaybe that things outside our audible range could impact/influence things inside of our audible range?
- 542354234235 1y agoI imagine it would be like if there were Rosetta Stones of text, written with a language you could read and a language you couldn't. For your purposes, discarding the text you can't read would be fine and you wouldn't lose anything. But if you were ingesting a bunch into an LLM, the additional text would give the LLM more context and help it make connections and relate words more accurately, even if you never were going to have it output anything in the language you don't understand. The inaudible sounds add context and additional datapoints on how the audible sounds are related.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- miki123211 1y ago> Try asking any of them “Am I speaking in a low voice or a high voice?” in a high-pitched voice, and they won’t be able to tell you. I wonder how much of that is LLMs being bad, and how much is LLMs being (over) aligned not to do it. AFAIK, Chat GPT Voice mode had to have a lot of safeguards put on it to prevent music generation, accent matching (if you sound Indian, it shouldn't also sound Indian), and assuming ethnicity / biasing based on accents. It doesn't seem that impossible to me that some of these behaviors have been aligned out of these models out of an abundance of caution.
- sbrother 1y agoI don't think it's just safeguards; they really don't seem to understand pitch at all. I tried asking ChatGPT's advanced voice mode to recognize a tune I was humming, and it insisted it was Beethoven's 5th -- multiple times. I think it must have basically tokenized my humming to "dun dun dun duuun".
- bigzyg33k 1y agoadvanced voice mode operates on audio tokens directly, it doesn't transcribe them into "text tokens" as an intermediate step like the original version of voice mode did.
- cubefox 1y agoBut they behave just like models which use text tokens internally, which is also pointed out at the end of the above article.
- bigzyg33k 1y agowe don't know if that's due to inherent limitations of the tokenisation of audio, or a byproduct of reinforcement learning. In my own usage, I noticed a significant degradation in capabilities over time from when they initially released advanced voice mode. The model used to be able to sing, whisper, imitate sounds and tone just fine, but I imagine this was not intended and has subsequently been stunted via reinforcement learning. I don't find the articles argument that this is due to tokenisation convincing.
- quinndupont 1y agoY’all need to learn about the history and development of spoken language and writing. Writing isn’t just a copy or derivation of writing. LLMs work because of the conceptual characteristics of writing (consider the distinctions between ideographic, logographic, alphabetical…). What a sloppy mess! Read some Wittgenstein and Goodman, but especially Derrida who calls this logocentrism.
- daxfohl 1y agoAnother interesting thing here is that the model presumably has some understanding of the passage of time. That's one thing that can be odd about chat models, in that they will respond the same no matter whether you respond a second later or a month later. I think even for text models, "streams" could be useful. Perhaps if the LLM sees too long of a pause after explaining something and asking a question, they could interject a "do you need help?" or something. Pure chat GPTs don't have that ability.
- daxfohl 1y agoI wonder if a linear-space, constant-time model like RWKV or S4 would work better here. For audio, I wouldn't think you'd need long range context, and all-to-all mapping seems like overkill. Maybe a transformer could be running in parallel, but much lower frequency, where the linear model feeds it "summary" tokens once per second, whose information would mostly be "text", but also some hint of emotion and other cues. Then the output of this could be fed back to the linear model so that it would know what it was saying and with what emotion. Basically the transformer would be the low frequency long range context thinker (and feeler), and the linear model would translate that to and from phonetics. They'd be trained in parallel, so those transformer tokens would attain meaning at training time, not something that would have to be pre-defined. So it'd still be purely phonetic e2e, no direct translation to text. It could even end up being a good way to compress text for LLMs, since low-value words might have smaller representation in the token. Probably would never reach the level of text based LLMs for logic and code and such, but that somewhat parallels humans anyway; it's pretty hard to explain an algorithm in detail in plain conversation.
- tehnub 1y agoWrite this paper please!
- daxfohl 1y agoIf anyone wants to buy me some GPU time I'd be happy to try it out! Fair warning: my only experience in deep learning thus far was training a CNN to count dots on an image, which worked semi reliably up to 8, when the image was perfectly square black "dots" on a perfectly white background.
- lxe 1y agoI've been messing around with Higgs Audio that actually uses the delay pattern. It has to apply it and then unapply it after the generation. I noticed it's actually really hard to chunk and stream audio correctly when you need to apply and reapply these patterns essentially to the "entire" output.
- mmaunder 1y agoThanks for posting, I wasn't aware of Kyutai and it seems your work is perfect for something I'm working on.
- croemer 1y agoTypo: "not even a the length of one word"
- vvolhejn 1y agomerci, will fix tomorrow
- deleted 1y ago[deleted]
- Rickasaurus 1y agoI wouldn't mind so much if they cheat on the way back but listen in earnest. There are use cases like teaching language where having the AI understand the sounds carefully matters a ton.
- casey2 1y agoI can't wait for LLMs to actually understand how they and you are speaking. It's going to be so cool when an AI can correct your second language pronunciation or laugh at you for making a silly sound. The usecases and value will explode when that happens 100%
- orena 1y agoHow many epochs did you train with ? 100k hours is not a lot for an LLM, Feels like bitter lesson
- vvolhejn 1y agoI train for 1M steps (batch size 64, block size 2048), which is enough for the model to more-or-less converge. It's also a tiny model for LLM standards, with 150M parameters. The goal wasn't really to reach state of the art but to show how the performance of a single language model architecture can be vastly different when you just change the tokenizer.
- singularfutur 1y agoTo get around state of the art, how many parameters would be needed with your approach?
- liqilin1567 1y agoOut of curiosity, would it be possible to attach pitch, emotion, tone info as text-based metadata to each word during ASR, so that the asr output retains these metadata?
- Razengan 1y agoMan, one of the best uses of all those AI algorithms based around finding similarities between stuff, would be to give you actually relevant recommendations for music. All the streaming services are shit at it. They can't do much beyond shallow similarities or hardcoded recommendations that are probably just based on manually-entered keywords like the genre etc. Has that already been done? Or is it yet another of those what-could-have-been utopian things that got crippled before it was born because of corporate overcontrolling/overcautiousness (not being able to train on copyrighted music) Maybe some open-source project could do it? (I don't even feel confident in asking AI if a music-recc AI exists because ChatGPT 5 didn't know ChatGPT 5 was out, and Claude still thinks iOS 26 isn't out yet..sigh)
- rldjbpin 1y agothe OP is quite an interesting team to watch regarding open-weights* voice-related efforts. this is a nice read to understand the core of their approach. quite unfortunate, however, their approach to accessibility. unmute [1], which uses the approach discussed in this post, runs quite well with claimed feature of adapting to any voice provided you have a 10 second recording. this is not made available to public at all, despite an issue raised since july. [2] given the pace of the industry, it is a shame that we need to look elsewhere for using an otherwise well-designed tooling. [1] https://news.ycombinator.com/item?id=44109610 https://news.ycombinator.com/item?id=44109610 [2] https://github.com/kyutai-labs/unmute/issues/99 https://github.com/kyutai-labs/unmute/issues/99