5 ms·
Llama 3.1 Omni Model
- nickthegreek 2y agoThe speed looks very nice. I just recently setup LMStudio + AnythingLLM to try out local voice chat and its still a little slower than I'd like but the PiperTTS voices are nicer than this.
- opdahl 2y agoAny demos showcasing it’s performance?
- potatoman22 2y agoThere's one on Huggingface https://huggingface.co/ICTNLP/Llama-3.1-8B-Omni https://huggingface.co/ICTNLP/Llama-3.1-8B-Omni
- opdahl 2y agoThank you. Obviously it doesn’t sound human but that’s extremely impressive for an 8B model. Compared to the Moshi model also on the front page now, this model seems to be more coherent, but maybe less conversational?
- twobitshifter 2y agoThere is a demo video on the page
- dingdingdang 2y agoDoes any of the model-runners support this? Ollama, LM Studio, llama.cpp?
- LorenDB 2y agoThe TTS voice in the demo clip sounds remarkably like Ellen McLain (Valve voice actor). https://en.m.wikipedia.org/wiki/Ellen_McLain https://en.m.wikipedia.org/wiki/Ellen_McLain
- spencerchubb 2y agoSounds like it's trained on LJ Speech dataset, which is one of the best datasets and very commonly used
- londons_explore 2y agoCan this play sounds that can't be represented in text? Ie. "make the noise a chicken makes"
- hansenliang 2y agoasking the real questions
- indigodaddy 2y agoVery interesting question actually
- deleted 2y ago[deleted]
- 8note 2y agoAs in, a cluck? But can it both say the word cluck, and make a clicking sound?
- evilduck 2y agoIf it can create sounds associated with any nonphonetic word spellings, I can't see why it would struggle with onomatopoeia.
- oezi 2y agoCan it understand those sounds? And distinguish correct and incorrect pronunciation of words and accents?
- DrSiemer 2y agoAlmost certainly not. Sounds like an old school vocoder, made to produce human speech and nothing else.
- OJFord 2y agoBwaaaaaak bwakbwakbwak
- twoodfin 2y agoI’m not clear on the virtues or potential of a model like this over a pure text model using STT/TTS to achieve similar results. Is the idea that as these models grow in sophistication they can properly interpret (or produce) inflection, cadence, emotion that’s lost in TTS?
- fragmede 2y agoReally Yeah that's the point. Without punctuation, no one can tell what inflection my "really" above should have, but even if it'd been "Really?" or "Really!", there's still room for interpretation. With a bet on voice interfaces needing a Google moment (wherein, prior to Google, search was crap) to truely become successful (by interpreting and creating inflection, cadence, emotion, as you mentioned), creating such a model makes a lot of sense.
- Reubend 2y agoEssentially, there's data loss from audio -> text. Sometimes that loss is unimportant, but sometimes it meaningfully improves output quality. However, there are some other potential fringe benefits here: improving the latency of replies, improving speaker diarization, and reacting to pauses better for conversations.
- bubaumba 2y ago> I’m not clear on the virtues or potential of a model like this over a pure text model you can't put pure text with keyboard on a robot. it will become a wheeled computer. actually this is a cool thing as a companion / assistant.
- a2128 2y agoThere's a lot of data loss and guessing with STT/TTS. An STT model might misrecognize a word, but an audio LLM may understand the true word because of the broad context. A TTS model needs to guess the inflection and it can get it completely wrong, but an audio LLM could understand how to talk naturally and with what tone (e.g. use a higher tone if it's interjecting) Speaking of interjection, an STT/TTS system will never interject because it relies on VAD and heuristics to guess when to start talking or when to stop, and generally the rule is to only talk after the user stopped talking. An audio LLM could learn how to conversate naturally, avoid taking up too much conversation time or even talk with a group of people. An audio LLM could also produce music or sounds or tell you what the song is when you hum it. There's a lot of new possibility I say "could learn" for most of this because it requires good training data, but from my understanding most of these are currently just trained with normal text datasets synthetically turned into voice with TTS, so they are effectively no better than a normal STT/TTS system; it's a good way to prove an architecture but it doesn't demonstrate the full capabilities
- cuuupid 2y agoWish there was training or finetuning code, as finetuning voices seems like a key requirement for any commercial use.
- aussieguy1234 2y agoNot bad for 3 days training, voice output quality needs some work, it will be interesting to see what effect more training will have.
- drcongo 2y agoAm I the only one who trusts a GitHub repo much less when it has one of those stupid star history graphs on the readme?
- cvzakharchenko 2y agoSo it's not STT -> LLM -> TTS? If I scream Chewbacca noises as input, will the model recognize it as nonsense, or will it interpret it with some lousy STT as some random words?
- a2128 2y agoIt's not, but it probably won't recognize it as nonsense. According to the paper, > we construct a dataset named InstructS2S-200K by rewriting existing text instruction data and performing speech synthesis It has only been trained on questions spoken by TTS, it has never seen (heard) nonsense. Most likely it'll just hallucinate that you asked some question and it'll generate some answer instead of asking if you're good. There's just not many audio datasets with real voices, there's no audio version of StackOverflow to be scraped
- schrodinger 2y agoWhat about every movie that's been made? Although it might need to stick to those more than 100 yrs old to avoid copyright law?
- vintermann 2y agoI used to have fun with that. Set Google Translate to Chinese (Or some other language I don't speak, though tonal languages seemed to work better), make some vague noises into it, and get out coherent but crazy phrases in English.
- deleted 2y ago[deleted]
- barrenko 2y agoI assume one can't finetune this further?