5 ms·
This shows the real utility of the Groq low latency inference. That delay in responding makes it much less impressive (it’s still really impressive obviously!)
by eigenvalue 3y ago
This shows the real utility of the Groq low latency inference. That delay in responding makes it much less impressive (it’s still really impressive obviously!)
- modeless 3y agoThat delay will be eliminated very soon. IMO low latency natural voice conversations are going to be bigger than ChatGPT. It's going to blow people's minds when they can converse with these AIs casually just like with their real life friends. It won't be anything like Siri or Alexa anymore. Here's a demo from a startup in this space. Still very early. https://deck.sindarin.tech/ https://deck.sindarin.tech/
- tomcam 3y agoI’ll be thrilled when Siri can spell my wife’s name correctly after 13 years of continuous usage and explicitly training it in her name. Admittedly her name is wildly complicated and totally unknown to the software folks at Apple: Ada Also my radio-trained voice is so generic a caller every week-ish assumes I am a bot, so I’m pretty sure the problem isn’t me enunciation or accent.
- margorczynski 3y agoLLMs can't really spell as the smallest "block" of information they operate on are tokenized whole words. Same issue with e.g. arithmetic.
- ta8645 3y agoBut they should be able to select the correct token for homophones, which amounts to the same thing.
- jxy 3y agoCan we stop this kind of misinformation? Training a model to map a token to individual letters are no harder than training a model to be fluent at English. Arithmetics with small number of digits are achievable as well. You can just try a small 7B model yourself. If you don't know where to start, try the mistral instruct v0.2, and this is how it goes, > [INST] Spell out the following word letter by letter: margorczynski [/INST] m - a - r - g - o - r - c - z - y - n - s - k - i > So, the word "margorczynski" spelled out letter by letter is: m-a-r-g-o-r-c-z-y-n-s-k-i. The text between `[INST]` and `[/INST]` is the input. The text after `[/INST]` is the output.
- margorczynski 3y agoIs Karpathy lying saying that word tokenization brings such problems that can be seen in many LLMs? https://twitter.com/karpathy/status/1657949234535211009 https://twitter.com/karpathy/status/1657949234535211009 I'm not arguing that you can't use single chars just that many of the issues parent discussed are caused by this.
- imtringued 3y agoThe easy solution is to create an additional dataset that is token aware. I.e. you take 1% of the dataset and take random tokens and split them into smaller tokens while expecting the same answer on the character level. This should force the model to learn multiple token representations of the same character strings.
- tintor 3y agoLetter-by-letter tokenization increases inference and training costs and latency (as you need more tokens)
- ben_w 3y ago> Admittedly her name is wildly complicated and totally unknown to the software folks at Apple: Ada Aye. I was surprised this morning when it decided I had was talking about a "Mark of chain". 1/3rd of the time it hears "bedroom 100%" as "bedroom off". When cooking dinner today, I asked for a "ten minute timer", it responded "for how long?" then confirmed my "ten minute minute timer". Still better than Alexa, which kept telling us it couldn't find «kitchen» on Spotify even though we didn't even have Spotify. And way better than the voice control on Mac OS Classic; back in the late 90s/early 00s, it interpreted 75% of my attempts to use it as "tell me a joke" (it wasn't even a good joke), and ignored 20%.
- educaysean 3y agoI must say the demo did nothing to improve my opinion of the current state of voice-based AI conversations.
- batwood011 3y agoHey there — I’m Brian, the founder of Sindarin, the company behind this pitch deck. This demo is pretty bad compared to what we currently have in development. We’ve been in code freeze in prod for over two months to get our substantially improved engine finished. It’ll be out in a few weeks, and it’ll blow this version away in every way that matters. Thanks for checking us out!
- penjelly 3y agodemod this. The ability to interrupt the language model is very cool, However, i notice, it failed to move onto the next slide often. It could never get to the final slide without explicit mention to go there, and when i got to the last slide, i asked to go back to the first slide, it would say "ok lets go to the last slide" everytime, these are probably more control issues than language model issues but i thought id point them out, just in case.
- amorriscode 3y agoTotally agree on this take but the founder posted above you saying they have greatly improved their engine. Should be interesting to see!
- joshstrange 3y agoHeck “talking” to ChatGPT is already pretty freaking impressive but even faster responses (and handling interrupts) would make it even better.
- leetharris 3y agoAbsolutely. People on X keep making the mistake of assuming cloud / network latency is the problem here. The vast majority of America is within 10ms of a data center. That's nothing. The current challenge for most interaction is ASR -> prompt processing latency. This will be improved with multimodal models on specialized hardware like Groq.
- tomp 3y agoIME I can get about 0.2s to get the first chunk from Mistral (i.e. Mistral API, using Mixtral model (`mistral-small`), not Mixtral on Groq) (and note the Mistral sends larger chunks, unlike ChatGPT which sends individual tokens) and another 0.6s or so to get first voice chunks from PlayHT measuring STT latency is harder, I need to implement a local VAD model first to properly measure it, but I think it's on the order of 0.5s So this has nothing to do with Groq, really. ChatGPT is just slow (too slow for realtime voice communication).
- hobofan 3y agoUnless the only thing you want to do with the robot is talk, you need to do a lot more reasoning and execution planning first (= multiple LLM round trips; tool calling) before you even know whether talking is the correct action to take. So the naive time-to-first-chunk estimate will be way off.
- andoando 3y agojust add a hmmm before every response
- fragmede 3y agoWhich we humans do all the time. Okay, like, so, hear me out, alright? See, what’s really going on, yeah, is…
- fennecbutt 3y agoWhich is so cool because it's an evolutionary/language thing. Why do we add junk words while we think? I think it's probably because we're social animals, we want to hold that person's attention as we think as periods of silence are likely to make them become disengaged. But who knows really.
- gfodor 3y agoThe bottleneck here is the multimodal vision processing, at least if my experience building this kind of thing is any indication. Afaik Groq has not demonstrated the speeds they have for this. (Obviously they'll be better than OpenAI, but it still may be slow enough to leave people disappointed.)