4 ms·
I have lots of questions about how important latency is since you may be replacing many minutes or hours of a person’s time with undoubtedly a quicker response
by gpapilion 2y ago
I have lots of questions about how important latency is since you may be replacing many minutes or hours of a person’s time with undoubtedly a quicker response by any measure. This seems like a knee jerk reaction assuming latency is as important as it’s been with advertising.
I’m not convinced latency matters as much as groqs material tries to claim it does.
- frozenport 2y agoI guess its tool calling? When you chain the LLMs together?
- w-ll 2y agoWhen has latency ever not mattered? Let alone 'chat' use cases, but holding a reponse up for N*1.2 longer than it could holds all sorts of other resources up/down stream.
- ben_w 2y agoWhen it's already faster than I can absorb the response, which for me as an organic brain includes the normal token generation rate of the free tier of ChatGPT. If I was using them to process far more text, e.g. summarise long documents, or if I was using it as an inline editing assistant, then I'd care more about the speed.
- qeternity 2y ago> When it's already faster than I can absorb the response Streaming a response from a chatbot is only one use-case of LLMs. I would argue the most interesting applications do not fall into this category.
- ben_w 2y agoNumber of different use cases (categories) I'd agree; I'm not so sure about use (volume)… …not yet anyway. Fast moving area, lots of blue water outside the chat interface.
- boroboro4 2y agoName one use case where there is a difference between latency of 200 t/s (fireworks.ai mixtral model) and 500 t/s (groq mixtral)? Not throughput and not time to first token, but latency. Groq model shines at latency, not at the other two.
- michaelt 2y agoDepends on your application. For example, if you're a game company and you want to use LLMs so your players can converse with nonplayer characters in natural language, replacing a multiple-choice conversation tree - you'd want that to be low latency, and you'd want it to be cheap.
- beepbooptheory 2y agoBut are people really going to do this? The cost here seems prohibitive unless you're doing a subscription type game (and even then I'm not sure). And the kinds of games that benefit from open ended dialogue attract players who just want to pay an upfront cost and have an adventure. (All the sudden having nightmares of getting billed for the conversations I have in the single player game I happen to be enjoying...) If there is a future with this idea, its gotta be just shipping the LLM with game right?
- Const-me 2y ago> If there is a future with this idea, its gotta be just shipping the LLM with game right? That might be a nice application for this library of mine: https://github.com/Const-me/Cgml/ https://github.com/Const-me/Cgml/ That’s an open source Mistral ML model implementation which runs on GPUs (all of them, not just nVidia), takes 4.5GB on disk, uses under 6GB of VRAM, and optimized for interactive single-user use case. Probably fast enough for that application. You wouldn’t want in-game dialogues with the original model though. Game developers would need to finetune, retrain and/or do something else with these weights and/or my implementation.
- michaelt 2y agoI understand there are games using LLMs for NPC dialog, yes [1] > If there is a future with this idea, its gotta be just shipping the LLM with game right? Depends how high you can let your GPU requirements get :) [1] https://www.youtube.com/watch?v=Kw51fkRiKZU https://www.youtube.com/watch?v=Kw51fkRiKZU
- beepbooptheory 2y ago
- verdverm 2y agoGoogle won search in large part because of their latency. I stopped using local models because of latency. I switched from OpenAI to VertexAI because of latency (and availability)