3 ms·
The point of speed is to increase throughput. What the point of all this speed, if overall throughput is still so low? This doesn't work for my use case at all
by eveningtree 1mo ago
The point of speed is to increase throughput. What the point of all this speed, if overall throughput is still so low?
This doesn't work for my use case at all (code generation).
These bursts of speed might work well for workflows that need bursts of quick decisions, followed by silence. But these workflows have needed provable determinism to som extent, so I haven't been using llms for those use cases. And I don't see myself using llms for them in the future too.
- brookst 1mo agoReminiscent of race to sleep: not suitable for sustained workloads, but for bursts ones it’s a good approach.
- wongarsu 1mo agoMight be usable for short-context utility workloads? Generate the title of your chat session based on the first three messages at the speed of light
- 70rd 1mo agoThroughput is useful if you want to generate a lot of transcripts for RL. It's for making Qwen better, not for actually using Qwen.