3 ms·
I was wondering whether this was any good for programming, but it is too fast for its own good. There is a limit of 450,000 tokens per minute. I hit this limit
by gpugreg 1mo ago
I was wondering whether this was any good for programming, but it is too fast for its own good. There is a limit of 450,000 tokens per minute. I hit this limit in about 90 seconds and burned through $1.10 while doing so. This is because cached tokens count towards the token limit.
For comparison, I ran the same task with DeepSeek-V4-Flash, which finished in 172 seconds and cost $0.024 with a final context window size of 55217 tokens, while Qwen3.8-27B was not even close to being done with a 64178 context window.
This is a very efficient way to burn your money, but I would not recommend it for programming.
On the positive side, I got a $5 signup bonus, so it wasn't my own money.
- d2p 1mo ago> There is a limit of 450,000 tokens per minute. I hit this limit in about 90 seconds I'm confused. If it's 1500t/s, isn't that only 90k per minute? How do you hit a 450k/minute limit?
- gpugreg 1mo agoCached tokens count towards the limit as well. For example, if your context window is 50,000 tokens, it takes 9 requests to reach that limit without generating a single token.
- perching_aix 1mo agothen it's basically useless lol, wtf, this has to be a defect
- selcuka 1mo agoIt's PR: https://news.ycombinator.com/item?id=49556302 https://news.ycombinator.com/item?id=49556302
- nullbio 1mo agoCached tokens counting toward the limit is ridiculous.
- Pxtl 1mo agoCould this also be coming from the problem that Qwen3.8-27B's default mode being "extra-high reasoning level"?
- irthomasthomas 1mo agoWithout prompt caching this becomes more expensive than fable 5.1 after turn 50, assuming you start with 40k tokens and add 2k per turn.
- eveningtree 1mo agoThe point of speed is to increase throughput. What the point of all this speed, if overall throughput is still so low? This doesn't work for my use case at all (code generation). These bursts of speed might work well for workflows that need bursts of quick decisions, followed by silence. But these workflows have needed provable determinism to som extent, so I haven't been using llms for those use cases. And I don't see myself using llms for them in the future too.
- brookst 29d agoReminiscent of race to sleep: not suitable for sustained workloads, but for bursts ones it’s a good approach.
- wongarsu 29d agoMight be usable for short-context utility workloads? Generate the title of your chat session based on the first three messages at the speed of light
- 70rd 29d agoThroughput is useful if you want to generate a lot of transcripts for RL. It's for making Qwen better, not for actually using Qwen.
- jurgenburgen 29d agoI don’t think I understand. Why would faster token generation burn more tokens? The LLM should not be generating anything in between tool calls so the only difference should be that the human waits less between turns.
- codygman 29d agoQwen 3.8 on xhigh defaulr needs 128k context minimum or you'll spend most of your time compacting context. Also make sure you use the instruct temperatures/etc for implementation.
- gabri200 29d agoTo get a good coding agentic system you need to use big context (Specs and conversation context can't be condensed every minute), so you need to use prefix caching, and the price for the hit cache tokens can't be the same that miss cache or the final price could be insane.