3 ms·
Yeah, this is pretty accurate. Some providers are basically scams too. I encountered one provider for GLM-5.2 which ran at 200 tps (absurdly high), and was so b
by wren6991 23d ago
Yeah, this is pretty accurate. Some providers are basically scams too. I encountered one provider for GLM-5.2 which ran at 200 tps (absurdly high), and was so broken that it would issue 20 full reads of the same file in one turn and quickly jump up to 300~400k context and charge me for the prefill. Went straight on my deny-list, but they got a few dollars out of me first.
Another common annoyance is having a request go to a provider that dribbles out ~1 tps (even for small models like DeepSeek V4 Flash). If you cancel the request, you still get charged for the prefill and the handful of generated tokens. If you don't cancel the request, you might be waiting 10 minutes for the turn to finish.
The overall experience is pretty good, and it's the best way to try new models, but they don't appear to do any real vetting or apply any quality standards to their providers, and occasionally it bites you.