4 ms·
Did you miss the part of this article where the benchmarks specifically call out Fireworks as one of the worst in their tests: > Fireworks scored 46% on TAU, a
by nacs 21d ago
Did you miss the part of this article where the benchmarks specifically call out Fireworks as one of the worst in their tests:
> Fireworks scored 46% on TAU, a 30 point gap
Another surprise was DigitalOcean being bottom of barrel too.
Companies are apparently willing to risk their brand name by being deceptive about these heavily quantized/flawed model-serving.
- svachalek 21d agoDigitalOcean (at least through OpenRouter) is pretty reliably bad in my experience. Fireworks can be great, but it depends on the model and on the day.
- tandr 20d agoDigitalOcean is bad not just through OR... Direct accesss through DO gave me such a miniscule context and maximum tokens, it was pretty much unusable. So, accessing it through OpenRouter gave bigger context length, but model behaves like it was seriously damaged - tool calling was producing paths that had / replaced with some other symbols, files were not found etc., and this is just for things that errored out, I have no idea how bad reasoning was. I had to blacklist DO completely on OpenRouter.
- FergusArgyll 20d agoIt's not always deceptive. Bugs in inference can be very subtle (well, bugs everywhere can be subtle). The providers should be benchmarking their offerings daily