3 ms·
if you want to control the routing, you'd lock down providers I don't want to control routing. I want the model to work how the giant, prominent "BENCHMARKS" s
by bbor 25d ago
if you want to control the routing, you'd lock down providers
I don't want to control routing. I want the model to work how the giant, prominent "BENCHMARKS" section says it works, not randomly have a 100x error rate.
If OpenRouter is a marketplace to pick a provider while avoiding huge problems, it is terrible at that job. It surfaces literally none of that info in the top-level list, the graphs below are mislabeled and useless at best, and doesn't notify you of this horrifying situation anywhere, even in passing. There's not even a way to compare providers, AFAICT -- you can only compare models.
This is quite literally the point of OpenRouter.
Their tagline is "better prices, better uptime, no subscriptions". The first two of these directly and inherently contradict your understanding -- neither would be possible if OpenRouter was just a fancy way to change something in their GUI rather than changing the target url of your gateway.
- embedding-shape 25d ago> I don't want to control routing. I want the model to work how the giant, prominent "BENCHMARKS" section says it works, not randomly have a 100x error rate. Why do you care about the public benchmarks at all? The way companies and effective individual developers use OpenRouter, is that you first create your evaluation framework/benchmark, for your specific tasks and use cases, and make that real easy to run and use various models and providers with it. Then you run this to gather data. Then you use said data to figure out what works and what the quality/cost tradeoff you want to make is. Then you lock that down in production while you keep iterating on your benchmark to make it match with real-world use cases and keep adding the new models that pop up. I don't think anyone serious is just willy-nilly making individual requests against OpenRouter and similar platforms, get a "feel for a provider" then use only that provider. Not only would it be wildly inefficient, but also you need hard numbers to compare so you can make informed choices. For this process and workflow, OpenRouter is great, because adding/changing providers and models is essentially changing two strings, rather than having a adapter for each platform you want to try out. If you just want best accuracy requests from SOTA models for your agent you run locally or whatever, then don't use OpenRouter, it doesn't make much sense, but use the provider the model maker has available, as almost all of them run their own endpoints.
- bbor 25d agoIf openrouter is only for people who "make their own benchmark" in a mission to roll something that's usually made by scientists with large budgets, it should say so and thus fade into deserved obscurity. I don't think anyone serious is just willy-nilly making individual requests against OpenRouter Despite your confidence, that is indeed the basis of this massive corporations entire business plan. If you just want best accuracy requests from SOTA models for your agent you run locally or whatever, then don't use OpenRoute If OpenRouter is only for bad accuracy, they should say as much and fade into deserved obscurity.
- embedding-shape 25d ago[flagged]
- infecto 25d ago> Hmm, yeah, good and condense version of what my previous comment said. I'm much impressed by your reading ability. 10month old account with 20k karma. Low value rubbish postings as a professional user. Sad.
- embedding-shape 25d ago[flagged]
- deleted 25d ago[deleted]
- porridgeraisin 25d agoI think you're focusing only on the general coding agent aspect of LLMs. > If openrouter is only for people who "make their own benchmark" in a mission to roll something that's usually made by scientists with large budgets, it should say so and thus fade into deserved obscurity. That is the way LLMs have to be used for highest reliability. While the term stochastic parrot has been co-opted by unreasonable LLM skeptics, that is indeed what LLMs are. You have to have grounded evals that check outcomes if you want to use them reliably - or a human in the loop works too. The more general your family of tasks, the less likely you can make automated evals. So for "general" coding agents, you need a human in the loop that can verify and it's not that easy to write an eval. But if you have specific tasks, then you can spend the time to make a eval, and then you can optimise the way you use the LLM and get extremely good success rates. It's not like it's black magic. Nor does it need large budgets. > that is indeed the basis of this massive corporations entire business plan. No. Individual developers using codex (for extremely underspecified general engineering) needs human in the loop, is not amenable to evals but is only a fraction of all LLM usecases.
- lelandbatey 25d ago> There's not even a way to compare providers, AFAICT That's not quite true. The only thing they don't show per-provider is benchmark data, cause I don't think they are doing continuous benchmarking of each model from each provider, as I assume they feel that's too expensive. You can see hugely detailed breakdowns for near-time metrics per provider for any model by visiting the page for that model on Openrouter. For example see the page for Qwen 3.8 27B: https://openrouter.ai/qwen/qwen3.8-27b https://openrouter.ai/qwen/qwen3.8-27b Some of the killer stats they show per provider: - Pricing: Effective price accounting for cache hit rate, by provider - Performance: Throughput in tok/s, latency, E2E latency, tool call error rate, structured output error rate, and more; all per provider. - Uptime: You have to click on the provider to see their specific uptime, but doing so does show the last-7-days uptime, and you can click to see more.
- Gracana 25d agoThe thing they don't show is the one we really need, especially because model providers can skimp on quality (run lower quantization, lower kv cache precision, etc) to improve their pricing and performance. I agree that it's probably too expensive to keep running the benchmark, but we need some way to hold the providers to a certain standard, otherwise every user has to discover the problems on their own.
- numlocked 25d agoWe are doing continuous benchmarking of each endpoint, for each provider, and it is very expensive :)
- bbor 25d agoNo I know, I was in the GUI as I wrote that lol. As the other person said: if they vary this much in quality, not including that way above updtime and performance is absurd. What would you use a fast, always-up, broken endpoint for?