7 ms·
The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin
- deleted 5mo ago[deleted]
- andai 5mo agoSo basically, Hy3 is the cheapest decent model on OpenRouter, unless you use DeepSeek as the provider for DeepSeek V4 Flash, in which case DeepSeek's insane caching wins out. (And Hy3 is close-ish on the benchmarks.)
- 0xbadcafebee 5mo agoYou need to use DeepSeek API directly to gain the extra caching benefits. The DeepSeek provider on OpenRouter is only the 5th-cheapest for V4 Flash, so you have to specify DeepSeek provider when calling OpenRouter. But DeepSeek's API discounts on its models only applies if you call DeepSeek directly. So anyone using OpenRouter to call DeepSeek models is actually losing quite a bit of money.
- beacon294 5mo agoZDR is also on by default and deepseek is not ZDR.
- NitpickLawyer 5mo ago> The DeepSeek provider on OpenRouter is only the 5th-cheapest for V4 Flash You might have the default settings on your account, which limit Deepseek as a provider. If you disable that feature you see them on openrouter as well (and they serve it at the same cost as their own API).
- 0xbadcafebee 5mo agoI just checked my settings and I have everything enabled. https://openrouter.ai/deepseek/deepseek-v4-flash?sort=price https://openrouter.ai/deepseek/deepseek-v4-flash?sort=price (per-1M price) shows DeepSeek provider as #5. https://openrouter.ai/deepseek/deepseek-v4-flash/pricing?sort=price https://openrouter.ai/deepseek/deepseek-v4-flash/pricing?sor... (effective price) shows them as #3. The effective price will change your total cost since each provider has a different price for input vs output vs cache, so what's #1 and #5 for one person could be #5 and #1 for somebody else, depending on their workload. However, I just double checked, and OpenRouter's pricing page for Flash v4 with DeepSeek provider shows a cache hit rate of $0.0028, which is the same as on DeepSeek's official API pricing page ($0.0028), so they do seem to be the same price, (assuming DeepSeek is able to pin your specific OpenRouter requests to the same DeepSeek server). OpenRouter adds 5% to that cost, but still it might be cheaper than the other providers. Also just found out OpenRouter has a new feature "Response Caching" where they can cache identical requests and return them immediately with no billing. The entire request must be identical, though, not just a prefix, and you have to enable this feature. I don't know who would need to send multiple identical requests, but it's better than nothing?
- NitpickLawyer 5mo agoInteresting, it seems we have some providers offering dsv4-flash cheaper than ds themselves. For the full model it's the other way around, all 3rd party providers are 2x+ more expensive.
- 0xbadcafebee 4mo agoThe cheaper ones are fp4 and fp8 whereas I assume DeepSeek provider is unquantized, so that probably accounts for it. DeepSeek also doesn't necessarily have the cheapest hardware, other providers could be using it as a loss leader, etc
- throwa356262 4mo agoI belive no sane provider, antropic and openai included, serve BF16. Side note: I suspect Antropic was experimenting with changing quant level based on server load a few months back which is what caused that major quality drop we saw then.
- Aurornis 5mo ago> Two new models are now beating LLM darling Claude in terms of token usage and by more than 50%? Time for a reminder that OpenRouter leaderboards only show tokens sent through OpenRouter, which most Anthropic API users don’t use.
- svantana 4mo agoI would think that's true for all the models on OR. The data is skewed for sure, but it's interesting none the less.
- killingtime74 4mo agoAre you next going to say YouTube rankings don't take into account videos that aren't on YouTube and Spotify rankings don't take into account songs that aren't on Spotify?
- 9cb14c1ec0 4mo agoThat doesn't mean it can't be used as a market signal. These 2 things can both be true at once.
- TurdF3rguson 4mo agoI'm pretty sure the popularity came from being free at some point
- smartbit 4mo agoNotice that Hy3 Preview usage didn’t go down after the free period was over https://openrouter.ai/tencent/ https://openrouter.ai/tencent/ The list of apps using Hy3 Preview shows Hermes Agent causing 65% usage over the last 3 weeks https://openrouter.ai/tencent/hy3-preview/apps https://openrouter.ai/tencent/hy3-preview/apps Hermes Agent 72B OpenClaw 10B OpenHands 9B Claude Code 8B Kilo Code 8B
- 0xbadcafebee 5mo ago> it makes sense that a cheaper model would prevail, but only if it offered similar quality You're trying to think logically, which has no place in an AI discussion. :) People just jump to whatever the latest model is. Plenty of people also prefer price to "quality" (which is very subjective). It's new, it's cheap, so people use it. It's likely people will stop using it when something else is cheaper and/or newer.
- olmo23 4mo agoSince my employer pays for it, I just select the latest and greatest.
- vessenes 5mo agoSince there’s only one inference provider it could be a recycling/ad experiment. The similar usage between trial and paid periods would be explained by this as well.
- haeseong 5mo ago[dead]
- simonw 5mo agoFirst model I've tried that gave me back HTML with a "Change Pelican Color" button: https://static.simonwillison.net/static/2026/hy3-preview-pelican.html https://static.simonwillison.net/static/2026/hy3-preview-pel... (Transcript: https://gist.github.com/simonw/c2a0d8ecd3056a2681319eae8fc3f7ce https://gist.github.com/simonw/c2a0d8ecd3056a2681319eae8fc3f...)
- fragmede 5mo agoHaha does it get bonus points for the extra button, or does it fail because html != SVG?
- dodslaser 5mo agoAny bonus points for the color sre immediately subtracted because the "animate wheels" button leaves the wheels stationary and makes the sun rotate.
- MostlyStable 5mo agoI wonder if it is actually animating the wheels as well, but just managed to match up the spin rate to the gap size.
- fragmede 5mo agoHy3 is a Scandinavian model, and is leaking that out via Norse mythology about Sol being a wheel!
- cicko 5mo agoThat depends on the perspective. If you're on the Sun, the wheels rotate around you.
- postepowanieadm 5mo agoROTFL
- Garlef 5mo agoJudging from the dotted trajectory lines, it even "thought" about giving the bike a wobble. (But maybe that's just my interpretation based on something else going wrong in the animation)
- simonw 5mo agoOpenRouter rankings frustrate me, because they show the total number of tokens but they provide no indication of how many unique users a model has. Which means if a surprise model tops the leaderboard one week we can never be sure if it was because a single whale user pushing billions of tokens a day switched to it, or if it represents a genuine community trend towards that model.
- senordevnyc 5mo agoAgreed. My little solo dev SaaS app’s production pipelines push almost two billion tokens a day.
- senordevnyc 4mo agoHaha, I never tire of the AI haters downvoting stuff like this. Down with reality!!
- daveguy 4mo agoOr, everyone finally realizes that token burn is not the same as productivity. Maybe they just down voted for the questionable spending brag.
- dotancohen 4mo agoQuestionable spending aside, GGP is providing information about how a specific metric may not measure what people think it measures. There is value in that comment.
- senordevnyc 4mo agoWe were talking about whether these metrics are meaningful. I was just pointing out that even a tiny one-person company can burn a lot of tokens. As to whether the token spend is questionable, the number I quoted is for my production AI pipelines, not for coding. And my customers (and profit margin) seem to think the spending is valuable.
- 5mo ago
- bandrami 5mo agoFor the life of me I will never understand the thought process that leads you to say "we don't really know who developed this LLM but I'm going to feed all of my business's data to it"
- ddalex 5mo agowhat can it do ? it's just a big set of numbers, if you trust the host that's good enough
- what266262 5mo agoIf you are ok with everything being fed into it being stored forever I guess it’s no problem. I don’t see how you trust them if you don’t know them.
- Dylan16807 5mo agoWho is "them" here? The developers and the hosts are not the same.
- bandrami 5mo ago(And either one is a threat vector)
- ddalex 4mo agowhere would it be stored ? it's just a big set of numbers.
- Mashimo 5mo agoIf you Code open source projects anyway, might give it a spin.
- est 5mo ago> I'm going to feed all of my business's data to it Your business data is probably worthless, even considered harmful for the pretrain corpus. Your interactions and decision making process are most valuable parts of the whole business.
- freakynit 5mo agoThis was originally a 400+B param model which was later reduced to 295B considering it as the "optimal zone". https://www.mdshare.online/s/uend0pj3og_A_rgcxzINf https://www.mdshare.online/s/uend0pj3og_A_rgcxzINf
- zone411 5mo agoI’ve tested this model on four of my benchmarks: https://github.com/lechmazur/buyout_game https://github.com/lechmazur/buyout_game 10th out 36. https://github.com/lechmazur/pact/ https://github.com/lechmazur/pact/ 14th out 25. https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/ 60th out 81. https://github.com/lechmazur/debate https://github.com/lechmazur/debate 16th out of 29.
- CamperBob2 4mo agoWould be interesting to see the 27B dense Qwen 3.6 model thrown into the mix.
- baxtr 4mo agoGood stuff! Is there a reason you change the leaderboard graphs for the third and fourth one? Also: would be great to have an overview page with a summary over all test, like a total score or similar.
- Sandworm5639 4mo agooh, I love the connections benchmark. Just curious, can you share what are those hardest puzzles that even the top models can't crack? sometimes when I find the puzzle absolutely undecipherable I like to ask LLMs to solve it, and I haven't seen them fail yet.
- thot_experiment 5mo agoTried this extensively in OpenCode, never used it once since Gemma 4 came out, got into thought loops and did stupid edits I didn't ask for more often than the local 31b model. One of the worst "frontier" models I've ever tried.
- BoorishBears 4mo agoThis article got me messing with it, and I'm loving it as a post-training target. Training on ~1B tokens on 8xB300 and the first checkpoint halfway in learned really well. Tencent might be struggling with agentic work, but the base knowledge is there.
- cicko 5mo agoHow is it a "mysterious" model? It's Tencent's Hy3?
- theanonymousone 5mo agoMy question as well. Isn't Tencent a very well-known company? Maybe the mystery is in the model itself?
- alecco 5mo agoPSA: Don't use OpenRouter for DeepSeek V4 as it messes up you caching. Use DeepSeek API directly and you'll get 2x to 3x more cached tokens.
- numlocked 4mo agoCan you share more? I'm with OpenRouter and we would love to address this! We don't see this in our own testing, I don't believe -- but will share this feedback and dig in.
- bwfan123 4mo agoHere is some data from my experience using both deepseek v4 flash directly, and deepseek v4 flash via openrouter. Directly: 135M input tokens - $0.57 (134M cached) Via OpenRouter 6M tokens - $0.81 (caching stats & inp/out not reported) Caching is a huge win with using deepseek directly.
- alecco 4mo agoJust try. In a case last week it was ~3x and I tried multiple providers: deepseek, gmicloud/fp8, novita/fp8, and another one I can't remember. It was a large job where at least 2/3rds of the start of the prompts was exactly the same (literally a static string). Then I read somewhere (I think X) that OpenRouter adds stuff and breaks caching (telemetry? headers? can't remember). So I stopped the job, switched to actual DeepSeek provider, and voilá, caching 3x more tokens per request (on average).
- alecco 4mo ago> switched to actual DeepSeek provider I meant actual DeepSeek API.
- phainopepla2 4mo agoI am experiencing this using Opencode. Caching works fine via Deepseek API but not so good via Openrouter
- SV_BubbleTime 4mo ago
- lithiumii 5mo agoWhat's so mysterious? Isn't it from Tencent?
- gmerc 4mo agoVery mysterious: https://huggingface.co/tencent/Hy3-preview https://huggingface.co/tencent/Hy3-preview
- segmondy 4mo agoHigh token usage cuz it's free doesn't count
- minimaxir 4mo agoThe post goes into that issue. Throughly. The numbers at the beginning of the post are weekly aggregate values well after the endpoint was paid-only.
- segmondy 4mo agoThe post is wrong, it's still free, see - https://openrouter.ai/tencent/hy3-preview:free https://openrouter.ai/tencent/hy3-preview:free it's free in kilo.ai https://kilo.ai/models/tencent-hy3-preview-free https://kilo.ai/models/tencent-hy3-preview-free It's free in a lot of places.
- minimaxir 4mo agoThe first endpoint was closed. If you actually try and call it from the API you get this response: > Hy3 preview is no longer available as a free model. It has transitioned to a paid model. Continue using it here: https://openrouter.ai/tencent/hy3-preview https://openrouter.ai/tencent/hy3-preview The Kilo Code may have free traffic but if you check the numbers is still inconsequential relative to the trillions of tokens through OpenRouter.
- ravirdp 4mo ago[flagged]
- sheepscreek 4mo agoFYI - DeepSeek has NOT announced its own coding platform. That app is an independent project. It says so in the footer as well: “Independent open-source project · not affiliated with DeepSeek”