8 ms·
Hy4 preview
- zem 1mo agoI was briefly impressed that https://hylang.org/ https://hylang.org/ had released a 4.0 version!
- snthpy 1mo agoMe too!
- minimaxir 1mo agoHy4 apparently has ludicrous traction on OpenRouter already (https://openrouter.ai/tencent/hy4-preview https://openrouter.ai/tencent/hy4-preview), with trillions of tokens processed in a couple days: more than GLM 5.3 in a week. That said, it's relatively cheap with a 5% cache cost when everyone is still doing 10%/20% cache costs, so Hy4 may be more compelling.
- cyanydeez 1mo agoi'd be curious if openrouter is just being gamed by these publishers by paying for the exposure. wouldn't trust they dont do Capitalism like the rest of the AI field.
- tokai 1mo ago>dont do Capitalism like the rest of the AI field Like lobbying the US president to harm their competitors?
- realo 1mo agoI would suggest "lobbying" is not the correct word to describe all the corruption going on in the current USA administration cesspool.
- noir_lord 1mo agolobbying/legalised bribery hard to say where one ends and another begins at times.
- CamperBob2 1mo agoWell, it sure as hell isn't capitalism.
- blackqueeriroh 1mo agoLmao that’s exactly capitalism
- andrekandre 1mo agoi mean, theres capitalism as the ideal, and there is capitalism in practice, so maybe you are both right...
- CamperBob2 1mo agoWhere in the Wealth of Nations does a Trump appear?
- realo 1mo agoOh. Administration corrupted up to it's very core? Check. Nihilism of anyone not part of the proper color, gender, whatever agenda? Check. Unlawful surveillance? Check. Sending totally innocent citizens to prison with many of them dying mysteriously? Check. Killing innocent people in the streets simply because they dare protest peacefully? Check. Welcome to North Korea! Oups. Confused. Welcome to the GREAT US of A! Where True Capitalism is practiced.
- CamperBob2 1mo agoYou need to go back to school and demand a refund if you think any of that is "true capitalism." Better yet, just go back to Reddit.
- drob518 1mo agoOf course they are. Of course they do. Nobody should be surprised by this.
- npn 1mo ago[flagged]
- martinald 1mo agoI wrote about this a couple of weeks ago. It's actually often the biggest cost and it tends to be hidden away on most platforms! https://martinalderson.com/posts/watch-out-for-cache-read-costs/ https://martinalderson.com/posts/watch-out-for-cache-read-co... Btw I still haven't came across any decent model that is <$0.01/MTok cache costs apart from deepseek thru their official API (even with the price increases). Seems like a bit of an opportunity for someone to take - drop cache read costs significantly.
- dakolli 1mo agoThat's because Deepseek invented the paradigm of prompt caching, they are the SOTA when it comes these techniques. Despite them open sourcing all their research, nobody beats them. edit: I do wish openrouter would let you sort providers by Cache Hit % and Cache cost. These are the only things that matter to me at this point when choosing a provider.
- minimaxir 1mo agoYou can click the table headers to sort Ascending/Descending.
- dakolli 1mo agoYou can sort by the cost cache cose, but you cannot by cache hit %. You have to click on the provider and see what their cache hit % is. A provider could have a super low cache cost, but a 50% cache hit percentage, making the cheap price of cache read's meaningless.
- andai 1mo agoWait, what does that number mean? I thought it always uses the cache price when the prefix matches.
- dakolli 1mo agoWhen the prefix matches a request sent to the same Providor. The thing is the TTL is different for each provider, some cache for 5 minutes some cache for 1hr. Its ideal to only use one provider per agent session / and per model with the best cache hit % if you care about costs.
- Dinux 1mo agoWhich explains why almost none of my request go though
- redox99 1mo agoIt's very likely tencent games those stats, buying their own tokens.
- eli 1mo agoOpenrouter tracks what apps are using the model and the top ones for hy4 are all different coding harnesses. I guess it could be fake but seems more likely people are just trying it out. Hy3 was a very strong and underrated model.
- redox99 1mo agoThe speed at which hy4 usage increased on openrouter, especially considering its not a cheap model, doesn't seem organic to me. Its already serving as much tokens/day as the incredibly cheap and good GLM 5.3 flash, which had a crazy marketing campaign as ox alpha? Also those top 5 apps are just 1.58B tokens out of 1.54T tokens from yesterday. Negligible.
- joegibbs 1mo agoIf you’re Tencent you can just plug it into some field somewhere that lots of people see right? Like how Meta could put their model on Instagram search
- bobby_coder_55 1mo agoUnfortunately codebuddy login is not working for me in the United States of America
- vcryan 1mo agoI used Hy3 quite a bit for the type of tasks it was suited for. Excited about this. My one concern over Hy3 was speed. In theory, it could be served much faster as a smaller model but it was relatively slow everywhere I could get it (including from Tencent directly) but also several other inference providers.
- Topfi 1mo agoIn my evals, I saw an unprecedented jump between preview and final release on Hy3, from unusable to competitive. Did you see similar in preview vs release version?
- vcryan 1mo agoOh yes! I forgot about that. Yes, you can see this in benchmarks about hy3 preview and hy3 release still today because they measured them separately - it was significant.
- usernomdeguerre 1mo agois it just me or are the bar charts in the blog post strange? Higher numbers don't seem to correspond correctly to their actual height?
- feynmanquest 1mo agoNoticed that as well
- alanfranz 1mo agoProbably AI generated. But, what bars are clearly off? I couldn't spot any.
- pixelesque 1mo agoLooks okay to me. The first column has both the Hy4 and Hy3 scores overlaid on one another (Hy4 is darker blue and the taller one), with both scores written below the top of the respective bar - maybe you're seeing that?
- onesandofgrain 1mo ago[flagged]
- yogthos 1mo agoSeriously, without China we'd just have two parasitic companies hoarding this tech and deciding whom and how is allowed to use it.
- Kuyawa 1mo agoClaude and OpenAi are not allowed in Venezuela, so I thank China too and I swear to god I'll never use them and will be rooting for chinese models forever
- aforwardslash 1mo agoIm not particularly fan of the chinese, but no chinese model asked for my citizen card yet to complete a task. And apparently no chinese provider uses persona to manage this kyc information. OpenAI does, in EU space. Just saying.
- unethical_ban 1mo agoVague and without substance. It easily passes as sarcasm, which means criticism but without any commentary, else it is sincere... but doesn't have any commentary.
- jorl17 1mo agoI experimented with Hy3 for a project and was surprised with how good it was. I don't know if it's good for coding, but as a general purpose agentic model, it was only beaten by deepseek4-flash in our tests. It was so close to deepseek behaviour I kept thinking it must have been forked from it.
- alexfortin 1mo agoFor the last few days I've been experimenting with the _free_ version of Hy3 offered by Opencode Go and I was also surprised to see how (relatively) good it is on coding tasks too. The free quota from Opencode Go is also surprisingly generous, I perhaps hit limits one or two times and I've been using it _a lot_ for implementation tasks (using e.g. GLM-5.3-flash for working on specs and planning next steps).
- deleted 1mo ago[deleted]
- Zigurd 1mo agoIs anyone here working on a problem for which current generation LLMs are inadequate, but that could possibly be solved by the next release of a first tier LLM? Or is it like bicycles? Unless your problem is named Tadej, you don't need a $13,000 bike.
- poincareball 1mo ago[dead]
- tokai 1mo agoA spanish rock solved that problem for free.
- nozzlegear 1mo ago¿Como?
- Rexxar 1mo agoHe just had an accident on "Vuelta a España" : https://www.theguardian.com/sport/2026/aug/29/cycling-tadej-pogacar-pulls-out-vuelta-espana-crash https://www.theguardian.com/sport/2026/aug/29/cycling-tadej-...
- nozzlegear 1mo agoOh I had not heard about this, thanks for the link.
- wiether 1mo agoReading OP's analogy I was like "even him don't need this bike now..." I feel bad for him as a human, but as a cycling fan I'm glad that we'll have an interesting WC in Canada
- RGS1811 1mo agoFor me personally, the answer is no. Fable is adequate to do basically anything I want to do. My perspective, broadly speaking, is that we've saturated most of the benchmarks because we've largely saturated our capacity to verify models' work at scale. What's left is context-bound verification, i.e. the problem of ensuring that output matches intent and ambiguities in prompting were resolved correctly. Further advances in autonomy do not make that latter verification problem easier. If anything they make it harder as the output per task becomes more complex and therefore more taxing for a human to verify. The solution to that (to my mind) would be not a better model but a basic shift in architecture beyond the current paradigm and into a setup where agents have durable, plastic memories and undergo contextual individuation over time. But at that point agents start to become quasi-persons and not tools.
- fastball 1mo agoI wish model providers would stop committing chart crimes in their releases. - if you're gonna order the rest of the bar chart by rank, order your model accordingly. - if you're gonna highlight a winner in a table of benchmarks, don't highlight your entire model row in the table. Etc etc
- mirekrusin 1mo agoRead websites through llm.
- deleted 1mo ago[deleted]
- jimbob45 1mo agoI wonder if this is being reinforced via LLM because they see every other modeler doing the same thing.
- nullbio 1mo agoI'd bet there's a correlation between benchmaxxing and chart crimes. Companies who try to deceive perceptions via the charts are more likely to cheat at the benchmarks too, I'm sure. That's assuming ill intent, of course - which is often the case for charts related to model releases, but not necessarily always the case.
- petcat 1mo ago> Tencent has released and open-sourced Tencent Hy4 preview, a next-generation large language model with 770B total parameters and 49B active parameters, and a context window exceeding 1M tokens. There are no open source models, at least not useful ones (yet) [0]. Open weight is not the same as open source. The current "open weight" models are just opaque binary blobs you can run on your own computer instead of through a web API. [0] https://allenai.org/ https://allenai.org/ Imagine thinking that running a Photoshop binary on your own computer instead of through a SaaS web app means that it's "open source". Of course you think that's ridiculous.
- mirekrusin 1mo agoYou can open source dataset without all the details how it was assembled. Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under. Hy4 is released under OSI approved Apache License 2.0.
- kennywinker 1mo agoParent poster is technically right - open “source” implies the source used to make something is open. The model source is training data and code, not just weights. But the reality is, the weights are a useful artifact that you can use to create derivative works. So, dismissing it as a photoshop binary is as technically wrong as calling it open source.
- Alpha3031 1mo agoIIRC Nvidia claims to release enough data that it should be possible to fully reproduce Nemotron, so even if it's not as good as the current best models, GPT 5.1 or Opus 4.1.was still useful right? I guess it depends on what you wanted to do with them.
- LtWorf 1mo agoSo windows is open source because the binaries are a lossy compression of the original source?
- XCSme 1mo agoI tried benchmarking it, but it keeps timing out/rate limiting, so the current provider(s) are unusable.
- simonw 1mo ago> [...] Let's maybe add a helmet? It could improve riding theme, but may obscure head. Maybe a small cycling cap or helmet? The user didn't ask; can add red helmet? Might be cute. But pelican with big beak; a helmet might obscure. Better maybe no. > Maybe add sunglasses? no. > Maybe add water? no. https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fcb69816b3fb940f2782569a82a523af1 https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- Aboutplants 1mo agoHalfway through it states “Let's mentally compose SVG.” Is this a common thing? I’ve never seen it before, the “mentally compose”
- simonw 1mo agoIt's common for models to produce a draft of the SVG part way through their reasoning. Here's Qwen3.8-Flash-Next doing that, for example: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6ba7cbfc1a9336986703b41f7fccd73a https://tools.simonwillison.net/markdown-svg-renderer#url=ht... The open weight models let you see the full reasoning trace. Models from OpenAI, Anthropic, and Gemini tend to obscure or summarize the reasoning traces so you can't see exactly what they're doing. Here's Gemini 3.7 Flash which looks like it's doing something similar: https://tools.simonwillison.net/markdown-svg-renderer.html#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6779a22d5e7bb6bdf29936f1600a5259 https://tools.simonwillison.net/markdown-svg-renderer.html#u... One of the step summaries includes this: > I'm now detailing the pelican's anatomy within the SVG. I've sketched the main body outline, including coordinates for the tail, chest, neck, head, and massive beak with a pouch. I'm focusing on the position of the eyes and considering the positioning of the wings, with the foreground wing on the handlebar for a confident look.
- kakadu 1mo ago[flagged]
- 1mo ago
- sezaidemirer 1mo agoCongratulations, it turned out great!
- codethief 1mo ago> Notably, Hy4 preview also contributed to its own development process, participating for the first time in the automated optimization of training methods, data strategies, evaluation frameworks, and low-level operators. The model proposed approaches, ran experiments, and iterated based on the results, with the resulting code, logs, and feedback feeding into subsequent rounds of exploration. This established an early-stage recursive self-improvement loop. This reminds me of one of the predictions from https://ai-2027.com/ https://ai-2027.com/ . Only that there it's "OpenBrain" doing this, not the Chinese. And the authors of that paper were also slightly wrong about "Mid 2026: China Wakes Up": China woke up already a while ago. And: > But China is falling behind on AI algorithms due to their weaker models. The Chinese intelligence agencies—among the best in the world—double down on their plans to steal OpenBrain’s weights. No need to steal anything, they have already caught up. And then there's this prediction for February 2027: > Officials are most interested in its cyberwarfare capabilities: Agent-2 is “only” a little worse than the best human hackers I think we're past that point now, too…
- bredren 1mo agoIf the distillation "attacks" created useful inputs to open weight models, ai-2027 was directionally correct that the Chinese would find ways to extract IP from western firms. (Scaled account creation and grinding outputs etc is not a dramatic story element as spies, though!) Whether the distillation has constituted "attacks" or has or will meet the bar of "stealing" IP is not super interesting to me, though.
- vlyan 1mo agothe chutzpah of calling it an `attack` or `stealing` is super interesting tho.
- unrented7977 1mo agoIt's actually turbo boring and predictable. Capitalism has long since standardized on out and out lies to influence public perception and government action.
- vatsachak 1mo agoI'm liking where LLMs are headed: They can do the difficult small level optimization, the boring but tedious code but cannot be tasteful. That means I'm more valuable and more productive. Good stuff
- handfuloflight 1mo agoThis guy gets it.
- Flere-Imsaho 1mo agoThis is basically the conclusion the creator (DHH) of Ruby on Rails has come to: https://lexfridman.com/dhh-david-heinemeier-hansson-transcript https://lexfridman.com/dhh-david-heinemeier-hansson-transcri... It's all going to be who has the best and most tasteful ideas. Interesting times indeed.
- throaway2525634 1mo agoI, for one, welcome our new Chinese overlords.
- 087532379864 1mo ago[flagged]
- andsoitis 1mo ago> open-sources link to source code?
- jychang 1mo agohttps://huggingface.co/tencent/Hy4-preview https://huggingface.co/tencent/Hy4-preview
- yipinwong 1mo agoI am going to bring up graph issue for everyone of these announcements. They all suck. They shoulda put their stick where they belong, not at far left. It just makes comparison to Deepseek 90% of them time as Hy4 has nothing to show off.
- ls612 1mo agoOpen Weights is where the action is at in the past couple months, I’d have to think the US frontier labs are getting nervous. Like Anthropic hasn’t released anything pushing the frontier since “the event” earlier this summer.
- joshheitzman 1mo agoMaybe's its a problem with the hosting at novita.ai but I didn't got much useful out of this model as a coding agent.
- coder543 1mo agoNovita does not offer Hy4-preview on either OpenRouter or their own model list. Maybe you confused it with Hy3.
- joshheitzman 1mo agoYou are correct.
- pilotcat 1mo ago[flagged]
- jamienk 1mo agoGenuine Q about word optimization/token density: If we create a stripped-down vocabulary with greater token density to use less resources and to resolve ambiguities earlier in the semantic process, aren't we creating NEWSPEAK and dragging along the worst aspects of it? The ambiguity and multi-valence of words is what creates more connections between words, increases the directionality of associations, and expands the potential subtlety and depth of meaning. By paring down (or requiring verifiability) we make it harder to say certain things, or at least make it harder to unintentionally say something that makes MORE or DEEPER sense than what we intended. If the token density becomes extreme, you're left with something like a calculator. Maybe this is the ultimate path toward better coding? But the worse path toward better genuine thinking?
- dnautics 1mo agoI don't think so. It's pretty clear that LLMs use the higher level layers for reasoning, so a bit of logorrhea very possibly enriches the result quality.
- nbush 1mo agoThis is one of the dangers. AI boosters would say that humans already do this compression and it was accelerated by mass media and then the internet, and that model memory + context can be broad enough that compared to human capabilities the opportunities for depth and variability are even greater. But I think we know which way this optimization usually goes. Even the notion of a "fine-tune for subtlety" is a contradiction.
- vatsachak 1mo agoNah reducing token length means that we're just reducing English down towards a programming language like a nice demi-glace
- algoth1 1mo agoIt's my understanding that the llm is not literally thinking those words, they are just the conversion of the matrix multiplication results (numbers) into the tokens. So the matrix is "multiplying" concepts and directions to come up with the final answer - which produces a somewhat readable reasoning trace. As far as the llm is concerned the reasoning trace could be random (to us) symbols. In fact, the reasoning traces are not necessarily optimized for readability as much as they are an emergent property of the way a reasoning model is trained
- xeonax 1mo agoWe humans should adopt grug, instead of claudish. It seems simple to understand. And has this melancholical feeling
- Yash16 1mo ago[dead]
- realty_geek 1mo agoHas anyone tried CodeBuddy? Is it worth trying out?
- DarmokTanagra 1mo agoHow funny will it be when the US economy collapses because of all the money poured into AI at the expense of pretty much everything else only for China to come out ahead anyways. Personally I hope this is China's "Star Wars" moment, the current US admin certainly seems easy enough to manipulate into catastrophic own goals.
- jimmydoe 1mo agoit's not Star Wars. china doesn't have to start any war, just let us finish (or get finished in) the wars...
- DarmokTanagra 1mo agoI think you misunderstood what I meant by "Star Wars". https://en.wikipedia.org/wiki/Strategic_Defense_Initiative https://en.wikipedia.org/wiki/Strategic_Defense_Initiative
- jimmydoe 21d agoI didn't. I'm just making fun of the words.
- bearjaws 1mo agoChina basically has been running a "do nothing, win anyway" campaign since the Trump era. More and more people view the US as a bad ally, and China is stepping in to help in many places. People should get out more, go visit South America, Dominican Republic, etc. They have cheaper AC, cheaper cars, cheaper appliances, all because they can import from China without tariffs. This is not to say they have worse quality, I would argue the opposite, much of what they import is at least as good or better than what you get in the USA. It will be no surprise to me when we end up forced to use equal or worse American AI for a premium in price, while the rest of the world moves on.
- DarmokTanagra 1mo agoSpeaking as someone living in Asia who visits the US every year or so, you are falling more and more behind and I don’t see a way for you to catch up without a drastic societal reformation. The EV market alone should scare the average citizen, but it seems like a non issue every time I bring it up. I think the US economy has painted itself into a corner and this all in bet on AI is its last real play before the house of cards collapses.
- formvoltron 1mo agoso it used enflame hardware? The policy of the US to block Nvidia should have been called. The "Let a thousand flowers bloom" executive directive. A very stable genius
- zyralab 1mo agoSeeing the model use "caveman speak" in its thoughts to save tokens is hilarious. "Why use many word when few word do trick?" is actually a legit tech optimization now!
- scirob 1mo agoJust got benchmarks done for German langauge eval index i help maintain. Hy4 is a big improvement over h3 but still below Deepseek pro and Significantly below GLM 5.3 Flash . hy4 ranks ~14th overall https://dach.peerbench.ai/compare?models=tencent%2Fhy4-preview,z-ai%2Fglm-5.3-flash,google%2Fgemini-3.7-flash,deepseek%2Fdeepseek-v4-pro-0813 https://dach.peerbench.ai/compare?models=tencent%2Fhy4-previ... https://dach.peerbench.ai/models/tencent/hy4-preview https://dach.peerbench.ai/models/tencent/hy4-preview