13 ms·
GLM 5.2 and the coming AI margin collapse
- LoganDark 3mo agoI hope cheaper inference eventually means faster speeds at the lower tiers. I don't want to settle for 100 t/s, but I don't want to pay $10 per prompt either
- copperx 3mo agoWhich raises the question, which are the fastest frontier models? are the enterprise hosted Anthropic models faster than what Anthropic serves? Somehow no one talks about LLM speed.
- tough 3mo agoOAI has announced an upcoming 750tok/s 5.6 served through their cerebras acquisition
- LoganDark 3mo agoYeah, Cerebras is the one with competitive speeds nowadays but they cost an absolute fortune. Also they don't host good models publicly. Good to see OpenAI leaning into them, can't wait until these speeds are available by subscription
- dcl 3mo agoThat is going to be absolutely wild for whoever can access/afford it.
- manquer 3mo ago> cerebras acquisition Partnership you mean?, Cerebras went public and are trading at around 45B in market cap. While OAI could in theory cough up that kind of money, it would massively hamper their existing committed capital outlays.
- tough 3mo agoYes sorry, i got confused some how and mixed up the partnership announcement [1] with an acq one, maybe i should get some cerebras stock then, ty for the pointer 1. https://openai.com/index/cerebras-partnership/ https://openai.com/index/cerebras-partnership/
- LoganDark 3mo ago> Somehow no one talks about LLM speed. When I've raised speeds about local inference I've been told 60-75 t/s is perfectly usable. It makes sense that people aren't talking about speed yet since you either already have a response fast enough to wait for, or you go do something else and check back in a few minutes. I would love to wait for the latter type of tasks though, because those are typically the ones that require the most work from me to verify and I don't want my attention divided with multitasking.
- esafak 3mo agoGLM 5.2 has a Fast variant at 200-400 tps.
- KronisLV 3mo agoThey have a vision MCP to make up for the model itself not having the capability natively: https://docs.z.ai/devpack/mcp/vision-mcp-server https://docs.z.ai/devpack/mcp/vision-mcp-server I also found their web search to be mostly okay. Furthermore, in case this is of interest to anyone, if you use their ZCode harness then you get bigger Coding Plan quotas: https://zcode.z.ai/en https://zcode.z.ai/en Used it for a bit, it sits somewhere between OpenCode Desktop (still new but nice) and Claude Desktop (recent versions are good). As for GLM 5.2 as a model - with max thinking it’s generally satisfactory, somewhere between Sonnet 5 and Opus 4.8, better than DeepSeek V4 Pro for sure. Pricing wise, the subscription doesn’t seem as good as expected. I spent like 60% of the weekly limits of the Pro (50 USD) plan in one day, only because each 5 hour limit only gave me 20% to spend, otherwise it’d be 80-100%. Not even doing anything crazy, just parallel long form work on 2 projects with about 96% cache rate and at most 3 parallel code review sub-agents. Their Max (100 USD) subscription would last me the whole week, but so does Anthropic for the same money and so would OpenAI. Off-peak is more palatable but I can’t just twiddle my thumbs at 9 AM to 1 PM local time. Proper savings would show up with the Max plan and yearly billing, but that’s more of a tough sell.
- esafak 3mo agoThey also have GLM-5V-Turbo. https://docs.z.ai/guides/vlm/glm-5v-turbo https://docs.z.ai/guides/vlm/glm-5v-turbo
- zackify 3mo agoI switched to yearly Cline pass because it was too cheap haha
- richardfey 3mo agoI can't find on their website some indication of what kind of usage I can get out it, otherwise I'd be interested.
- zackify 3mo ago$6 a month I plan to use deepseek v4 flash mainly which should provide closer to 5x the usage on the cheaper ones but no set number
- budsniffer952 3mo ago>the least understood upcoming shift in AI economics. Then proceeds to talk about something in the AI news every day. Hey, did you guys hear? Open source models are cheaper and their quality is increasing! So, first, by no measure is GLM5.2 as good as Opus. Second, yes, open source models will put pressure on margins...eventually. Everyone knows that. But do you think today's AI business model is the same as tomorrow's?
- A_D_E_P_T 3mo ago> So, first, by no measure is GLM5.2 as good as Opus. Depends what you do. Complex tasks, poorly-defined tasks, sure. For relatively simple tasks, though, or very well-defined tasks, it's just as good and usually a lot faster. It also has a more neutral character and is somewhat less adversarial than Opus. (Opus is always "Let me push back on that..." whereas GLM is "sir, yes sir!") I use both and I appreciate both. If Opus disappeared tomorrow, though, I wouldn't cry -- I'd be able to adapt to a GLM-5.2-only life real quick.
- stingraycharles 3mo agoI think the point is that if you’re doing simple, well defined tasks then Opus is overkill and you’d want Sonnet instead. Meaning, GLM5.2 is Sonnet-quality, not Opus-quality.
- AlotOfReading 3mo agoI think it's interesting to note that in one year we've gone from they're not even close [0] to arguing whether open models are only as good as sonnet or opus. [0] https://news.ycombinator.com/item?id=44623953 https://news.ycombinator.com/item?id=44623953
- stingraycharles 3mo agoI see the exact same discussion as we’re having right now there; people stating that local models aren’t as good as the state of the art, but good enough for certain tasks.
- fny 3mo agoI'm not convinced raw costs matter: 1. Compute costs collapsed since the advent of Cloud and yet hyperscalers still have fat margins. 2. Many open source office suites exist yet none compete with the ubiquity of gsuite or office. GitHub, Slack are similar examples. 3. Both Windows and macOS dominate the home desktop space despite free alternatives existing for a long time. 4. Many formerly open source infrastructure components like Redis and Elastic Search have Apache equivalents, but they still command healthy margins. I understand the arguments for a margin collapse, but I don't see any historical analogues. It seems that enterprises will pay top dollar for service guarantees, integration, and someone they can sue. It's nobody gets fired for buying IBM all over again.
- 3abiton 3mo agoThe target audience is different. Coding is mainly a trade of the tech savvy, who like many on r/localllama users do not hesitate to deply on 16GB Vram gpus. Even if so, it is estimated that within 2 years we will be able to run Claude 4.8 on consumer hardware give the rate of improvement of open-weight LLMs, which will put more financial pressure on "paid" labs. It's just a matter of rate of improvement which is shrinking between open-closed models.
- fny 3mo ago[dead]
- tuvix 3mo agoA lot of those things you mentioned have sticking power because they’re familiar to folks and migrating to something else is a big deal. I can’t imagine most people would be able to tell the difference between Sonnet and GLM 5.2. If the infrastructure around the model you’re using doesn’t change, then swapping models is extremely easy.
- arikrahman 3mo agoI agree with swapping models making it easy. With openrouter, I just change the provider. With reasonix harness, cache hits are basically free. And that's with unsubsidized American providers like Digital Ocean or cloudflare.
- softwaredoug 3mo agoI also think we’ll approach a point where increasing intelligence is not really going to suddenly improve most work tasks. I bet that’s already happened actually. We’re oohing and aahing about models, when the ones a few versions ago did a good enough job for most of the dumb coding, etc we do
- dofm 3mo agoThe thing is they are inventing new things people will want to do. But for example, "loops", fully hands-off agentic coding etc., seem really unlikely to get much traction because that just isn't how software is designed within its producer/user community. Requirements evolve in use, and fully hands-off LLMs simply cannot be trusted to only change the things you ask them to change, so I don't think it's likely that products will, in the main, be developed that way. And if you don't need that fully-hands-off stuff, then the models that run on at least reasonably modest desktop hardware are surprisingly close to being enough.
- aetherspawn 3mo agoI don’t think this is true. All the models prior to Fable were honestly dumb as rocks, and Fable is too sometimes, but at least it’s helpful now and not a hindrance. The future of AI most definitely involves making something twice as good as Fable that is virtually its own employee, and not on reducing inference costs because to be honest Fable isn't actually that expensive. The real utility behind an AI model (imagining that it can be made twice as good as it is now) would be being able to scale a small business up and down instantly without hiring (to implement a new feature or whatever), which is costly and time consuming these days.
- Havoc 3mo agoI’ve been on a GLM coding plan since they launched ~year ago and it’s been at „good enough“ since the start. Tangible behind absolute SOTA but like you say most coding isn’t rocket science.
- throwdbaaway 3mo agoSeems like a pretty pointless post that still centers around output tokens. In agentic coding, cached input tokens is 90% of the API "cost". It doesn't require GPU compute, and DeepSeek has shown that it can be done 50~100x cheaper with MLA/CSA/HCA, and a whole bunch of disks. This should collapse the margin.
- nozzlegear 3mo agoAren't the American AI labs desperately struggling to find a market beyond just agentic coding?
- le-mark 3mo agoI have heard but don’t have first hand knowledge that at least one company (financial services BPO) has moved most of their previously manual processing to llms. The person I talked to wasn’t forthcoming with any detail. This is what we’d expect to see though.
- dgellow 3mo agoAll AI labs. Not just Americans
- throwdbaaway 3mo agoThe current top comment in https://lobste.rs/s/ua1gxl/glm_5_2_coming_ai_margin_collapse https://lobste.rs/s/ua1gxl/glm_5_2_coming_ai_margin_collapse correctly zoomed into cached input tokens, but landed on the opposite conclusion: > That is, for your $100/month fee, you get $3600 equivalent of API usage. This is presumably because Anthropic has figured out some clever things to do with model routing and input caching, and also can subsidize with investor money and take a hit on their operating margins. My take: this is exactly what Anthropic wants everyone to think. In reality, 90% of that $3600 are for cached input tokens, that can be made to cost next to nothing, as shown by DeepSeek.
- throwdbaaway 3mo agoWhile we are all speculating, Boris kindly provided some guidance in https://news.ycombinator.com/item?id=47880089 https://news.ycombinator.com/item?id=47880089 > The challenge is: when you let a session idle for >1 hour, when you come back to it and send a prompt, it will be a full cache miss, all N messages. We noticed that this corner case led to outsized token costs for users. In an extreme case, if you had 900k tokens in your context window, then idled for an hour, then sent a message, that would be >900k tokens written to cache all at once, which would eat up a significant % of your rate limits, especially for Pro users. Using the current Opus pricing, that pre-lunch 900k tokens should roughly consist of: 720k input tokens = 0.72 x $5 = $3.6 180k output tokens = 0.18 x $25 = $4.5 900k 1h cached writes = 0.9 x $10 = $9 500M cached input tokens = 500 x $0.5 = $250 $267.1 in total, with 93.6% from cached input tokens. The portion that requires GPU compute is about 3% of the total. Post-lunch, the 900k tokens should consist of: 900k input tokens = 0.9 x $5 = $4.5 900k 1h cached writes = 0.9 x $10 = $9 So Anthropic is fine with the $267.1 accumulated over 3~4 hours before lunch, but not fine with the $13.5 incurred immediately after lunch. Why? The only plausible explanation is that the actual cost of caching is way less than the API pricing. If you use a coding plan, Anthropic doesn't really care about your cached input tokens usage. Indeed they want you to show your ccusage screenshots. On the other hand, if you pay by API tokens, the margin is huge for cached input tokens. Only when you do something that requires a lot of FLOPs, e.g. the post-lunch 900k input tokens, the cost becomes real.
- montroser 3mo agoThis article only promises to get into "the coming AI margin collapse" in a yet to be published "part two". This part only makes the point that GLM 5.2 is pretty good (no shit).
- esafak 3mo agoIt truly is a pointless article.
- benjiro29 3mo ago> I'd be very surprised if it wasn't more than 50% cheaper for nearly all workflows, for a very similar level of quality. If your using pure API ... providers like neuralwatt cut that cost down even more by using energy as the actual cost. So GLM 5.2 is more expensive then GLM 5.1 on their service (those thinking tokens), compared to API costs, its dirt cheap. And way more tokens then the zai subscription delivers. We are seeing a move towards more realistic pricing on actual consumption based usage. Be it DeepSeek, Xiaomi (MiMo), or zai's GLM via neuralwatt. The main issue facing subscriptions a-la-carte usage, is that a lot of the heavy hitters really drain the resources. And that as a business model can not survive without ... a) increasing the prices. b) everything goes to actual token/energy usage based billing but with more realistic pricing, and not the bloated API prices that are focused on companies. We shall see what the future holds but things will change.
- felixfurtak 3mo ago> It turns out that nearly every agentic session does a lot of web searching for looking up items This is why Google will win the race over most of its competitors. They own search.
- Applejinx 3mo agoIf they did I wouldn't have had to go to DDG. It's not like it's a big jump over what used to be. I left claw-marks in Google Search, if they drove me off they're in trouble, because I didn't want to accept reality for quite some time.
- redrix 3mo agoI wonder if this is an alternative (and better) revenue stream vs ads for search engines: Offer a competing web search for LLMs as an alternative to Google, and charge enterprises and LLM providers for it. I know Brave do this already. Not sure about DDG (I wonder if their agreement with Bing would allow it?)
- felixfurtak 3mo agoBuilding a good search engine is expensive. Perhaps not as expensive as AI build out. Market share is currently Google (91%), Bing (4%), Yandex (<2%), Baidu (<1%), Brave (<1%) Google can and do already monetize automated search from AI models. Heck, if they wanted to, Google could turn off search and make you go through their AI model to get information. Imagine that. That's how powerful they are.
- jazzyjackson 3mo agoKagi assistant IMO does a great job giving relevant material to the LLM. It's a pretty neat way for a search engine to charge a premium, to offer a good model on top of their results.
- cmrdporcupine 3mo agoWhich race? As an information-providing "oracle" type model, maybe. For practical agentic tasks? Not even close. Gemini is blatantly incompetent at tool use in an agentic harness. Even their own.
- codedokode 3mo ago[dead]
- 0xbadcafebee 3mo agoYes, margin on model inference is high with some providers. If you just wanted inference (at cost), you'd buy a GPU, or rent one from AWS or Microsoft. But you're not paying OpenAI/Anthropic for inference. You're paying them for a platform. Every feature OpenAI/Anthropic bake into their applications, models, online services, etc - anything that isn't pure LLM text generation - is a custom integrated add-on service that LLM weights do not include. Even if open weights became cheaper and better than OpenAI/Anthropic, most people would still pay for OpenAI/Anthropic, because they give you things the weights alone don't give you. Comparing Z.ai GLM 5.2 to Claude Code w/Opus 4.8 is like comparing Linux Kernel 7.0 to Microsoft Windows 11. If you don't know much about computers, you'd say these are the same things. If you know a lot about computers, you know the latter has a thousand extra things that make a huge difference in what it does out of the box. Which one you use speaks to what kind of customer you are. Sure, GLM 5.2 doesn't have vision; but an AI power user can plumb together any VLM with the text generation of GLM 5.2 in most AI harnesses, just like a Linux power user can combine the Linux kernel with KDE Desktop. Most people don't use Linux and KDE, because it's unpopular, difficult to use, hard to get support for. Instead they pay for Windows or Mac, because there's lots of support, with a giant company pouring money and effort into filling all the usability gaps, making it seamless. Most people don't pay for the cheapest possible thing. They pay for the thing they can afford that improves their life while making it easier. An open weight alone is almost completely unusable by itself (like the Linux kernel), compared to an AI platform (a completely usable system). If you're constantly wondering about when open weights will reach parity with OpenAI/Anthropic, you're a Linux person. If you just pay $20/$50/$100 for OpenAI/Anthropic without thinking about it, you're a Windows/Mac person. There is nothing wrong with either of these groups, but they are fundamentally different, and always will be. An LLM weight is simply a different category of thing than an entire AI platform/provider.
- sailfast 3mo agoHow long will that $4.40 rate persist? Until we know more about the real unit economics it will be damn near impossible to rely on steady inference costs or make them predictable at the enterprise level. Gonna be a wild ride for awhile.
- markasoftware 3mo agoMultiple providers (who need to make a profit) offer the same 4.40 rate for glm-5.2. It's not subsidized. Deepseek's 0.86 or whatever is likely subsidized but alternate providers offer it for a price comparable to glm-5.2.
- est31 3mo agoGPU/RAM/etc prices could continue to rise. If the world leaders decide it's time to build the robot armies, then that could price out the civilian uses for GPUs.
- throwa356262 3mo agoAccording to deepseek themselves, their current rates are NOT subsidised. They have published tons of articles dedicated to performance and efficiency engineering. Feel free to have a look...
- markasoftware 3mo agoWhy is no other inference provider offering similar prices then?
- throwa356262 3mo agoHow long did it take vLLM to implement deepseeks sparse attention from the r1 paper? Does ananyone outside deepseek have a working code for the v4 compressed attention mechanism? Has any other provider managed to bypass CUDA and program the compute engines in their native assembly language to get 10% more performance out of them? There is your answer.
- zuzululu 3mo agoi would use glm 5.2 if the servers weren't in china i mean i guess my employers wouldn't know the difference but i'd like to play it safe and keep everything in america
- tarpitt 3mo agoits open-weight. I think you can find a host for GLM-5.2 in the USA
- kristianp 3mo agoIf you look at https://openrouter.ai/z-ai/glm-5.2#providers https://openrouter.ai/z-ai/glm-5.2#providers there's about 28 providers, including z.ai and Alibaba. Most outside of China. I've never seen so many providers for a model on there before, glm 5.2 is popular.
- zuzululu 3mo agothanks I see cloudflare has it. I will give glm 5.2 a try
- _pdp_ 3mo agoIMHO, cheaper inference means higher costs overall :) because everyone will use more thus driving up the investment required to stay current or to compete. Switching models is also kind of easy but not plug-and-play. Most harnesses out there do very poor job with the open weight models. Unlike Opus, GLM 5.2 ends up in loops and hallucinates a lot more. If your harness is built on the expectation that the LLM will perform well, then switching to GLM 5.2 will be an uphill struggle. We had to refactor our harness and introduce more defences because of GLM. The cost savings are substantial. Obviously it really depends on your workloads but it is noticeable cheaper for agentic work. Coding - I don't know. We do have some coding agents on GLM 5.2 and what I noticed with some landing page experiments that the results between GLM and Opus are identical - they might be using the same training data? Obviously Opus is still substantially better model. I don't think there is an argument to be made here but GLM 5.2 is cost effective and really good too. Overall, we switched all of our internal agents to GLM 5.2 and because it is Open Weight we are in talks to get the model from certain geo locations giving us more freedom as well as extra protection. Overall I think this industry will be in much better place because of GLM 5.2 and whatever open-weight models come next.
- cma 3mo agoAre you running unquantized GLM-5.2 and getting in loops or quantized?
- spyckie2 3mo agoIt’s important that none of these entities can collude to price fix. Having China be the competitor ensures that. Basic microeconomics is still the easiest way to understand token economies. How is it not a competitive market (where profits go to zero?). Anything A or O does to keep more margin, any competitor can copy or choose to undercut, and undercutting has the benefit of collecting training data. So what is going to stop gross profit of tokens going to zero except for collusion/price fixing?
- intrasight 3mo agoYou left out the one that will: federal government industrial policy
- twelve40 3mo agoSo the federal government industrial policy is the thing that supposedly will keep the prices on "A and O" high in the US while the rest of the world will get comparable AI competing to get cheaper and cheaper?
- regularfry 3mo agoTerrible ideas get executed all the time, despite the problems with them being well understood.
- CuriouslyC 3mo agoBasically, the US govt will say that foreign models and providers are a security risk and ban them. If the US has shares of Anthropic/OAI due to a sovereign wealth fund, it'll be billed as domestic industry protectionism too.
- intrasight 3mo agoYup. Just like electric cars. Actually, more like network gear - so only non-western countries.
- HDThoreaun 3mo ago
- redrix 3mo agoThe fact that these Chinese models are getting close to “Opus-grade” despite costing 6x-8x less is huge. As the token bills start to come in, those economics will be harder to ignore (regardless of the origin of the LLM); especially as there will be many CIOs sweating over their quick and costly AI initiatives showing little ROI. My hope is that the EU also steps up their own competition in the frontier model space so that it’s not just China v USA.
- AgentMasterRace 3mo agothey're not near opus at all, anyone using the models in a real working environment will tell you the same thing. on paper they have impressive benchmarks, but that's not realistic to actual use.
- shrinks99 3mo agoI've been using GLM 5.2 a lot this past week, it's been replacing Opus 4.8. I mostly do front-end web development and haven't noticed much of a quality difference. Sure, "it's just frontend", but that's actual use enough for me to take it seriously.
- drbscl 3mo agoI think it depends on your use case. For my personal projects (a mix of webdev & some Rust desktop apps) it's honestly very close to Opus 4.8 (which I use in my day job). I don't feel like I'm missing out after cancelling my personal Claude subscription, whereas I used to feel that way a few months ago.
- samuelknight 3mo agoInference has been decreasing in cost by about 10x per year since 2023.
- maxglute 3mo agoHow fast is glm 5.2 in western hosts? It's doing everything I want it to, but going through PRC host it takes like 5-10 times longer. Not sure if that is nature of modest or PRC computer infra/routing.
- vachina 3mo agoGLM feels faster and more reliable in my experience. Anthropic and OpenAI models would hem and haw or straight up timeout during peak times.
- dbalatero 3mo ago> Of course, this was a hugely poor read of where the costs actually lie in AI. Training - while no doubt capex intensive - is a fixed, up-front cost. You spend hundreds of millions to train a model, then you are "done". I don't understand this point that people make. If you're consistently needing[0] to train new models and the cost of training relative to the % improvement seems to go higher, isn't this just a constant cost that you continue to bear? The footnote seems to allude to this, but then sort of waves it away anyways. Also are there continuing incremental training costs to keep models relevant? Or do they only have knowledge of events up to the day they were trained? [0] needing, because you have competitors and people expect more and more.
- blourvim 3mo agoThese models rely on knowledge that are embedded in their weights, if a new library is released, a new linux version comes out, some new protocol succeeds the previous one, you want your llm to know about it. Sure you can just add that into the context window, but that has its own problems. Unless new research, there are a few which look promising, gives a new method, training is going to be a constant cost sink. On top of this, if you stop training, it is 6 months until someone releases an open weights model and now you are competing to give the lowest price for the same product. Also we can't forget that this is a business that *has to* be in the global labor industry, not just a tech tool, they have to have much better models to justify the trillion dollar evaluation
- rbranson 3mo agoit does not require training a model from scratch for it to be updated. the entire LLM training process is iterative. essentially each step (there are hundreds/thousands for a training run) yields a complete, usable model. tokens with updated data can be added on top essentially at any time in the future.
- dbalatero 3mo agoThat's good to know, but I think maybe the bulk of the original point is still concentrated in the "we always need new models to stay relevant" piece.
- gnarbarian 3mo agothe economics of this are a little counterintuitive. is there a market saturation point for intelligence? how about for software? it seems like the more you have the more you want because you're trying to do more things. as the models get smarter I get busier because I'm doing more things...
- yogthos 3mo agoThere's definitely a saturation point depending on the complexity of the problem you're solving. For example, any model can write a small shell script to resize a video with ffmpeg for you right now, so it doesn't matter whether you're using a local Qwen model, GLM, or Fable. They'll all do a roughly comparable job and you'll end up with a working script that does what you need. Then you have things like CRUD apps, where a model needs to write some SQL, make a service endpoint, serialize some JSON, etc. Here a local model might have a bit more trouble juggling all the pieces, but any hosted model will do just fine. If your day to day job involves working on CRUD apps, then it's basically a solved problem now. The cases where frontier models matter are when you're solving genuinely complex problems, but that's not what most people are doing day to day. So, paying an order of magnitude for a model that has capabilities to solve problems outside the range of problems you actually work on becomes a waste of money. There's going to be a market for these models from people who really do work on complex things on regular basis, but the question is how big that market is. Additionally, open models keep getting better, and GLM 6 or DeepSeek v5 could end up being another big jump in capability where they fully close the gap with Fable. At that point, even more of the market becomes covered by these models leaving truly complex cases on the frontier. Another thing to consider is that most big problems can be broken down into smaller ones. That's the basis for how programming languages are structured. We have primitives which are arranged into functions, that get bundled into classes or namespaces, and so on. So, you don't need an infinitely capable model to solve big problems. You just need to be able to break large problems into smaller ones, and a model that's smart enough to decompose a problem to the point where it becomes tractable.
- gnarbarian 3mo agoif you give me a smarter model I will find a more complex problem for it to work on that will require an exponential number of sub-agents. the smarter the model the more capable it is to marshal resources to accomplish the goal. this means you don't need a smart person to be able to take advantage of this. a smart model will be able to do the same.
- yalogin 3mo agoI don’t understand the argument here. The article doesn’t describe a collapse or the breadcrumbs for it. The only argument I can put together is companies hosting the open source models in house or use some service like Amazon that could potentially host them and so replace the frontier models. Data center and specifically infra to host llms is still the main sticking point given the security concerns about data going to china. The article doesn’t make these arguments coherently
- aussieguy1234 3mo agoI would not be unsurprised if the US govt steps in to prevent this. They'll do anything to stop China getting ahead in the AI race. There's the sanctions already implemented, next step might be giving these companies government funding, just like they do with military companies.
- yomismoaqui 3mo agoGood luck trying to enforce that outside of the US.
- aussieguy1234 3mo agoI posted this some time ago https://news.ycombinator.com/item?id=48759668 https://news.ycombinator.com/item?id=48759668 Singapore seized a mansion due to Nvidia chip smuggling. So there are some countries that will enforce sanctions.
- kzrdude 3mo agoDon't Anthropic and OpenAI both already have military contracts? They are already growing into that fabric I think.
- lqstuart 3mo agohow do massively negative margins "collapse"
- bluegatty 3mo agoAs long as the SOTA models are 'ahead' then there will be a big premium.
- blagui 3mo agoGLM is the model that will sink the frontier labs. Recall last year deepseek? And 18 month's later? What changed?
- ddxv 3mo agoA year ago I wasn't using Deepseek. Now I am. I guess what changed is which models people are using most for coding.
- tw1984 3mo ago> GLM is the model that will sink the frontier labs. this is the claim you are making here, no one else claimed that. two obvious issues here - 1. GLM itself is a frontier lab, ranked No.3 in the world in July 2026, ahead of Google, Meta and xai. GLM is not going to sink itself. 2. GLM won't sink OpenAI, it will significantly restrain OpenAI's profit margin. OpenAI will still be able to get stupidly high market cap, but not trillions, hundreds of billions will be far more likely.
- maxdo 3mo agoin cursor benchmark glm5.2 is on par with gpt 5.5 medium and sonnet for the same task from results and cost perspective. The speed of generation for both gpt 5.5 medium and sonnet 5 will be dramatically faster. source : https://cursor.com/evals https://cursor.com/evals I don't get the hype. It's near SOTA model that is not deepseek of this world. It an expensive to run model, and under certain tasks it is comparably cheap as closed source ones.
- ilaksh 3mo agoI think the profits depend on how well they manage their fleet purchases (or possible sub-leasing?) to get high utilization without overloading or idle racks. Because accelerators like H200, B300 etc. are highly parallel and designed to run like 200 or maybe 300 sequences at once (depends on the model, just guessing). I assume they finance the hardware and that cost per device or rack is the same whether each unit is handling 10 requests or 150 requests (aside from electricity). And probably international customers factor into it to get good utilization over more of the night time. And it likely is something that they look at quarterly more seriously than monthly. The biggest risk to profits might be a downturn in business that causes some portion of the financed AI accelerators to go idle or get low utilization for some weeks (that they can't sublease).
- jaggederest 3mo agoSomeone on HN made a comment in one of these threads that we could bake the weights into something like Cerebras's wafer scale chips and serve essentially the entire world off a single wafer, which is a pretty wild thing to think about. You'd have to make new hardware any time you trained a model but that seems really worth it.
- ilaksh 3mo agoWell, Taalas has that kind of technology, but the chip they demoed is probably 20-100 times smaller than necessary since it's only an 8b model. But let's say they could someday scale that up to a much larger model, 72 large chips per wafer and each chip can do 1000 LLM requests at once (Vera Rubin?). So it's roughly the equivalent of an NVL72 rack. You might be able to serve something like 50000-60000 requests at once. So I think it's more like handling a small city's worth of customers per wafer than the world if you had that. I believe in less than 5 years we will get to that, but the model size and/or number of agents is going to keep going up also.
- Atotalnoob 3mo agoYou’d never be able to update it’s knowledge. LLMs need retraining to incorporate new knowledge. Baking them into wafers means they will be out of date by the time they finish the first wafers.
- AgentMasterRace 3mo agoI don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , different libraries and cut corners. Fable said about 80% was implemented wrong. I had GLM 5.2 do the same, and it performed exceptionally better, but when it got stuck on something it would be trial and error mode going forward and have zero foresight for future issues that might occur due to fixes it was trying. the model severally lacks prompt understanding, and testing .
- vachina 3mo agoOpus is good but not consistently good. That’s a problem. I’m paying the same but not getting the same results.
- jmyeet 3mo agoI think OpenAI, Anthropic and SpaceX are going to envy the dinosaurs because there's not asteroid coming for them, there's three: 1. There will be no moat around frontier AI models in the future. China is going to make sure that happens. It's a national security interest for them. DeepSeek was the first shot across the bow for that but it won't end with them. There are other labs and there are non-Chinese actors too. The stratospheric valuations depend on there being that moat; and 2. Nobody seems to be considering what the next generation of AI hardware is going to do with current hyperscalar investments. We're about to go through this with the B100/200 move to R100/200 but a lot of the investments are probably slated for that next-gen. But what about 3 years from now when the hypothetical X100/200 comes out and doubles FLOPS and halves performance-per-watt. What will that do to existing investments? Some people are delusional and think that they'll get 10 years out of GPUs when 10 year old GPUs (eg V100) are sold for scrap and 5 year old GPUs (A100) cannot run DeepSeek v4 Pro. And people think the A100 is going to get another 5 years of use? No; and 3. Local LLMs are coming for remote usage. You can buy a 5090 PC for less than $5000 currently but you're limited to 32GB of VRAM, which will comfortably run 31B models but nothing really larger. Go to $12-13k to upgrade to an RTX 8000 Pro and you have 96GB of VRAM, which will run larger models (but certainly not, say, DS v4 Pro or even Flash). You have shared video memory products rapidly coming from NVidia's aggressive market segmentation. Things like Strix Halo and DGX Spark have severe limits on memory bandwidth (<300GB/s compared to 1.8TB/s for a 5090/6000 Pro and 3TB/s+ for server grade HBM3e/4 based GPUs). Macs could be real interesting in this space butr they lack the raw FLOPS with the M5 generation. But what will this local hardware look like in 2-3 years? I think people will be shocked at how much better it will be with the Apple M7 Pro/Max generation (2028 expected) and the RTX 6000 cards at that time although I fully expect NVidia consumer GPUs to still top out at 32GB of VRAM to maintain that segmentation. And I look forward to what the next generation of the AMD Ryzen AI Halo platform will look like if they really try. All of this adds up to these three companies needing to cash out before the music stops (IMHO).
- regularfry 3mo agoOn 1 and 3, the obvious move is to shift the bulk of the harness behind a new API that's not based on raw LLM access. Then they get to hide secret sauce behind that API and all three go from commodity to premium while simultaneously being able to try out whatever tricks they can get away with to reduce their own inference costs. I'm almost surprised this hasn't happened already.
- TacticalCoder 3mo ago> Where it gets really scary for the frontier labs is how easy it is to migrate to open weights models. Both Z.ai and Fireworks offer both an OpenAI compatible and Anthropic compatible endpoint. This makes it absolutely trivial to use with Claude Code and Codex. Yes the ease of switching is greatly appreciated. Now the reason I tolerate Claude Code in my tmux sessions is because apparently Anthropic ain't playing nice with the subscription plans and other harnesses. But I'm evaluating pi.dev atm and it looks amazing. To me being able to rid of that piece of vibe-coded underperforming, characters-modifying, turd that Claude Code is a big motivation to switch to GLM (I'll probably keep my OpenAI subscription as OpenAI repeatedly said they were cool with other harnesses). It's also quite obvious that Claude Code is receiving new vibe-coded slop features after vibe-coded slop features in an attempt to lock you in. To anyone thinking about switching to GLM: I'd say at least evaluate pi.dev and see if that wouldn't be an opportunity to kiss Claude Code and its "gameloop that converts characters from a headless browser to other characters to show in a terminal at 60 fps" goodbye once and for all.
- sroerick 3mo agoPi.dev is great and with only a little customization made even previous gen open weights feel superior. It also doesn't feel like they're trying to sell me on transhumanism all the time. It also doesn't get mysteriously downgraded. It's just consistent, even before 5.2. 5.2 is great in a lot of ways - but it's best quality is that it gives some pushback and isn't nearly as synchophantic
- WhitneyLand 3mo ago“Z.ai provides a replacement MCP for web search, but it's pretty awful and slow” I’ve had good results with Tavily so far, might be worth checking as an alternative for agent search.
- Obertr 3mo agoMetaphor i like is that it will be as cheap as electricty? Do you know who is supplying your electricity or which factory it runs on? probably no, bc its a commodity and mostly settled and there is so many energy resources. some are alternative some are coal mines. And they all fight in the supply demand trade for energy which is happening real time ( think open router here) And eventually the consumer wins bc of the abundance. I think greatest example of abundance of cheap infinite intelligence will be not glm5.2 but DeepSeek V4 Pro max with $0.435 per 1M input tokens and $0.87 per 1M output tokens
- cdavid 3mo agoThe metaphor breaks quickly because as a first approximation the electricity "quality" does not depend on the provider, and will not change overtime. That's not true of LLM output.
- Obertr 3mo agodisagree. you don’t care if your task was solved by 130 iq or 160iq model. what matters is that it surpassed let’s say basic 120iq barrier and price that’s why glm5.2 is a drop in replacer for most of the population. not fable 5 really
- scritty-dev 3mo agoBraintrust which is a really solid eval tool/platform just compared it to Opus 4.8 to see if it could preserve exact long context retrieval under prod serving constraints and it did really well. I think 6-12 months before OSS has Fable-esque models
- itkovian_ 3mo ago‘640kb of ram should be enough for anyone’
- RichardChu 3mo agoIs this AI written (or edited)? The word "genuinely" appears 4 times on the page. Man, I hate how often people/LLMs use that word now. Maybe other people gloss over it but it's super distracting to me.
- blondie9x 3mo agoDisclosure - Fireworks kindly gave me some free credit to experiment with GLM to help write this article.
- hn1986 3mo agoneeds to be higher up
- 01100011 3mo agoI'll agree but from the other direction. AI continues to absorb my job as a senior systems software engineer (c/c++) and after a couple months I've only spent a few hundred dollars using gpt-5.5/5.6 and codex. I have no idea what people are doing to burn so many tokens but for me this is laughably cheap and every day I discover new capabilities. I don't care if costs go up or down, it's so cheap for what I get that I don't care.
- throwawayffffas 3mo agoAsking claude to implement a single feature that takes under 30 minutes consumes 10-30 dollars of tokens in api costs.
- poisonborz 3mo agoWhy would you do that? Why not use it to create a detailed plan and let it implement by a low cost model?
- jeremyloy_wt 3mo agoAnd if the engineer bills for $100/hr or more, the trade off is worth it
- mynameisbilly 3mo agoWhy is everyone still operating under the assumption the current token costs will remain so heavily subsidized? We could see $200-400/hr in token costs once these companies need to turn a profit
- Paradigma11 3mo agoBecause then people will just flock to open weight models that are even cheaper and Anthropic/OpenAI will go bankrupt. Then somebody will buy those companies for cents on the dollar and run a profitable token selling business for today's prices without doing any new research. New model releases will slow to a crawl but that's it.
- jillesvangurp 3mo agoI think the fixation on numbers of tokens and dollars per token is missing the point a bit. LLMs are quite useless without good tools. The article calls out search as one of them. And it's important. If you are coding, the tools are relatively easy: they are mostly open source and don't have a lot of authorization logic around them. Anyone with access to python and some access to a half decent AI model can pull together a decent agentic coding tool. There are many examples out there. But if you look at the overall market, there's a rapid shift happening to non-coding tools and non programmer users starting to become very active. This kicked off beginning of the year with Claude Cowork. OpenAIs Codex and ChatGPT (they both have the same plugin infrastructure) is doing a lot of the same things. I've talked to a lot of non technical business users in recent months. There's a growing amount of people who definitely have zero interest in programming starting to use these tools and getting value out of them. This is going to rapidly scale to essentially most white collar users. Programming tools are becoming a side show to this market. The difference here is that these people need connections to all their favorite protected data SAAS silos: MS office, Sales Force, Outlook, Gmail & GSuite, Calendar, SAP, Oracle, etc. The moat here is very different: it's mediated access to these silos in a compliant way. Anthropic announced a solution in the form of some MCP features. Those features boil down to getting access to all your favorite silos, if you sign in with the right identity provider. What's the right identity provider? The one that's whitelisted by the data silos you are locked into. Okta seems to have weaseled themselves into a position of power here. And it's all the other usual suspects. We'll see who is going to "win" that race but I bet it's going to be a pretty exclusive club with zero outsiders from China on that list. You can hack your way around some of those limitations. But doing so in a compliant way is going to be tricky. And that's before you consider who's going to pay for this and what they are going to insist on. Corporate IT departments & data security policy compliance basically. What's the moat here? Secure & compliant access to all your favorite silos. Here in the EU that also includes data residency. The difference between sending all your data to Silicon Valley or Beijing is that of getting stabbed or getting shot. If it leaves the EU, you have a huge compliance issue. Most of the juicy corporate LLM usage is going to have to be fully compliant. I.e. hosted and controlled in the EU. This will be the same across the world. The least important choice right now is which model you use. The most important ones are about where those models run and what tools the models running there have access to and how that is governed. On paper, OpenAI, Anthropic, MS, and Google are pretty well positioned here. Not necessarily in that order. Most others are still figuring it out. But they'll have a moat of data center ownership in the right regions + mediated tool access that works out of the box.
- seydor 3mo agoBut what if they introduce a tarriff per token
- joka88xj 3mo ago[flagged]
- Alan_JoshyMJ 3mo ago[flagged]
- deleted 3mo ago[deleted]
- Kavon2992 3mo ago[flagged]
- pixlmint 3mo agoLast month, I cancelled my Claude Pro subscription and instead used those 20$ to purchase Openrouter Credits. Most of my knowledge-seeking questions can be answered by Gemma4, for basic code editing, Qwen3.6 27b is enough, and for really difficult tasks, GLM5.2 doesn't leave me hanging. I'm by no means a heavy AI user, so I'm even saving money going the API Credit route and relying on the smallest possible model depending on the task complexity.
- NostraDavid 3mo agoWhat do you use as interface to OpenRouter? I, too, am looking into using an API to see if I can reduce costs (I use OpenAI + Github Copilot, currently). TensorX instead of OpenRouter (because it's in Europe, and EURouter wanted 15% more money from me :P), but I'm not sure if I want to change a configuration in vscode every time I want to switch the model in the Claude extension (and having an API key in my settings feels iffy too >_>)
- k8sToGo 3mo agoI use Openwebui
- pixlmint 3mo agoI have OpenWebui hosted on my homelab, but you can also just have it live on your machine in a docker container. I honestly just embrace the iffy feeling. Openrouter has very good telemetry (which is partly why I went with them) and it'd be pretty easy to notice when someone other than myself uses my keys. For the little agentic coding I do I like to use OpenCode, and if I need to ask a question in my editor I use CodeCompanion (neovim AI chat plugin). I quickly went to check what OpenCode does with the API key, and it doesn't seem to store it in the user config, so that's at least something. But yeah, really recommend OpenWebui as a ChatGPT replacement (though there are a lot more alternatives out there, I just already knew owui from when I was playing around with local models)
- khimaros 3mo agopi.dev
- peepee1982 3mo agoThere is mention of GLM 5.2's poor web search capabilities, but I see that as a harness responsibility. I've set up my own SearXNG instance on my VPS and integrated it into Pi alongside the webfetch tool, and GLM 5.2 has so far been great at finding things. I asked it to give me the current news from an Austrian online newspaper that's difficult to parse because of its aggressive ad overlays. Both ChatGPT and Claude failed in their native chat apps. GLM 5.2 in Pi was clever enough to search for the RSS feed and gave me a detailed overview. The lack of vision is a real shame, though. I've implemented workarounds in Pi that are okay, but they're not as good and the whole experience feels awkward.
- xienze 3mo ago[dead]
- epolanski 3mo ago> I expect for most professional use the very lax terms around training and data retention will make this a difficult sell I couldn't care less whether a chinese or american company reads my crap code. I'm not working on state secrets but warehousing software for specific clients on a machine that has access to nothing but crap enterprise code.
- impjohn 3mo agoThis take may be frowned upon in HN, but I suspect it's much closer to the median sentiment.
- typ 3mo agoUnlike the belief that frontier AI is expensive due to a high margin, and going to be expensive if there is no competition. My understanding is that, under certain circumstances (which is most likely true), the price will be driven down just because of profit seeking. The frontier LLM labs run on a huge fixed cost and very low marginal cost. They need the economies of scale to make sense of the business (an incentive to expand their user base as large as possible). Imagine that you want to buy a few B300s to run GLM 5.2 and rent the service out to other people. How could this business be viable and sustainable in the first place? You need as many customers as possible. If you charge everyone $1000, you find fewer customers who can afford it. It rots the ROA if the servers are not utilized 100% (you would better buy less compute instead). Also, the marginal cost for onboarding a new customer is low. And it's getting even lower when you have more customers. You wouldn't leave money on the table (especially for your competitors) if you want to maximize your profit. By this logic, all frontier AI labs are incentivized to lower the price to maximize their customer base, profit, and ROA.
- Gareth321 3mo agoI agree, but there is prestige to consider. Many people are motivated to buy the best, even if it's much more expensive. "We're building a mission critical application here. Sure the API costs are much higher, but it's worth it."
- echelon 3mo agoI could spend $1000/day with Fable and it would be worth it. It has much deeper systems thinking, enabling me to trust it to follow my instructions and not fuck things up. 1. That confidence and quality is worth the price. 2. We're accelerating at lightning speed now. If you don't spend, someone else will and they'll eat your cake. We're nearing the point where you could spin up an entire YC startup in a day. That changes the economics of everything.
- lgl 3mo agoI hear this all the time lately. Things like AI X or LLM Y can create a fully working company like "BigCorp XYZ" in N days/hours etc. But is speed of creation really the golden goose here? A few skilled and motivated individuals could also do (and have been doing) that. Sure, maybe they take a few months instead of days or weeks, but AFAIK, having a product is just a tiny bit of the battle, finding customers, product market fit, and actually growing it is where the gold is so I'd argue that you'd be better off building the product with a $100 day LLM and spend the other $900 on marketing. AI won't automatically make everybody business gurus and every LLM generated company a unicorn.
- synapsehire 3mo ago[flagged]
- jFriedensreich 3mo agoI am confused by this article landing on front page, it does not seem to contain any new insights but besides mostly reading ok falls apart when mixing up model and harness comparison in an amateurish way. Why would the subpar search tool zai provides be relevant for comparing models? They did not even mention the >capability< to use said tools but talk about MCP/ search providers as if thats not an implementation detail.
- inigyou 3mo agoWe'll keep saying the same things about every new free model just like we say the same things about every new frontier model but nothing really changes.
- devinabox 3mo ago[flagged]
- throwthrowuknow 3mo agoI don’t see it. GLM 5.2 seems noticeably worse than Opus and especially GPT 5.5, the poor vision capabilities are also a massive strike against it since these are a huge improvement in the frontier models that can make all the difference when working on anything visual. Running it locally is its biggest advantage but for a lot of use cases that isn’t needed and is a burden to set up and maintain.
- devinodowd 3mo agoI wonder if anyone has actually measured the difference in verification time between these models. A senior dev in a high cost of living area costs the company something like $200 an hour. If a cheaper model produces code that takes an extra 20 minutes to debug or verify because it missed a subtle edge case, you have already lost any savings from the lower API bill. It feels like the real moat for labs like Anthropic is the level of trust a reviewer can have in the output without reading every single line. Curious how much people trust GLM over something like Opus? Is there much of a marginal difference here?
- tokoi 3mo ago[flagged]
- senordevnyc 3mo agoThis is exactly how I see it, and why I’m often willing to pay Fable rates. Yes, I might spend $300 on a feature instead of $20, but it’ll be done in a tenth of my time, thanks to faster iteration in both the planning and post-impl verification / iteration stage. Plus a mistake could easily cause a bug that costs me a customer worth four or five figures of LTV, and the feature might easily add tens of thousands in marginal revenue over the coming years. So really $300 is nothing, it’s barely worth worrying about.
- zurfer 3mo agoSomehow the blog post seems naive. Yes GLM 5.2 is good and cheaper per token, but margins are a result of supply and demand. Now demand for quality and quantity of tokens is increasing at least quadratic or cubic (more users * more tasks * more tokens per task). On the other side you have real infrastructure constraints on the supply side. Openai and Anthropic have large commitments and contracts that enable them to get access at a scale of compute that is not obviously going to be available for open source model hosts. And you see it, glm 5.2 inference is less stable and higher variance than any of the bigs labs. Why is SpaceX not hosting glm 5.2? because they make more money with renting out to Anthropic and Google.
- alansaber 3mo agoAgent systems only increase the gap between frontier and open models. Open models still experience more tool call failures, run longer loops, and get stuck more often. Until that's resolved (and it's obviously technically possible) people will be forking out for a better agent experience.
- alexmercerdev 3mo ago[flagged]
- a_c 3mo agoRegarding the lack of vision part, if you are using Claude or opencode, I've made a skill[1] that let's you talk with any models in Claude/opencode mid-session. You ask "Have claude opus to look at this PDF for a second opinion" during a session of claude with GLM5.2 or opencode with GLM5.2 It doesn't need to pass whole conversation history as context (unlike /model), you can ask follow up to that forked model (which sub agents in claude doesn't support AFAIK), and you can ask models from opencode while using claude. [1] https://github.com/kmcheung12/second-opinion https://github.com/kmcheung12/second-opinion
- msephton 3mo agoFWIW I've seen subagents remain open for followups on latest Claude.
- a_c 3mo agoWill give it a try thanks for the heads up. Mind if I ask if you are referring to SendMessage? I was testing on Claude code 2.1.196. SendMessage was not available. Skimming through their change log didn’t seem to have anything related to SendMessage. There is “ Fixed SendMessage silently misrouting when a re-spawned agent reuses a previous agent’s name — the tool now detects the mismatch and asks the caller to retarget” from 2.1.199. Not sure if they are related
- msephton 3mo agoI have no idea what things are called but somebody showed me how they could start a bunch of subagents and steer each with additional messages; some closed after their task was completed, but one stayed open for followups and continued working after being prompted by the main agent who was acting as triage.
- Dantlv 3mo ago[flagged]
- moktonar 3mo agoThe problem is that the more AI eats labor, the more you can hike the costs until you pretty much can match the salary of the workers you replaced, with some margin, enough for the user to accept the cost. That’s what will happen in the next decade IMHO. price = base expense + what user accepts to pay
- chrisss395 3mo agoWhere does the harness come in to play? I'd love to use GLM 5.2 for general chat, but I don't know of any harness that offers an experience close to ChatGPT or Claude (e.g., history, skills, projects) without requiring a PhD to set it up.
- xienze 3mo agoYou can point Claude to another model if you want. See ANTHROPIC_BASE_URL: https://code.claude.com/docs/en/env-vars https://code.claude.com/docs/en/env-vars
- glimshe 3mo ago"There is no doubt that using Z.ai's official API and subscription is almost certainly a non-starter, with their terms being at best weak and the deep connection to Mainland China." This is the key statement in the article. I think people don't realize that these "open" weight models exist because giving away your product at a loss is a time honored marketing strategy. There's nothing guaranteeing that the next iterations will be open (remember "Open"AI?). The Chinese labs are profit seeking companies. If they can't recoup their investment through API use, they won't be able to train more models. But if the argument is 'who cares, training models will be so cheap anyone will be able to do it ',then check the comment elsewhere on this comment section about free alternatives for consumer and enterprise software. Oh... And the variation 'what we have today is already good enough for everyone' argument is just another incarnation of '640Kb should be enough for everyone'.
- lukax 3mo agoWell, Microsoft just started offering Kimi K2.7 through Copilot hosted on Azure. https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-available-in-github-copilot/ https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-av... Cursor Composer 2 and 2.5 are also fine tunes of Kimi K2.5 It looks like politics don't matter when it comes to economics.
- nttylock 3mo ago[flagged]
- wg0 3mo agoEveryone is declaring GLM 5.2 as something that's really a big deal. I don't know about that but based on my own experience with Deepseek v4 Lite alone (with high effort) I have no doubt in my mind that anyone claiming such great things about GLM 5.2 must be true because Deepseek v4 already is really awesome.
- traceroute66 3mo agoThe blog author complains of "lack of/poor web search capabilities" in GLM, but you can always use it against an MCP of which there are many. For applications where I am not concerned about my queries being passed through a US provider, I have had success with exa[1] There are also other ways to give it context without web-search. For example the various MCPs that make `man` pages available. I've also found GLM to be quite strong for coding tasks without the need for web search. So it also depends what you're doing. [1] https://exa.ai/ https://exa.ai/
- jeremyjh 3mo agoIf you use a good harness or add the right tools and plugins both the image and web search issues mentioned are non-issues. oh-my-pi (omp.sh) handles images for text models out of the box - as long as you have any vision capable provider enabled, it will be used when you paste images to a text model. Rather than let it guess I configured it to use MiniMax M3 for this task (as well as other utility tasks like code exploration & library functions). opencode has plugins that do the same thing, but I haven't used it since picking up omp and haven't tried them. In open harnesses you can also configure your search provider(s) separately from the model provider - if you've got a ChatGPT sub you can use just their websearch for example. I've been using Kagi's API and found its cheap enough not to matter to me at all. As for slowness, I'm not sure I'm really seeing that in terms of wall clock time. The author says GLM uses more tokens for reasoning but doesn't explain how they know that - frontier models don't provide nearly the entire reasoning trace. I have the suspicion that the author is not aware of that fact. I use Opus with Claude Code for work and I find it subjectively slower because I can't read its CoT trace. That is another HUGE benefit of GLM: I can't tell you how many times I've seen it start to go sideways in its CoT - usually due to something I didn't tell it - and I just stop it and give a course correction rather than wait a whole turn. Overall I agree with the takes from the article and frankly its sad how much cope I see on Twitter (and even here) from people that think AI coding is busted once subscription subsidies are dropped. GLM is already good enough and cheap enough to use it at API rates - but it is MUCH more expensive than other open models that are also very nearly good enough. In twelve months I'm confident you'll be able to get equivalent results at API rates for less than $1 per million output tokens, and more likely that will happen in six months. Deepseek v4 Pro is already almost there (and at only $0.85/MM) - and at least on benchmarks its already better than GLM 5.1 which I was happily using quite a lot before 5.2 dropped. I haven't tried Deepseek since I already have a z.ai pro sub that I locked in for $30 - at $72 its a lot less compelling.
- stack_crafter 3mo ago[dead]
- apwheele 3mo agoAgree with the web search point. (I would like for Perplexity to start to offer more models out of the box integrated, like they do now with OpenAI/Gemini models.)
- cmiles8 3mo agoCompanies like Amazon taking out loans to fund more AI infrastructure, coupled with AI companies that are massively overvalued and burning cash, coupled with companies saying there’s no ROI from AI investments, coupled with the margin from selling AI models trending to zero paints a nasty picture that can’t continue much longer.
- davedx 3mo agoMeanwhile: > China’s Ministry of Commerce has led meetings over the past month with major AI companies, including Alibaba, ByteDance, and http://z.ai/ http://z.ai/, to discuss measures that would restrict overseas access to cutting-edge AI models, including models that have not yet been released. > The discussions reportedly include not only closed-source models but also open-weight models. > Future regulations could take the form of a tiered framework based on technological capability. Basic open-source AI models may be managed through a filing system, high-performance models may be subject to security reviews, and the most sensitive frontier models may be banned from public release or restricted to use within China https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/ https://www.reuters.com/world/beijing-is-looking-curbing-ove...
- alecco 3mo agoYep https://news.ycombinator.com/item?id=48816025 https://news.ycombinator.com/item?id=48816025 And EU leadership completely destroyed Europe's future by betting on depending on US and Chinese models. https://pleias.ai/blog/fable-eu https://pleias.ai/blog/fable-eu
- myrmidon 3mo agoThis is pure speculation at this point tbh. You could just as well read the european approach as a bet that frontier models will be unable to keep a significant edge over open competition (and thus not worth throwing subsidies at, because any economic advantage is fleeting at best). Looking at the data and related past experience, this looks like a pretty solid bet (despite the "risk" being hard to quantify).
- sajithdilshan 3mo ago> You could just as well read the european approach as a bet that frontier models will be unable to keep a significant edge over open competition And that's a bet they will lose 100%. Once the Chinese starts imposing export bans/controlling the access to their models, Europe would be at the mercy of US/China to allow them access or just rely on miserable mistral
- Pakvothe 3mo ago[flagged]
- wartywhoa23 3mo agoAn entertaining thread on just what a dismal rat race the life of a clanker shepherd has become.
- Roark66 3mo agoThe author mentions lack of good Web search. I've been using slightly modified crawl4ai and searXNG together with firebase for the rare sites that insist on throwing wrenches in the works of my LLMs. I also have my fork of metamcp that replaces firebase MCP spec with my own that tells the model to use crawl4ai and SearXNG instead. I've been using this wia Librechat with every commercial and open weight model I tested. The search is way better than OpenAI and what ClaudeCode uses, but Gemini is way faster. That will change soon as I'm planning to put these instances in a DC with gigabit pipe. Firebase is not cheap, but it retrieves everything, bypasses captchas and so on.... As long as one uses it for 1% of Web queries the cost is manageable.
- s8kur 3mo ago[flagged]
- dparkmit 3mo agoI run a 2 Claude setup (one architect, one coder) and have been using it extensively. But I don't like the fact that I have to pay $200/month to use it. I'm going to try GLM with my current setup.
- pier25 3mo agoInference margin is irrelevant. That’s like a gas station saying they have 90% margin over pumps but still losing money.
- ddp26 3mo agoPeople have been making claims about the commoditization of llms since chatGPT, and they've been wrong every time as quality and prices and differentiation have increased.
- segmondy 3mo agoGLM5.2 is not the concern, it's that there's many good Chinese models and labs, glm, qwen, deepseek, kimik, mimo, longcat, minimax, hy3, Ernie, etc They have figured out how to train, plenty of them and are consistently doing so
- seydor 3mo agoIt wasn't really a secret, the main constraint is lack of compute resources
- DIHCAPITAL 3mo ago[flagged]
- mountainofdeath 3mo agoHistory is littered with the corpses of companies that had exceptional but expensive products that were replaced with cheap, good enough products.
- gexla 3mo agoAdd me to the list of skeptics. 1) A lot of companies simply won't trust Chinese models 2) These companies will still trust a direct frontier provider over a non Chinese company hosting Chinese models. 3) Frontier model providers working with cloud hosting providers for data residency. 4) Hosting your own models is expensive (hardware, people to setup and maintain, risk of little to no return.) And smaller companies likely won't go this route. 5) As mentioned in the article, a good but cheap model can still be expensive if it takes longer to do a task. I imagine I could save days doing work with Fable vs using Chinese models. ETA: And loads of demand. Not sure the pricing situation is yet settled. As Chinese models get more popular, they may have to raise their prices as well. US models may struggle more with meeting demand rather than getting margins squeezed.