5 ms·
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is s
by dalemhurley 1mo ago
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.
Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).
Codex is slightly better than Claude Code.
Good on Sam Altman getting back to basics and turning OpenAI around.
- kroaton 1mo agoI think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
- tonyhart7 1mo agothey don't have moat in hardware either Chinese counterpart like CXMT and Huawei is begin producing their own chip You cant block an entire nation level effort with tariff
- astrobiased 1mo agoI think the moat that China has is energy costs. It's taking learnings from the Bitter Lesson. If you role up scale and compute to the next level, it's energy resources. China has it and sharing open weight models is an effective means of removing the tech moat. This idea has been floating around for a bit now (I'm not taking credit for it).
- spartacusnacho 1mo agoThey also benefit from the commodification of software/knowledge work since they own manufacturing
- rgbrenner 1mo agoIt's not energy costs. The US produces about 70% more electricity per capita. Chinese households do pay less than half what US households pay for electricity, but that's because the NDRC sets prices below costs for households. They make it up by charging industry more, and the industrial electricity prices in China are roughly 34% higher than in the US.
- haldujai 1mo ago> The US produces about 70% more electricity per capita. And consumers use 4x as much per capita. Industrial generation per capita China comes out ~2x > industrial electricity prices in China are roughly 34% higher than in the US For which industrial customer and where? Chinese compute hubs are on par to slightly cheaper on pure electricity costs. Conversely the US makes it more expensive with interconnect and upgrade fees as well as hefty take or pay contracts. A 1GW datacenter in VA for example would add 5-10c kWh and a 12 year take or pay deal
- thinkthatover 1mo agoper capita seems the wrong metric given the difference in population sizes and America's wealth. They have roughly 4x the people and have added 10x new power capacity in the last 10 years, not to mention lapping us in renewable and long distance transmission lines added.
- VirusNewbie 1mo agoIf there was no moat, nvidia and meta would have SoTA models too.
- seunosewa 1mo agoMeta is awfully close.
- dansquizsoft 1mo agolol! Good one...
- nwienert 1mo agoWent from years behind to months pretty quick.
- amazingamazing 1mo agoIt is not in nvidia’s interest to be too good at model creation
- david-gpu 1mo agoWhy not? Commoditize your complement, and all that.
- angulardragon03 1mo agoAnd if they get too good, they risk harming or otherwise killing their golden geese (their customers), who they are heavily invested in.
- david-gpu 1mo agoHow? Imagine an open-weight model comes out that is somehow better than proprietary solutions. Now the marginal cost for the consumer is just the cost of renting the inference hardware, without having to pay the overhead of the owner of a proprietary model. And because it is cheaper, more customers want to use it, and Nvidia will sell the providers the inference hardware that they need.
- davidguetta 1mo agobringing the price down b.c. competition != no moat. There's not 100 frontier labs, it's not like airline companies
- haldujai 1mo agoAbout the same, 5-10, when you consider major (aka frontier) airlines. Actually not a bad comparison. Both burn massive amounts of up front capital to protect an oligopoly in the hopes their commodity product eventually pays off.
- keeganpoppen 1mo ago[flagged]
- asa123 1mo agowhy so much negativity and certainty?
- Razengan 1mo agoThe "moat" is the "harness", the app. For most people, the app IS the AI. And even for its wonkiness, ChatGPT has had the best UX/UI of them all. The way to win the AI wars in the eyes of the common folk is through the frontend, to be the Apple of AI, as it were.
- tw1984 1mo agothis basically says you don't believe there is real AI.
- Razengan 1mo agoRead the second line guy There are people all over the world who have no computer skills but they use ChatGPT on their phones daily They don't know/care shit about models and all that For them, if the app sucks, the AI sucks.
- m3kw9 1mo agoThey have a lot of moat, i'm not sure what youa re talking about. Only amatures are using Qwen, open source stuff that is 3-8 weeks behind. Plus OpenAI has some verticals that keep people in there.
- scronkfinkle 1mo agoIn what way do they have a moat? A cursory look at https://artificialanalysis.ai/models/gpt-6-astra#intelligence https://artificialanalysis.ai/models/gpt-6-astra#intelligenc... it lands at 61, only a single point above glm 5.3 while costing significantly more. The only moat they appear to have is by hoarding compute, and the current trajectory of hardware shows that isn't permanent either for very long
- bitexploder 1mo agoI wish people could see how some of this reads. You are an “amateur” using a model 6-8 weeks behind? Really? Sigh.
- aurareturn 1mo agoI think it mostly shows that there is no moat You can argue that TSMC has no moat since Intel and Samsung are also able to eventually make a node as good as TSMC - just a few years later and at smaller scale. And no one would say that about TSMC. So there is clearly a moat there somewhere.
- coolandsmartrr 1mo agoYeah, I'm not sure if "no moat" analogy stands for chip manufacturing. Even if foundries acquire lithographic nodes, the procedures (temperature, duration, etc) are for them to figure out and are usually kept secret. This secret could be the "moat" that differentiates each foundry's operational capabilities.
- saithound 1mo agoNo. In the semiconductor industry, the "catch-up" player isn't normally spending less in absolute R&D terms. Comparing the R&D costs of creating GPT-4o vs. DeepSeek V3 (the latest gen for which we already have good accurate numbers) it looks like the latter cost 1/20th as much to create. If Samsung could catch up with TSMC for 1/20th of the cost, people definitely would say that TSMC has no moat.
- aurareturn 1mo agoWhy do you think Chinese models cost 1/20th to train?
- saithound 1mo agoThat's the ratio the widely published numbers give [1]. One does not have to believe the numbers [2], but those who do believe them are then justified to conclude that there's no moat. Which numbers you believe is of course going to affect whether you think there's a moat or not. That's largely orthogonal to your TSMC/Samsung analogy I responded to. If you think the "moatists" are wrong because they believe the wrong numbers, that's fine, but then there's no need for the analogy. [1] https://galileo.ai/blog/llm-model-training-cost https://galileo.ai/blog/llm-model-training-cost [2] https://medium.com/@theiand/how-can-deepseek-a-5-6-million-llm-outperform-openai-and-meta-38e995b35140 https://medium.com/@theiand/how-can-deepseek-a-5-6-million-l...
- treefry 1mo agoFrom my experience with complex coding tasks (AI infra), I don't think these open weight models are close.
- akie 1mo agoNot sure if I agree, I tried GLM5.3 and it was pretty decent. Ok, it's not Opus, but maybe it's Sonnet?
- upupupandaway 1mo agoTheir ads business is also doing well. Not "will recover all compute costs" well, but crossed $1b in a few months.
- jeffybefffy519 1mo agoIts funny, my experience with Sol has been awful. It really overworks problems and tracks into areas it does not need to... I just dont get how its good for some, and bad for others. It makes me suspect that the models performance is not even against problem sets and it really is just a probabilistic prediction machine. Which then makes me very skeptical of GPT-6 Astra, because if their big claim is Computer Use then it is probably bad in a bunch of other areas.
- embedding-shape 1mo agoIt is funny indeed, people sometimes with same amount of experience with software development, get vastly different experiences from different models and harnesses. > I just dont get how its good for some, and bad for others. If I were to listen to my hunch, it would tell me that it's all up to the prompts that ends up going over the wire (including all the bloat some people have), what workflow/process you use and what the existing state of the project is.
- oliviayii 1mo ago[dead]
- ragequittah 1mo agoYou have to bake the 'lazy dev'/'keep it simple stupid' mentality into your AGENTS.md and / or the skills you're using to design things. It will take things too literally sometimes so you also have to make sure you're being accurate. Best way I've found to use it is make it ask you clarifying questions about what you're trying to build and have it help design the shape of the thing. Then it writes the instructions in a format it understands. I've had Claude do the same thing where it goes off and spends 100% of my tokens on 3 functions and an ungodly amount of tests / scaffolding that do almost nothing when I gave it an underdeveloped idea.
- John7878781 1mo agoThis is what Google needs to do and is probably why Demis has stepped back a bit
- zachthewf 1mo agoI’ve found Sol performance to be incredibly spiky. It has tremendous IQ and can fix very difficult bugs. But it is horrible at design (both visual and system design), anything that involves thinking about users or UX, and massively overcomplicates almost all work.
- ghosty141 1mo agoI noticed the same. I wanted a simple crud webapp and suggested an insane techstack involving C#, Razor Pages, MSSQL and more. I went with my planned setup of python flask with an sqlite db which served me well for years. It's still incredibly important to have a human in the loop correcting design decisions and having good taste.
- jiggawatts 1mo ago> insane techstack involving C#, Razor Pages, MSSQL Is a very sane tech stack, you're just biased against Microsoft. Half the world's enterprise apps run on that combination, or a minor variation of it. Like Java it is full featured ("batteries included") but unlike Java it is relatively terse and actually pleasant to work with. Oh, and unlike Python, it is very fast, within spitting distance of compiled Rust and C++ web apps.
- kyleee 1mo agoThere are a million and one reasons to be biased against Microsoft, regardless of the fact that C# tech stack is decent
- ghosty141 1mo agoCorrect I'm biased against technologies that only run on a single OS for no benefit to the user.
- jiggawatts 1mo agoASP.NET runs on MacOS, Linux, and Windows. Microsoft SQL Server now (also) ships as a Docker container that runs on Linux. > Correct I'm biased against technologies that only run on a single OS for no benefit to the user. Do you ever use software that only works on Linux? Do you use an Android phone?
- Implicated 1mo ago> Sol is so much better than Fable 5. ... looks around ...
- andxor 1mo ago> Sol is so much better than Fable 5 I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else? Sol is a much smaller models and it shows. It often misses the forest for the trees.
- enraged_camel 1mo ago>> I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else? Same. It makes me wonder what types of things the person must be working on.
- resonious 1mo agoThis is perpetually an issue with the whole field of AI/LLMs. The experience is so personal. Every time I talk to someone about their use of LLMs for software engineering, I'm shocked by their approaches and experiences. They say "X model keeps missing things" when I rely on it heavily for being thorough. They say "Y always gives me the best results" when I can't stand it. People will see/think that I'm doing very well with my LLM use, and ask me what I'm doing. I tell them, they try it, then later they come back to me saying they just couldn't get it to work.
- ilikecode 1mo agoIt’s really inconsistent. There are sessions where it nails everything perfectly and I leave happy. Then there are sessions where every turn it corrects itself and changes it mind. One session recently I found it funny how every single time it did this one task it tripped over itself and killed its own connection. Like 20 times. It didn’t bother me I just found it odd how despite it being noted down in its state file it kept doing it over and over like some idiot. Literally they can’t learn from their mistakes yet.
- mf2hd 1mo agoThis is the job now, we are shepherds.
- bsndjdjdjdj 1mo ago[dead]
- fastball 1mo agoCodex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.
- carljungslabtek 1mo agoIt has had automode for a bit now. I use it every day at work.
- user43928 1mo ago"Approve for me" used to work without issue for me when GPT 5.5 was the latest model. Nowadays it routinely rejects "git push" to the project's repository.
- fastball 1mo agoAFAIK it's just yolo mode, which doesn't actually do any checks like Claude Code's auto-mode (which has a model checking all commands for safety).
- carljungslabtek 29d agoNo, it says “auto approved” after every evaluated command. Does it work? No idea, I can’t think of a time it ever stopped, but I’m also not ever doing anything dangerous with it. It does use another model to evaluate the commands though.
- drschwabe 1mo agoPut it in an isolated container and set it to YOLO
- superdisk 1mo agoI don't even put it in a container. I just yolo it. It hasn't done anything bad so far. My logic is that if it nukes my filesystem then I deserve it.
- ChadMoran 1mo agoSol better than Fable? What? I've found it to basically be on part with Opus and I max out 2 accounts on both providers every week.
- fnordpiglet 1mo agoCodex is missing a few things that Claude code has had for some time like defined plugin subagents and a few other things. But overall it’s fairly capable. The biggest gripe I have is that codex really restricts context window sizes and compaction leads to a lot of grounding work, and overall codex GPT is too literal in many situations - it’s follows direction slavishly, and when subagent reviewers are used, they tend to find increasingly obscure “flaws” on the instruction following impetus, and the harness agent takes them literally as issues to fix even when it leads to bizarre outcomes. For instance I’ve had several runs where it tries to end up building a hermetic system with sha hashing of everything (including operating system binaries and kernels, tool chains, etc) to certify test results are valid, etc. I have to sort of watch it carefully to be sure it’s not drifting into some insane yak shaving corner, which it will happily do for weeks on end. Claude has the exact opposite problem, especially opus-5, where I literally can’t trust it to print hello world without taking a shortcut, or just simply lying and saying it printed it when it didn’t, behind a giant wall of inscrutable text. I find it very ironic that Anthropic is the vendor of the lazy lying cheating model that does almost everything you tell it to it do. I’d really kill for something that balances instruction following and loop escaping behavior better. Fable 5.1 does seem a lot better, feeling more like 4.6 behavior, and honestly Sol has improved as well. I’m pretty psyched for the next generation, as I think the competition has heated up so much that things will improve really fast to the point of marginal utility opportunity being increasingly close to epsilon.
- swingboy 1mo agoYou can enable the 1 million token context window and adjust when it compacts in your config. > model_context_window = 1000000 > model_auto_compact_token_limit = 900000 I believe it does consume your usage a bit faster though.
- fnordpiglet 1mo agoUnless things changed this used to work but was blocked. You also used to be able to force in a model catalog to get around the token cap. Each, as of at least April, were closed off and ineffectual. I’ve not tried lately so you may be right.
- openaiscooked 1mo agoKilling Sora was one of the worst mistakes they ever made
- fooblaster 1mo agoplease tell us why
- openaiscooked 1mo agoRich media is where all the innovation is happening now and in the future. Text-to-text is dead, has been since Mistral 7b. Solved problem (you guys like that one don’t you) They also demoted themselves from “authority on AI” to “in over our heads” by bowing out in the pathetically defeatist way they did at the worst time possible (Hailuo/MiniMax/Vidu coming up) - they naturally completely missed the wave on audio with random companies like Singify taking that market for free. They just bowed out. They didn’t try. They didn’t try anything more than baseline text-to-text and they aren’t good at that (or code) either, compared to what others are doing. It’s a really bad position to be in if you’re trying to be an Apple or Microsoft. To have a mediocre product and then can’t even serve 75% of the mainstream use case.
- TheyTrappedYou 1mo agoYour right
- Otterly99 1mo agoI never used Sora but I recently tried Google Flow and the results are quite good. I have the feeling that Nano Banana 2 has been the image generation champion since it's release so maybe OpenAI feels that they cannot outcompete Google on this task?
- nimbleal 1mo agoGPT 2 is almost always better (80+% of the time?) than Nano Banana 2, albeit slower
- nsoonhui 1mo ago[dead]