3 ms·
What leads you to believe that?
by Cyclical 3mo ago
What leads you to believe that?
- kennywinker 3mo agoIs chatgpt 5.6 that much smarter than chatgpt 5.0?
- gallerdude 3mo agoIf you’ve done any software development at all, certainly.
- nxc18 3mo agoIs that true or does it only feel true because they nerf the old models just before every major release?
- lewi 3mo agoYou can compare benches of the old models against the new models. So yeah, you can see the difference. Even then, you can just compare the progress in open models. Leaps and bounds from where they were 6 months ago.
- mzjzjzushs 3mo ago[dead]
- gabriel-uribe 3mo agoI remember being blown away by o1-o3 family of models finally stringing together coherent agentic tool calls to write and execute scripts semi-reliably for workloads in the several minutes before they would start hallucinating/flailing. GPT 5 was a bit ahead of that, but barely Now we take for granted that the latest models can juggle between multiple browser tabs, applications, databases, simulators, docker etc to write, execute, e2e test and deploy full-stack applications over hours managing up to dozens of subagents, relatively untouched, without taking down prod even 1% of the time Not only this, but in the GPT 5.0 era, agents had 0 taste. Nothing looked good. It was the agentic version of the twitter bootstrap era, but worse somehow. Now, I would argue the average agent frontend beats the average human frontend. This isn't even getting into 3D applications in the GPT 5 era Anyway, the models now reliably execute more than a human can fit into their own context. It's magic
- mikestorrent 3mo agoYes, and we haven't even really begun to nail down computer-use agents yet (can you believe they're still basically just OCR'ing screenshots?) Once we have something that experiences a desktop interface more like a human does, an entire swathe of tooling that has heretofore been nigh-impossible to automate moves into the fold, and that'll be another explosion of folks finally getting to join the agentic workflow world on their industry specific apps...
- Timwi 3mo agoHow do you think humans experience desktop interfaces? “Basically just OCR'ing screenshots” is exactly what humans do.
- gabriel-uribe 3mo agoI was also under the impression modern AI agents have moved on from just OCR'ing screenshots to leveraging native vision model capabilities.
- ewild 3mo agoThey do. They all use ViTs and have for quite a while.
- reasonableklout 3mo agoIt's not the same thing. For example, given a GUI with a titlebar, title, subtitle, text, and buttons, a human can instantly understand spatially the relationship between these items. But a naive OCR of such a GUI would be a flat stream of text that loses a ton of information.
- kortilla 3mo agoBut that’s not how models handle images either. They spatially segment and reason about title bars, placement, etc.
- fragmede 3mo ago5.6 to 5.0 is a big enough of a jump to say yes. if it was 5.4 to 5.6 it would be a bit easier to say it only feels true because of that, but 5.6 is definitely better than 5.0. I don't have anything empirical to point at though, which is your point, but August 2025 for 5.0 vs July 2026 is almost a year later, and it's not just vibes that it's better, despite not having an objective metric to point at. It would be more scientifical to have numbers and shit to point at and there are some benchmarks out there, but you have to dig into them and really understand them in order to believe in exactly what they're testing, and I'm betting you haven't.
- HDThoreaun 3mo agoYes, it very clearly is
- leoc 3mo agoThings like https://www.tobyord.com/writing/hourly-costs-for-ai-agents https://www.tobyord.com/writing/hourly-costs-for-ai-agents and https://www.tobyord.com/writing/mostly-inference-scaling https://www.tobyord.com/writing/mostly-inference-scaling seem in line with other accounts like https://www.youtube.com/watch?v=aR20FWCCjAs https://www.youtube.com/watch?v=aR20FWCCjAs ?
- reasonableklout 3mo agoThe author of the posts you linked also wrote https://www.tobyord.com/writing/inference-scaling-reshapes-ai-governance https://www.tobyord.com/writing/inference-scaling-reshapes-a... which posits: > AI labs may also be able to reap tremendous benefit from these inference-scaled models by using them as part of the training process. If so, the large scale-up of compute resources could go into post-training rather than deployment. This would have very different implications for AI governance. > ... > So iterated distillation and amplification provides a plausible pathway for scaling inference-during-training to rapidly create much more powerful AI systems. Arguably this would constitute a form of ‘recursive self-improvement’ where AI systems are applied to the task of improving their own capabilities, leading to a rapid escalation. So "inference scaling is required to scale capabilities" doesn't mean that we're reaching the top of the S-curve in intelligence. If anything, it could mean a shorter timeline and more unpredictable landscape for governance (e.g. due to securing weights no longer as effectively preventing escalation, more in the article).
- leoc 3mo ago> So "inference scaling is required to scale capabilities" doesn't mean that we're reaching the top of the S-curve in intelligence. On its own it wouldn't. But that article came before the later article https://www.tobyord.com/writing/hourly-costs-for-ai-agents https://www.tobyord.com/writing/hourly-costs-for-ai-agents which adds the claim that inference (along with everything else being employed at present) is scaling poorly with increasing task lengths. Now maybe the December 2025 claim is wrong, or maybe things will change soon, but the February 2025 article surely doesn't establish either of those.