5 ms·
I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries
by FernandoTN 2mo ago
I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes.
For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...)
The real unlock will be, and you can already see it with GPT-5.6 and Fable-5, to delegate complex enough tasks that will take more than 24 hours to get done and they will not lose track. I'm not talking about a loop, but the actual intelligence to recover from these compounding errors that accumulate in dumber models.
We're still a long way from the intelligence needed to let one of these agents go ahead and supervise multiple layers of sub-agents underneath to do complex orchestration. The future looks very promising and exciting. Imagine having the possibility of a Frontier model orchestrating as many sub-agents as needed that are running on cheaper models like DeepSeek.
- copperx 2mo ago> these compounding errors that accumulate in dumber models While SOTAs handle these errors better, they compound in all models and there's a term for that. It starts with cluster and ends with an expletive. I wish I could, but I don't see the need for human steering going away soon if the task involves anything novel (see Terry Tao's chat).
- fuck_google 2mo ago[dead]
- svachalek 2mo agoThese are not 24 hours of inference with floating point errors accumulating; largely the system guards against errors compounding. Tool failures, compile failures, test failures, etc, push back against the model taking a wrong turn and force it to correct. Yes it's much easier to have a smarter model that goes straight to the correct answer first, but it may not be necessary or economical. There's a minimum bar for the model where it understands problems and knows the right step to correct them, and above that newer models give diminishing returns.
- mycall 2mo ago> it's much easier to have a smarter model that goes straight to the correct answer first That's basically ASI not AGI, if you agree humans are NGI (natural general intelligence) and make mistakes and wrong decisions in solutions all the time. Right steps with some wrong ones is acceptable though for AGI.
- xyzzy123 2mo agoThere are a lot of tasks that are hard for organisations to run consistently but require some intelligence - monitoring logs and metrics for anomalies and security events, backup audits, audit processes in general, ensuring document quality and consistency, database advice and tuning, customer experience management, process optimisation - that are not "long horizon" in the classical sense of each step depending on the last, but are the result of consistency and attention over a long period of time and a large amount of data. For this genre of task execution can run with limited horizon and is independent but would be too expensive to do with "us frontier tokens", I think for these, there is value in availability of cheaper tokens.
- Cookingboy 2mo agoThat doesn’t make sense. It’s not like SOTA models are error free, yet we still use them. You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?
- AussieWog93 2mo agoI think we put up with Fable's occasional hiccups because there's nothing better at the moment. I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides. If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entire $100 right there. Hell, you could argue that the best "coding model" that we have at the moment is the human brain, and people will gladly pay $10,000/mo for one of them. Arguing over $20 vs $100 for something that actually puts in work just seems insane to me.
- godwinson__4-8 2mo agoThe question low cost models will create: Why would you massage output? Fable 5 is still going to mess things up at any sufficient complexity. The advantage of low cost models with "good enough" intelligence is they can recursively correct. Why? Because it is cheap. Proper requirements and tests and subagents take away increasing amounts of work, at a cost that is not prohibitive. If you are reviewing code manually you might consider Fable 5 a worse option. As it articulates itself with higher confidence and you already know it is capable, you are may be more likely to miss a mistake. You know to be on guard with a junior engineer. Reviewing a senior who suddenly makes some weird stochastic mistake can be a lot harder. It would be like if the smartest human engineer you knew was capable of some random brainfart in the middle of their massive diff. Imo, much harder to deal with. Of course, we should keep in mind Fable 5 is only expensive today. It will be cheaper in the future. Autonomous, recursive prompting and improvement is the clear end state. Especially for entities that will always have the budget for that at the SOTA frontier.
- palata 2mo ago
- est 2mo ago> even small errors compound over time and can lead to catastrophic outcomes So, death sentence even to frontier models?
- dd8601fn 2mo agoI have a silly (but honest) question. What's an example or two of a > 24hr task that people are actually asking something to do? Like real life ones.
- trollbridge 2mo agoDecompiling / disassembling and annotating old software, making sure it can build cleanly back to the original binary, and then look for bugs or subtle issues. Another one I did was a printer data stream translator from an obscure format to PostScript/PDF (or just PNGs), complete with cups support, etc so these old apps can easily be hooked up. Flash is capable now of running long range defined-goal tasks like this.
- vehemenz 2mo agoI wondered this too. Also, the duration of the task depends on the quality of the prompt and the model used. I have a feeling a lot of these day-long tasks are bogged down by suboptimal tool use and on-demand python slop.
- canadiantim 2mo agoWorking through the total backlog of issues in a repo that may have accumulated from planning sessions
- spaceman_2020 2mo agoMajority of white collar work absolutely does not require sota models
- solenoid0937 2mo agoWhite collar work will require SOTA models up until the point where said models can automate white collar work entirely. Then, we might finally see a meaningful commoditization crunch.
- spaceman_2020 2mo agoI genuinely don't know any work that has been taken over completely by AI without humans in the loop managing things At this point, I don't even know if its possible
- deleted 2mo ago[deleted]
- onkarkdev 2mo ago[flagged]
- IrishLagger 2mo agoAmazing how many inaccuracies you fit in there. - You describe the breakdown in terms of time but it's more accurately a function of reasoning complexity. - You seem to assume that no intermediate evaluation is possible. - Often it is (e.g. the build breaks or tests start failing), allowing for course correction. There's definitely a cost to that but it can still be cost effective if the accuracy is "good enough" and the price difference significant. - There are numerous tasks that don't require Fable or GPT5.6 level reasoning to improve efficiency by an order of magnitude.
- yogthos 2mo agoWhat you're overlooking is that any large task can, and should, be broken down into smaller individual components that can be reasoned about and tested in isolation. This is literally the whole basis for how we do programming. You don't need a model that can keep track on a gargantuan tasks all at once. You need a process for breaking problems down into manageable chunks, and then assembling them into a solution. This is a problem that can be solved by a harness through steering and and having a decent agentic loop.