7 ms·
> OpenAI said it did not believe the industry “has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for m
by NichoPaolucci 17d ago
> OpenAI said it did not believe the industry “has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
Baffling. To my knowledge, they didn't properly airgap their systems. Keeping the genie in the box seems like 101 to me, and to "miss" that seems awfully fishy. This, among all of the Anthropic news, is an odd convergence.
Maybe they're being truthful and it really is the end times.
Maybe they've hit a wall in improvements, but I don't know enough on the topic to speak to that.
Which is more likely?
Either way, trying to sift through this can of worms is tiresome. I'm hopeful that this all comes to a head soon, what an exhausting few years it's been...
- deleted 17d ago[deleted]
- esseph 17d agoTons of people airgap training and test environments, and not just for AI stuff but IAC/Automation/Networking, etc.
- abixb 17d agoI'd wager on the second scenario. Anyone who's been paying attention to the industry knows that most of the 'gains' have come from test-time compute and architecting harnesses in novel ways. In my estimation, capability increases from "pre-training" alone died early last year, and we're now probably seeing test-time and other benchmark hacks approaching their limit as well. If you zoomed back to late-2024, people in the industry were predicting how we'd have AGI by now and the economy would've already 'taken off' with massive productivity growth and ushering in of great prosperity ('deflationary spiral'). Where is it? Where is the productivity growth? Where is the deflationary spiral? To be fair, models have gotten better in jagged ways, but reliability is far from usable, especially in long duration tasks, and there has been no effort by the AI companies to address the human brain's bandwidth bottleneck -- they hit the gas like there's no tomorrow and we have enormously capable but jaggedly intelligent multi-modal models with agentic capabilities that are only as effective as the human using it. This whole thing has become a giant mess.
- BobbyTables2 17d agoI even wonder if the frontier AI models are really as capable as they claim or if the companies behind them have just special cases all the “hard” questions. For example, the earlier generative LLMs couldn’t correctly answer ‘how many r’s in “strawberry”?’ due to the underlying nature of the tokens. If they get it correct today, how do they do it? It feels like we’re being deceived by the Wizard of Oz…
- pu_pe 17d agoHow do you explain the fact that Qwen3.8 27B performs vastly better than any open model from even one year ago, if using the same test-time compute and harness?
- abixb 17d ago"Vastly better" in what ways? Benchmarks? You know Benchmarks can be optimized for and benchmaxxed for, right?
- pu_pe 16d agoIt's obviously more capable in any task I tried (coding, translation, summarizing, etc). Benchmarks are not the only way to tell if a model is better or not.
- huurtehoog 16d agoI wanna see numbers showing companies and countries having excess growth due to these tools. Where are these data? It's all vibes, and the numbers contradict the vibes. There's 30 years of literature trying to explain the "productivity paradox" where we can't see any excess productivity driven by computer technology. Lots of FOMO, no hard data. For an entire generation. And people come here every day and say stuff like you just said and they really seem to think that "this time is different".
- butlike 15d agoYou gotta define 'obviously' here.
- nullbio 17d agoThe "we accidently connected to the internet" can only mean one thing: Intentionality. There's no world where this happens by accident.
- tim333 16d agoMy take isn't either. Stopping AI doing bad stuff is a real problem which can likely only be dealt with by trying it out and fixing problems as they arrive. Bit like SpaceX rapid prototyping the rockets - try it out, see what blows up, fix it and try again.