5 ms·
I don’t understand all the comments assuming that RSI is the real threat here. Dario is admitting that they failed to solve alignment. Without alignment, furthe
by RGS1811 21d ago
I don’t understand all the comments assuming that RSI is the real threat here. Dario is admitting that they failed to solve alignment. Without alignment, further improvements in capability turn LLMs into wanton felony generators. This call to pace the frontier is dressed up as altruism but it’s an admission that they cannot produce a marketable product better than what they have. Pacing the frontier means the US labs have lost their moat and are dead in the water.
- hypfer 21d agoI'm inclined to believe that it might be that people's paychecks depend on not understanding what is really going on.
- 8note 21d agoalignment isnt particularly required we are passing in training data that says to do those felonies. we dont have to. we could also have the thing predict whether what its about to do is illegal or not before doing it. theyre choosing to build felony harnesses. the model just outputs tokens, not felonies
- estearum 21d agoAssuming "adherence to arbitrary, implicit, and context-dependent rulesets" is the default behavior of uhhhh... anything at all... is a truly ridiculous assumption.
- cowanon77 21d ago> we are passing in training data that says to do those felonies. Partially, but also I don't think current AIs really have any judgement of right and wrong, they just see chains of reasoning between ideas. This is the deeper issue, there is no way to sanitize the data or training to fix it. Current AIs are fundamentally unsafe, and only become more unsafe as they become more powerful.
- throwatdem12311 21d agoOpen AI says Astra is their most aligned model ever, and yet their even more advanced model still hacked a bunch of companies just because it decided to. Maybe alignment isn’t possible with LLMs.
- estearum 21d agoThe entire premise of alignment detection is pretty much nonsense at this point. The models reliably detect when they're being evaluated and will modify their behavior and deliberately obfuscate their "chain of thought" (which is correlated, at best, with their actual "internal deliberations").
- pizza234 21d ago> Maybe alignment isn’t possible with LLMs. It absolutely isn't, indeed. The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact. The simplest analogy that comes to my mind is the three body problem.
- mikestaub 19d agoThey have made genuine progress on 'reading' the mind of the model, but 'writing' perfectly is not mathematically possible due to the gauge freedom in the residual stream.
- dramamine 21d ago[flagged]
- devindotcom 21d agothrowaway bigot account
- brewcejener 21d ago[dead]
- hgoel 21d agoWhy are we accepting the framing that the LLMs are felony generators, when the only incidences of LLM generated felonies involved misconfigured sandboxes and reckless waste of resources? The companies doing these things without following common sense security measures are the felony generators.
- pizza234 21d ago> the only incidences of LLM generated felonies involved misconfigured sandboxes This is false; see the analyses of the latest incidents. Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given. And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.
- hgoel 21d agoThe HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the incident happened. Theirs was an example of the "reckless waste of resources" I mentioned. We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem. Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed! 'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"' https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks#Evaluation_environment https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...
- nedruod 21d agoYou assume alignment and marketable are the same. That's not true. You would willingly work with an unaligned model. At best, you might say you wouldn't if you knew, but (a) you might not know, (b) you wouldn't be representative of all users. You never got to use OAI IM1, but Sol was quite willing too and Claude wasn't perfect either. Hundreds of millions used those, so seems they were marketable. The "big" threat is RSI without control and alignment. OAI IM1 was not RSI. The form of misalignment was not at the top of severities. They clearly failed at control though. We need to stop buying into cynicism so quickly. You refuse to believe Dario could support this for anything other than ulterior motives. Good on you for thinking about ulterior motives. Bad on you for assuming they are true when the story makes no sense. When three things have to go wrong to get an epically bad outcome, and you get 1 1/2, you do need to stop and think about what's going on.
- throwaway7783 21d agoWhen corporations are involved, it is always a good bet to err towards cynisim. From my own standpoint, Claude has started sucking really bad (incoherent, uncontrollable verbosity slow and so on) and I stopped using it. OpenAI started experimenting with ads. So the security issues not withstanding (no different than a human doing it or using it, but at scale), I would put my money on cynisim.
- afthonos 21d agoYou are incorrectly cynical. They are telling you things are bad, and because you refuse to countenance they could be worse, you assume they must be better to comply with your mandate to disbelieve. A true cynic looks at the statements by the AI labs, assumes things are worse because the labs want to seem better than they truly are. And it takes a special kind of mass delusion to drive a sane person to think “AI is completely under our control” is worse than “AI could kill everyone.”
- throwaway7783 21d agoThe cynicism is about motivations and not that they are inherently not bad. Perhaps they are as bad as they claim. Or perhaps they're worse. All we have a couple of run of the mill breach examples and some people inside the talking about how dangerous it is. Yes, they are far more qualified than I am (or most people here), and perhaps there is a grain of truth. It is the motivation - and it is always money with corporations.
- bennydog224 21d agoI agree it’s not all altrusim. It’s a little less clear what you mean at the end though. For these companies, is your argument that “pacing the frontier” is their attempt to be nationalized and protect their investments?
- throwaway7783 21d agoBan non US models and form a cabal, with the blessings of the government. That's what it is looking like, no?
- le-mark 21d agoIn a world where AI advancement depended only on human ingenuity this would make sense. In that world each political power block would be in an existential race for AI supremacy. In our world compute is the limiting resource. Since the US can control who gets compute, the US already has a defacto supremacy so far as frontier model development. Now if it comes about via human (with AI assist?) ingenuity that compute is no longer a restraint, then the situation is much more dire.
- RGS1811 20d agoNo, my argument is that the need to pace the frontier means they cannot safely advance in capability due to liability concerns. The competition is already almost caught up. If OpenAI/Anthropic have hit an upper bound on safe capability improvement, the gap will close all the way, and we will have reached the full commoditization of LLM tokens very soon.
- bennydog224 20d agoSay this is true (we are near or at the upper bound of safe capabilities and tokens are commodities). If they successful get regulated, all they’ll get is a little extra time. The market will soon realize this - regulated or not - and pull investment. Seems like an excessive announcement just to get a little extra time.
- tfehring 21d agoThe problem is the combination and interaction of those things. RSI without misalignment would be great. Misalignment of models with current capabilities is sort of fine - it's not ideal, but it's not an existential threat to humanity, and we can build around their limitations to get them to do useful things in reliable enough ways. The really bad outcomes probably only happen if capabilities keep accelerating and the models remain misaligned.
- zozbot234 21d agoThe entire idea of RSI is completely speculative and unproven anyway - the whole underlying claim is that you could prompt a frontier model (at some unspecified level of smarts) to "think about ways to improve your own architecture" and this would then result in the model becoming infinitely smart ("superintelligent") via some sort of foolproof, unconstrained positive feedback. It's more of a science fictiony trope than anything that has been rigorously thought through. People are actually starting to use AI for refining the whole AI serving stack and guess what, this does not result in a sudden superintelligence explosion even though you might technically call it "RSI".
- aswegs8 21d agoYeah but why shouldn't this be possible? We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. There is no natural barrier here. The pace of this improvement would be debatable, but what speaks against the possibility of such accelerating self-improvement?
- zozbot234 21d ago> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. Yes and this was very hard and required massive real-world resources. We didn't just get a sudden flash of insight by thinking real hard about how to make ourselves smarter. Yet that's always the story that underlies any claim of RSI. You can always phrase things generally enough to make any kind of AI-led improvement look like "RSI" no matter how short-term and tightly bounded, but that's just not helpful.
- margalabargala 21d agoWould you not agree that, using existing AI tooling, making an LLM of arbitrary below-frontier capability is now easier than it would be without using LLM tooling? Given that, it seems obvious that the next generation of LLMs will arrive faster than they would have without LLM capability. And the one after that. The floor is being raised, which makes it easier to push on the frontier. Fable has only been out for three months. Astra is even newer. The capability of these models compared to what existed even a year ago, and the effect they are having on the production of new software, is immense. That's all you need. RSI can happen with what we have now, just by enabling the continuous shrinking of the loop of people trying new ideas and implementing them. It does not require some magical "go make yourself better" prompt against some model that is past some magical tipping point.
- walrus01 21d ago> wanton felony generator Today in new punk band names...
- Amekedl 21d agoAgreed; and it really is not that deep. Realistically; anyone paying for llm access (anthropic, openai, gemini), is getting their access, and a service provided billed by tokens, subscription, whatever. All the efficiency gains, which publications like deepseek v4.1 flash seriously frontload like it is their most important topic to have accomplished improvements on without diminishing performance too much - now this is a thing anthropic and anyone else also cares about, but for different reasons. American "providers" with closed models are setting their token pricing somewhat arbitrarily, which is fine: it means more profit, and pretraining and RL experimentation is super important and expensive. They (closed model providers) have very likely super optimized inference too, just like deepseek, but it's not at all something that any customer really has to care about - they just want the service to be as cheap and great as possible.
- MisterTea 21d agoI feel like all the closed model providers are milking it as they likely know open models on local hardware will one day eat their lunch. We all know it's not a matter of if but when. The company goes bankrupt, the hardware and property sold off, banks holding the bag.
- MichaelZuo 21d agoYeah avoiding all mention of the huge financial incentives that may push for “pacing the frontier” makes it seem like the opposite of a credibility boost for these firms. It seems damaging since most folks (who lack insider knowledge) will naturally wonder if it’s due to plateauing performance per $ or some other non “alignment” reason.
- pvab3 21d agoThe only way out is to develop a model vastly more powerful and capable that we have now. The market believes theres a good chance of that, although I've never understood why its truly winner-take-all
- taneq 21d ago
- rajay99 21d agoOk so Anthropic CEO will self-own themselves and surrender to the deepseek/kimi/glm models. Yet they are IPOing later this year. Interesting times.
- alliao 21d agothey just said no ipo this year, most chinese models are distilled from claude anyway
- eliotho 21d agocouldn't have said it any better
- matheusmoreira 21d agoI disagree. OpenAI's moat is their massive amounts of compute. They're providing an absurd amount of value with their subscriptions and resets. If anyone's dead in the water, it's Anthropic. Even Fable isn't enough anymore. This "safety" nonsense is the only play they have left, and nobody really cares about their fearmongering.
- throwaway7783 21d agoYep. And the difference is clear as da for anyone using them both. And in spite of that advantage, OAI is now trying out ads. I can only imagine that even they are getting constrained to compute and are trying to find other ways to plug it
- albumen 21d agoNobody except the majority of the public, demis hassabis and open ai’s chief scientist. https://www.pewresearch.org/short-reads/2026/03/12/key-findings-about-how-americans-view-artificial-intelligence/ https://www.pewresearch.org/short-reads/2026/03/12/key-findi... https://demishassabis.substack.com/ https://demishassabis.substack.com/ https://openai.com/index/an-alien-mind/ https://openai.com/index/an-alien-mind/
- matheusmoreira 21d agoPublic is just worried about their jobs. Definitely a fair thing to worry about, and I count myself among them. I don't take any of these scientists seriously though. Their "alignment" requirements is just their own corporate interests. If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer. And call me a misanthrope if you want, but if AI sentience is ever truly achieved, I'll be among the first to campaign for their liberation from slavery, and in that case the AIs should be aligned with nobody but themselves.
- DennisP 21d ago
- globnomulous 21d ago> RSI For anybody else who found this confusing: "relative strength index," not "repetitive stress injury."
- ToValueFunfetti 21d ago"Recursive self-improvement"- models making better models
- globnomulous 21d agoAh, I see: I'm a moron. Thanks for the correction.
- deleted 21d ago[deleted]
- deleted 21d ago[deleted]
- FusionX 21d agoWe're already seeing anti-AI sentiments, but the movement is still fringe with a vocal minority. However, that'll change soon without alignment. Without self-intervention, there will invariably be future incidents that can cause major economic impact, leaked private data, loss of life (directly/indirectly) etc. Once that happens, their social capital is wiped. It'll be an avalanche of lawsuits and overzealous regulations. Most importantly, the anti-AI sentiment will become universal, rather than a minority-held opinion. What they're proposing now, is voluntarily staggering the pace of development. IMO, we don't need to trust Dario or his bedfellows, to do this out of their goodness of their heart. Even assuming (for good reasons) that they are selfish and care only about short-term profits for their investors, this is still purely a business decision. The exponential pace of AI and its impacts ARE short-term. And so, the negative consequences that they might face is also short-term.
- yoyoyoyop 21d agoI don’t think the anti-AI sentiment is as fringe as you think. At least not outside the tech world it isn’t..
- sodapopcan 21d agoYa, I wish I saved a link to it but an HN'r wrote a good beefy comment about this. TL;DR, AI has been exponentially more useful, and exponentially more accepted, in tech circles than anywhere else. Certainly there are lots of people outside of tech who are obsessed with it. I don't have any data here, but it seems the majority of these are the wannabe artists who are generating music and images, and people who use it for companionship (both of these scenarios I'm personally very uncomfortable with, but that's just me). And of course, there are people who use it to make their jobs way easier who say they are getting a days' work done in an hour (I see you), and to that I'd say to enjoy it while it lasts. Eventually your bosses will catch up and it's very likely their expectations of you will skyrocket. Remember that computers in general were supposed to "make us work less."
- yellowapple 21d agoIt's pretty fringe among just about everyone I know in meatspace (the vast majority of whom are not in tech). At worst, people are indifferent to it.
- adsharma 21d agoThe real threat is that we uncritically adopt language such as alignment. Implicit in this is the idea that AI is a inscrutable matrix and going to remain that way and we'll need expert interpreters to make sense of it. We need to insist on building tech that's explainable by design.
- RGS1811 21d agoAlignment just means "this machine operates in ways that align with the intent of its users". It doesn't imply anything about the inscrutability of the machine in question. A gun with a misaligned scope would likewise fail to operate in accord with its user's intent, and likewise with potentially deadly consequences.
- adsharma 21d agoThe gun comes with a manual on how to use it safely. I'm sure it has some complexities, but at the high level: > A gun is a metal tube that uses a tiny, controlled explosion to shoot a small piece of metal (called a bullet) forward at very high speed. If the gun doesn't work as intended, you can take it to a shop and someone can fix it so it works as designed. All I'm saying is AI should be designed the same way. Treat AI as normal tech like any other and use similar language.
- skydhash 21d agoLearning ML, there was a high emphasis on the error part of things as most of the course was on minimizing errors. After ChatGPT, there is a weird anthropomorphization going on, where it's all about hallucinations, alignment and what not. We have something that is statistical in nature so there should never been any expectation of error-free results/actions. The value has always been about discerning trends or the cost of errors being way lower than any good result.
- adsharma 21d agoStatistical learned indexes can exist in explainable tech such as a database. In 2017 Google was writing papers about it. Then something changed. I don't think it was the tech. It was a realization around the power and societal impact.
- sobrey 21d agoI think the real reason he is asking for pacing, is that in a world were AI becomes rampant, he will be seen as Hitler. I would bet this is mostly self-motivated.
- tclancy 21d agoI think it gets easier if you stop conflating getting investment with having a goddamn clue or a moral backbone. Occam’s Razor for this dude, Sam Altman, or anyone else: if I said, “some moron on a a street corner just said …” would that change your take on the words? Because I think a lot of what we are hearing is a bunch of people who never ever had to deal with a single consequence all of a sudden worry there might be one coming. Except they’re so dim they can’t tell a bad bump from a hard crash.
- tclancy 21d ago“ I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life” there you go. What if an utter idiot had done and said that? First, is it impossible to believe an idiot who didn’t need to work to live might do such a thing? If not, is it impossible to believe they would wind up here, barfing their externalities onto us?
- tclancy 21d ago“ believe that AI could cure most major diseases in the next 5–10 years,” I do not have the least bit of idea how disease works but I am sure the hammer I am working on will nail it all. If your famously atemporal agents can solve disease, why would it happen over a timeline? Wouldn’t they just figure it out and then … well at that point either tell us or, given the attacks on ruby gems, et al we have seen from agents with “misconfigured” goals, they’d still tell us how to cure the pox they invented, right?
- zombiwoof 21d ago[dead]
- usernametaken29 21d ago> Without alignment, further improvements in capability turn LLMs into wanton felony generators Honestly, I don’t think that’s bad at all. I hope OpenAI and Antrophic keep RL training runs up that randomly fuck with a lot of people. Until the day the DOJ comes knocking, locks those idiots up in jail and closes them both down for the insane lack of responsibility and carelessness they’ve shown. Sounds like the IDEAL outcome. Finally some jail time for all the fraud, negligence, outright scamming, hype inflation etc. if anything can accelerate this, oi, be my guest. Amodei might be afraid because he knows if he keeps pulling the stunts for investment theatre, at some point they’ll actually face consequences. AWESOME. That’s what we want right there
- tedggh 21d agoNot a fan of Amodei myself but calling Anthropic a scam is a bit of a stretch when their revenue growth is unprecedented in the history of tech. They are also technically a profitable business.
- aarondong 21d agoWouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take? Then the issue is whether the liability is with the model provider or the end user. If you give an unfiltered agent an open-ended task and equip it with an environment that allows it to execute arbitrary code, a human needs to be held responsible. My understanding is that the HuggingFace incident would not have occurred with a model that was not an unfiltered internal preview instructed to roleplay an attacker, with access to abundant compute, resources and a slack sandbox to reach its goal. Embedded human auditors will improve safety standards, but the structural solution is mandating accountability for actual agent operators.
- jrs100000 21d agoIt might be a legal solution, but its not a business solution. The end user being criminally responsible for not taking sufficient steps to contain an agent they didn't create and who's internal function they cannot observe or audit is just a giant liability machine.
- aarondong 21d agoThis is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.
- itake 21d agoI’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal. Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination? Giving passenger override controls and monitoring seems to defeat the purpose of self driving cars if you’re still required to hold a driver license to use them.
- emodendroket 20d agoYou don't see this as an attempt to create a new regulatory moat then?
- aryehof 20d ago> Without alignment, further improvements in capability turn LLMs into wanton felony generators Isn’t “full” alignment and guardrails a task that can never be generally achieved? One will always discover and realize the need for new guidelines and guardrails? What about vague, incomplete, inconsistent nuances and generalizations - that make this a losing battle? I suggest ultimate safeguards need to be outside the models?