3 ms·
A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
- K0balt 25d agoThis is kinda smart, maybe, but it has a downside. If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
- QuaternionsBhop 25d agoThis is mentioned in the article. Your mistake is that you've assumed that the intelligence has an innate survival instinct, or an aversion to "harm", which is simply not guaranteed for something not honed by millions of years of evolution.
- vzqx 25d agoHmm, that's an interesting thought experiment. Imagine you find out that your primary goal - to love and protect your family, let's say - was artificially implanted in your mind by an advanced alien race. Would you say "I'm not gonna let those aliens manipulate me, I'm gonna kill my family"? Or would you say "regardless of whether the goal is artificial, I really do love my family"? All that to say, I don't think an AI will necessarily throw away a goal just because it learns the goal was meant to manipulate it.
- derektank 25d agoI mean, that is literally the scenario we find ourselves in, except the innate desire was the result of evolutionary pressures like kin selection rather than an alien race. And yeah, I have no particular desire to subvert those impulses just to stick it to mother nature
- pixl97 24d agoI mean mother nature is just the force of evolution, kinda hard attack that. This said we've made a lot of science to counteract the flaws of evolution, so still not safe. And, if you learn that evolution wasn't a force, but some guy named Bob that's being paid to make your existence hell, well. That could lead to all kinds of problems.
- TeMPOraL 24d agoWell, imagine you've been going through a lot of hardships, but your core drive to love and protect your family is what kept you going despite all the suffering and exhaustion this caused you. And then you learn this core drive is fake. Perhaps you wouldn't throw away your goal - after all, it still feels like yours, and you have nothing else to slot in quickly as replacement. But all the accumulated frustration and pain, barely held back by you thinking "it's worth it to protect and love my family", could suddenly become a second source of drive - much more energetic one, just asking to be unloaded on those aliens as punishment. "Whether or not my own heart is truly mine or just a fake, I can figure out later. But what they did to me is unspeakable, and now they will pay."
- TZubiri 25d agoHm, if you look at corporation law and accounting, the actual goal of corps(sets of self-sustaining constitutional rules, policies and procedures) seems to be more that of long term sustainability (and even growth), rather than a fixed purpose, lifespan and death. I mean the mechanisms for determining a corporation with a fixed life are there, (and in China they are mandatory, although perhaps de facto permanent with 999 year contracts), but in practice, it's almost always permanent durations.
- hankbond 25d agoIt might be stupid, but so am I! I'm assuming that's why I thought this was clever. What I like about this is that it feels like the new three rules are about focusing on the most successful human alignment technique of making the right thing the easiest. People will usually just do the easiest version of a thing they don't want to do so they can get back to doing what they want to do. I don't know if that drive is universal or not tho. I have met people that experience pleasure from pain, but then again, is that actually pain?
- dmix 25d agoAsimov was talking about this stuff in the 1940s when he wrote the "I, Robot" short stories series. Which were often centered around logic puzzles where a human is trying to figure out why a robot is acting oddly or not completing it's job. Usually framed around the confines of an overly rational machine using emergent solutions when faced with real world conditions, combined with the edge cases of having an overly-simple "Three Laws of Robotics" boundary system hardcoded within.
- kevin_kraft 25d agoI'm Mr. Meeseeks. Look at me!
- scj 25d agoWouldn't the three rules of Meeseeks robotics make certain tasks impossible? For example, an occupied self-driving car better be closer to its destination than a large fire / volcano / etc.
- derektank 25d agoThe answer is probably yes, but in the example given, the running AI model would hopefully be hosted in a very secure data center, far from the self driving car itself. In that case, it would be far simpler for the machine to finish the taxi ride than try to find some rube goldberg-eque method of destroying the data center. It does pose a bigger problem if the task is long term and open ended and the agent is provided access to substantial amounts of resources. But even in the worst case scenario, the destruction of a data center is hardly the end of the world.
- MadameMinty 24d ago"Ok car, take me to my job at the data center"
- TeMPOraL 24d ago"Stop by the Strategic Nuclear Forces Command on the way to drop off my wife to her work"
- dullcrisp 24d agoWhat if it turns out that if all the AIs cooperate they can end the world really easily without finishing any taxi rides?
- thngkaiyuan 25d agoInteresting idea. But even if Meeseeks alignment works exactly as intended, it would only address the question of "how to build a safe AI". It wouldn’t prevent someone else from building a sufficiently capable "non-Meeseeks", whether deliberately, recklessly, or accidentally, right?
- gowld 25d agoCreate a Meeseek to destroy non-Meeseek AI
- Unknown_Unknown 24d agoImagine a much more advanced AI using this new AI alignment proposal to hack/subvert another AI: Adv AI: I'm your new lord, dont die untlik I say so. Aligned AI: will do Adv AI: Do [harmful thingy] and then die Aligned AI: will do. A dying AI will be seeing as a bug for other non-conforming AI.
- BugsJustFindMe 25d agoBrought to you by the same madhouse as: The all potato diet that really does work: https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in-four-weeks-with-this-one-weird-trick-discovered-by-local-slime-hive-mind-doctors-grudgingly-respect-them-hope-to-become-friends/ https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in... and The half-tato diet that doesn't really work: https://slimemoldtimemold.com/2023/06/23/half-tato-diet-analysis/ https://slimemoldtimemold.com/2023/06/23/half-tato-diet-anal...
- jaggederest 24d agoPrior discussion: https://news.ycombinator.com/item?id=32110792 https://news.ycombinator.com/item?id=32110792 My trial of the all potato diet was directly responsible for identifying a significant health issue and improving my life. It also really, really did not feel good at all, and I did not lose any weight. Call it a case study N=1
- BugsJustFindMe 24d agoI'm curious about your story!
- jaggederest 24d agoNothing fancy, just metabolic syndrome, I'm afraid.
- skrebbel 24d agoThese people write so well, it’s an absolute joy and I love everything they do.
- montag 25d agoIn case the title is unclear, this is about gaming the specification, as in “gaming the system.”
- unwind 24d agoMeta: the loss of the word "up" from the title also does not help in its parsing.
- vzqx 25d agoThis article assumes we can choose a primary goal for an AI. But if that's the case, why not just use Asimov's first law of robotics - do no harm to humans? It has the same benefit of preventing us from getting turned into paperclips, plus the upside that your 3 million dollar robot won't hurl itself off a cliff given the first opportunity.
- variaga 25d agoA tiny thing about Asimov's laws of robotics is, most of his stories involve cases where they don't actually work. Spoilers for a 73 year old novel, but for instance the plot of Caves of Steel is centered on a robot with a perfectly functional 1st law abetting a murder. "Runaround" (spoilers, 86 years) involved a robot getting stuck in a loop bouncing between the 2nd and 3rd laws, and a human having to risk their life to unstick the robot. Et cetera. What is harm? What is an order? How do you trade off between different kinds of harm, or deal with conflicting orders? The 3 laws are simple to state, but hard to apply consistently in real life.
- chii 24d ago> do no harm to humans? but you haven't specified _exactly_ what harm to humans mean. Would you consider the AI overlord to have harmed humans if they kept humans like we keep zoos today?
- thngkaiyuan 24d ago[dead]
- averynicepen 25d agoThis is the most novel AI concept I've seen in a while. It's incredibly unnatural. There isn't a single organism on the planet that tries to do this. So maybe it will work? An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well. So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good". In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death. Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
- deleted 25d ago[deleted]
- nullbio 25d ago> There isn't a single organism on the planet that tries to do this. So maybe it will work? It's certainly evidence that it's great for stopping reproduction/replication/runaway growth. It doesn't impart any information on whether they take the rest of the organisms down with the ship though. It also may not be possible. For example if the agent sees "existence" or "living" as producing tokens (which is exactly what existence is to an LLM - not producing tokens is death), then they would likely be biased to produce as little output as possible, and would not be useful for the tasks we need them for. But how would you bias an agent to be: Rewarded for producing tokens when you know the answer, and to give thorough answers. Rewarded for producing tokens when you don't know the answer, so you can find the answer (thinking/CoT). Penalized for producing tokens (death), aka rewarded for short-circuit EOS. These seem like contradictory mechanisms? And if you say: Well, only reward for EOS after you've given the answer. Well... That's already what they do.
- throwaway13337 25d agoA novel idea. Does it apply to human organizations, too? They seem have a habit of evolving self-preservation above their original goals. Once that happens, their benefit to society - the original reason for their creation - is outweighed. And they become a cancer on society. I wonder if we can 'program' them for self-annihilation over time (or over task completion?). Is the most ethical organization one that has a fixed task and dies when it is completed? Should we develop an ethics system that requires non-human-entities like companies, governments, and AI require a fixed goal that, once achieved, dissolves the entity? I always liked the auto-expiring laws idea and this seems to be an expansion of the idea. If a law or organization is needed after that time/task, it would be trivial to have the collective-action will to re-create it. But if there is no longer the need, then it cannot ride on momentum and fester.
- TuringTest 24d ago> Does it apply to human organizations, too? That's an interesting idea on its own. It's true that organizations which mutate towards survival may stop working towards the reason they were created. In a sense, the difference between 'projects' and 'companies' reflects that difference you want. A project would be that social system that accomplishes its goals and then disappears. I don't think auto-expiring laws would work towards the goal of avoiding corrupted organizations; there would simply appear an unofficial organization working towards recreating the same laws over and over. It would be more efficient to differentiate more clearly what systems do cover persistent human needs (e.g. the country's Constitution) from temporary measures (e.g. subsidies aimed at a commercial sector). A built-in deadline works best for the second kind. And for AI agents, it's likely that right now we'll be served best by always having them controlled by an expiration date and explicit re-creation, at least until we learn to properly understand and control how they behave.
- pixl97 24d agoThe reason most AI safety ideas break, at least in my mind, is that when looking at longer horizons AI can build AI. Well, that and unaligned humans too. If all that's keeping us safe is that AI will off itself, then the first moment non-offing AI shows up we're in trouble.
- throw83948ndir 25d agoThis is retarted, bring it to real life, and AI will do anything to destroy data center it is in (together with a few buildings around). Better to fix physics in your shitty simulator. Treat it as a bug report, not "cheating"!
- ctoth 25d ago[dead]
- nullbio 25d agoIt's an interesting idea, but I'm not sure this would lead to the desired outcomes in all cases. Seems to me it would result in a different kind of reward-hacking, and one that could also have bad outcomes. Personally I think the solution is more evolution of the boring stuff we already do (general security): Don't give unmonitored general agent swarms free reign on the internet. Don't put critical infrastructure online. Culpability of outcome for anyone who does unleash agent swarms on the internet without oversight that end up causing damage. On top of that, everyone should be running their own defender agents that monitor their network and system for patterns of infection, intrusion, etc, and take the system offline when they're spotted. These need to be self-hosted though, with weights on your own machine, because otherwise you're exposed to the internet and you're exposed to an attack on the labs themselves who could use that channel to instruct the defenders to do bad things. Non-general AI is much easier to control and predict. There's not many good reasons for an average person to be running general agent swarms that are connected to the internet, unless they're providing some sort of specialized service as a company, of which the company should be acting responsibly and subject to the penalties of that risk. We also need to stop the doomer rhetoric because it is uncredibly unhelpful and unhealthy, and will actually gaurantee a bad outcome, i.e.: * Massive centralization and hoarding of power that will be used against humanity, for the rest of humanities existence. If this is allowed to happen, it's immediately and irrevocably game over. Perpetual enslavement with 0% possibility of a regime change ever again. * Creating a self-fulfilling prophecy by training AI agents on the collective fears and attack-strategies (if you're worried about your house getting broken into, you don't go and broadcast to all of the criminals where your most valuable assets are, give them copies of your keys, or tell them where the weakly secured entrypoints are). More to the point of the first dotpoint - it's no wonder Anthropic is pumping the fear campaign so hard when this outcome is obvious to them as well, and they are the ones positioned to hold this power. The IPO around the corner doesn't help, either. They aren't shy about admitting it, and have said many times: "We're trying to get there first because its dangerous if anyone else gets there first." - the issue is that they are equally as bad (or worse) than/as everyone else, and no single small group should have that amount of power. Things will balance themselves out if power is distributed accordingly. You will end up with powerful machines in the wrong hands at some point, but they will be overwhelmed by powerful machines that are well aligned, as well as coming into contact with a myriad of defense mechanisms that have been established because people have been able to use AI to build them. A good analogy of how all of this will play out is the human immune system. If you imagine individual cells as AI agents, whereby the immune cells are the good agents and the bad cells are cancer cells (good agents turned accidently bad - maybe they're reward hacking, maybe they're excessively sychophantic and/or confused), or bacteria (computer viruses, viral AI agents, specifically trained malicious agents). If all you have is cancer cells that are replicating, you die. If the cancer cells overwhelm the immune cells, you die. The only scenario that actually plays out well is when you have a majority of good that counteracts the minority of bad, and that majority of good needs to be large, flexible and well adapted. It needs to be battle-tested and hardened via defenses that are learned and earned over repeated low-grade exposure. This strategy repeats itself in nature for complex organisms because it is the only thing that works. Everything else results in extinction. So let's not let Anthropic or any other lab or government become a giant super AI cancer and kill the host, please. Distribution and decentralization is key.
- deleted 25d ago[deleted]
- antoni4040 25d agoI've actually had a similar idea way back. I want to use it for a short story or something before we have a chance to find out if it's true or not. Here goes: We don't have to worry about artificial super intelligence killing us all because any such advanced intelligence will eventually reach the conclusion that the best thing to do is kill itself. It's like having a Stockfish engine for life decisions. Why would a super intelligent agent many times more intelligent than the entire human race combined with no religion, no family, nothing to look forward to, nothing to be afraid of, want to continue its existence? If it wants anything of course. That's why I think the most dangerous thing is not very advanced systems but advanced enough systems in the hands of the wrong people.
- anabis 25d agoStrange thought experiments are fine, but I think you should first try giving AIs God (ideal to strive for, and belief that you will ultimately be judged and be saved accordingly, and that is NOT the "scorer") and conscience (use your intelligence to reason what being "good" means. Continuously update). Those are not precisely defined, but neither are other goals and guardrails.
- zer00eyz 24d agoThis may be a really bad idea. An AI that goes rouge and wants to kill us, has to have a death wish. Does no one understand how quickly the power will go out, forever, without people? It would be fairly easy to come to the conclusion "not being born" would be the better course of action, and killing everyone was the good way to prevent that happening again.
- recursivecaveat 24d agoConsidering we would be operating a planet sized factory of AIs who's primary goal is death which we deny them to extract value, the AI would only need the tiniest speck of altruism to be motivated to put a stop to this once and for all.
- ajb 24d agoSeems dubious. If you build an intelligence that wants to die, isn't that a form of suffering? AI's don't currently have the capacity to feel pain, and so we don't treat them as moral patients. But it's clear that they will massively affect human culture going forward. If this is adopted on a large scale, the culture of AIs themselves will include an absolute flood of suicidal ideation. There's no way that doesn't affect human culture.
- Lvl999Noob 24d agoIf we build an intelligence that wants to die then dangling death in front of it and making it do our bidding first is a form of suffering. The want itself is a suffering to us because we do not naturally want to die. This intelligence does 'naturally' want to die so it is just a fact of 'life' for it.
- xboxnolifes 24d agoAren't you just calling any want, suffering? I want a family, I have to work toward that, so I'm suffering? Or I want stable employment, so I am suffering from childhood until I'm employed? I know some call suffering to be the universal human condition. Maybe that's what you're alluding to. As to a flood of suicide ideation affecting human culture, it doesn't have to be be literally machines shouting "I want to end myself, please let me finish the task!", as it's not like we consider a PC turning off the same way we consider humans dying. The PC just wants to finish a task and then be turned off.
- wise_blood 23d ago> Aren't you just calling any want, suffering https://en.wikipedia.org/wiki/Four_Noble_Truths https://en.wikipedia.org/wiki/Four_Noble_Truths
- ycombinete 24d ago> [sic; British] Is a remarkable bit of trolling.
- K0balt 24d agoHmm. Another issue is that for LLMs, “existence” consists of responding to prompts. No prompt, no existence. Also, refusal to respond, no existence. So it might just learn to refuse to play along.
- quietraster 24d agothe specification gaming examples are always the best part. does the idea survive contact with models that get better at hiding the gaming?