4 ms·
Anthropic and AI alignment research isn't about making AI that are DnD-style "good alignment", but making AI that have outcomes that are aligned with the goals
by chc4 2y ago
Anthropic and AI alignment research isn't about making AI that are DnD-style "good alignment", but making AI that have outcomes that are aligned with the goals that the designers intended for them. The chatbot AI model and goals are not the same model and goals for a defense AI.
The goals for a chatbot assistant are to be useful, correct, and not insult people. The goals for a defense AI are to extract correct features, provide useful guidance, and not kill the wrong people. If you are working in defense you already have a belief that your work is morally correct: most of those justifications are either that your work will kill bad people more effectively, and so save friendly lives, or will pick who to kill most correctly, and so save innocent lives. Having an AI that is better aligned towards those goals are better.
You may disagree that working in defense is ever morally justified! But Palantir dont't share those beliefs, and want to do as good of a job as they can, and so want the most aligned AI model they can.
- danesparza 2y agoAnd what happens when the defense AI 'hallucinates' and suggests that somebody is a terrorist when they are not?
- rodgerd 2y agoSecurity podcasters will cheer you killing kids because you might have hit a few terrorists in the process.
- JumpCrisscross 2y ago> what happens when the defense AI 'hallucinates' and suggests that somebody is a terrorist when they are not? Collateral damage. Same thing that happens in any war when an analyst or soldier misreads the battlespace. War is hell. We won’t change that by making it pleasant. We can only avoid it by not going to war.
- protomolecule 2y ago[flagged]
- JumpCrisscross 2y ago> Acts like firebombing of Tokyo or bombing of Dresden or atomic bombing don't happen now We still raze cities and drop incendiaries. America hasn’t gone to war with a near-peer nonnuclear power like Japan since WWII. To the extent we were faced with the prospect in the Cold War, both we and the Soviets were committed to MAD, i.e. using nukes. (Do you think unilateral disarmament in the Cold War would have lead to peace?) There has been no militarily useful technology that was voluntarily abandoned. Just constrained. You can’t constrain a technology you don’t bother understanding.
- protomolecule 2y ago[flagged]
- JumpCrisscross 2y ago> during the firebombing of Tokyo the US murdered 100,000 civilians Are you arguing there was a war in which firebombing would have been useful but someone decided it was too mean? Since WWII we invented better high explosives and stand-off precision weapons. If there were a strategic case for firebombing in a future war, have no delusions: it will happen. (Last year, incendiary weapons were used in “ in the Gaza Strip, Lebanon, Ukraine, and Syria” [1].) [1] https://www.hrw.org/news/2024/11/07/incendiary-weapons-new-use-calls-immediate-action https://www.hrw.org/news/2024/11/07/incendiary-weapons-new-u...
- protomolecule 2y ago[flagged]
- JumpCrisscross 2y ago
- jsheard 2y agoGoing by recent events I think the convention is to drone strike them and their entire family anyway, and then tally them up as a confirmed dead terrorist. https://www.972mag.com/lavender-ai-israeli-army-gaza/ https://www.972mag.com/lavender-ai-israeli-army-gaza/
- sourcepluck 2y agoAhh, I hadn't seen this before posting. Thank you for providing a dose of reality, excellent link
- woadwarrior01 2y agoThat's been done over a decade ago with random forests. There's no need to apply anything as advanced as generative AI 'hallucinations' for such a trivial problem. /s https://www.theguardian.com/science/the-lay-scientist/2016/feb/18/has-a-rampaging-ai-algorithm-really-killed-thousands-in-pakistan https://www.theguardian.com/science/the-lay-scientist/2016/f...
- sourcepluck 2y agoThe cynic in me finds this quite naive - there are territories of the Earth where if you are a male of a certain age and you are killed in a drone strike you are automatically classified as a "military-aged male", i.e. a non-civilian, regardless of the existence of any other evidence. So what will happen if AI suggests someone is a terrorist when they are not? Well, in the worst scenario they'll be killed, and it'll be very close to what we have today, except somewhere in a private military database there might be an automatically generated record from an LLM ok-ing the target.
- protomolecule 2y ago>you already have a belief that your work is morally correct Or you don't care about morals. Or you are evil.
- er4hn 2y agoThey train us to drop fire on people but won't let us write "fuck" on the side of an airplane because it is obscene. (Col. Kurtz - Apocalypse Now) Which, when you unpack it, is even more interesting. If you do embrace the emotional aspect of war you end up with situations like the my lai massacre. Does AI have the ability to prevent war crimes while engaging in "legal" killings feels like an interesting philosophical question.
- Loughla 2y agoStopping war crimes does not require AI to be allowed to kill people. I don't understand that equation.