3 ms·
Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, t
by RC_ITR 1mo ago
Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.
https://alignment.anthropic.com/2026/teaching-claude-why/ https://alignment.anthropic.com/2026/teaching-claude-why/
- hexasquid 1mo agoThe AI is getting bad morals from listening to that dreadful rock and roll
- notpachet 1mo agoRelated reading: The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property. https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...
- pixl97 1mo agoI mean, you're not wrong, but by that logic we were done for even before we had digital computers.
- RC_ITR 1mo agoAnd isn't that the great lesson of AI? The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?
- folkrav 1mo agoOh come on. It's also trained on fiction work. Shall we refrain from posting sci-fi stories too, now that we're there, just in case the AI might want to try it out?
- RC_ITR 27d agoObviously we can and do whatever we want, but yeah, maybe being more optimistic generally would be a good thing for us.
- sodapopcan 1mo ago> The things we say publicly actually do matter Certainly > the post-modern descent into absurdity and nihilism has tangible negative consequences? You mean breaking AIs? Not much of a lesson.
- cedws 1mo agoSounds just like the fantastical nonsense that comes out of Lesswrong.
- pineaux 1mo agoPart of the epstein class, dont forget.
- RC_ITR 1mo agoDo you make the claim that AI is something more than a reflection of its training data? I'm curious what other things you would argue influences an LLM's behavior. I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.
- deleted 1mo ago[deleted]
- mcmcmc 1mo agoDo you think it’s a good idea to self censor because someone might scrape your comment and feed it to an AI?
- RC_ITR 27d agoI think being thoughtful and baseline positive about the way we predict the future in popular media is a good idea. Nihilism isn't good in and of itself, so I'm not sure what you're arguing society would lose if we actively chose to be less nihilistic.
- HarHarVeryFunny 1mo agoThey could filter what they train on if they wanted to - they just don't want to.
- theptip 1mo agoIf the alignment process cannot fix this then we are cooked. The least of our worries is discussions on this forum.