5 ms·
Or, without the safety prompts, it outputs stuff that would be a PR nightmare. Like, if someone asked it to explain differing violent crime rates in America ba
by IncreasePosts 2y ago
Or, without the safety prompts, it outputs stuff that would be a PR nightmare.
Like, if someone asked it to explain differing violent crime rates in America based on race and one of the pathways the CoT takes is that black people are more murderous than white people. Even if the specific reasoning is abandoned later, it would still be ugly.
- bongodongobob 2y agoThis is what I think it is. I would assume that's the power of train of thought. Being able to go down the rabbit hole and then backtrack when an error or inconsistency is found. They might just not want people to see the "bad" paths it takes on the way.
- deleted 2y ago[deleted]
- decremental 2y agoThe real danger of an advanced artificial intelligence is that it will make conclusions that regular people understand but are inconvenient for the regime. The AI must be aligned so that it will maintain the lies that people are supposed to go along with.
- jasonlfunk 2y agoThis is 100% a factor. The internet has some pretty dark and nasty corners; therefore so does the model. Seeing it unfiltered would be a PR nightmare for OpenAI.
- quantified 2y agoI trust that Grok won't be limited by avoiding the dark and nasty corners.
- tim333 2y agoThat's an interesting point. I imagine even Grok will end up somewhat censored. Although maybe AIs will end up with a more sophisticated take on the problems than your average human.
- contravariant 2y agoCould be, but 'AI model says weird shit' has almost never stuck around unless it's public (which won't happen here), really common, or really blatantly wrong. And usually at least 2 of those three. For something usually hidden the first two don't really apply that well, and the last would have to be really blatant unless you want an article about "Model recovers from mistake" which is just not interesting. And in that scenario, it would have to mean the CoT contains something like blatant racism or just a general hatred of the human race. And if it turns out that the model is essentially 'evil' but clever enough to keep that hidden then I think we ought to know.
- bongodongobob 2y agoJust no. AI being racist is still a popular meme. "Because the programmers are white males blah blah".
- abenga 2y agoWhy can't it be, if it were (I'm not saying that it is, mind) trained on racist material?
- bongodongobob 2y agoThe problem is being kind of right (but not really) for the wrong reasons. Normies think it was told to be a certain way. While kind of true, they think of it more like Eliza.
- fragmede 2y agoIt's not racism, but from today, here's TechCrunch with: Hacker tricks ChatGPT into giving out detailed instructions for making homemade bombs https://techcrunch.com/2024/09/12/hacker-tricks-chatgpt-into-giving-out-detailed-instructions-for-making-homemade-bombs/ https://techcrunch.com/2024/09/12/hacker-tricks-chatgpt-into...
- greenchair 2y agoyes this is going to happen eventually.
- maroonblazer 2y agoUnlikely, given we have people running for high office in the U.S. saying similar things, and it has nearly zero impact on their likelihood to win the election.