5 ms·
The message is so obviously meant to insult people from an AI, that I suspect someone found a way to plant it in the training material. Perhaps some kind of att
by jasfi 2y ago
The message is so obviously meant to insult people from an AI, that I suspect someone found a way to plant it in the training material. Perhaps some kind of attack on LLMs.
- moffkalast 2y agoGoogle's AI division has been on a roll in terms of bad PR lately. Just the other day Gemini was lecturing a cancer patient about sensitivity [0], and Exp was seemingly trained on unfiltered Claude data [1]. They definitely put a great deal of effort into filtering and curating their training sets, lmao (/s). [0] https://old.reddit.com/r/ClaudeAI/comments/1gq9vpx/saw_the_other_post_about_claude_being_an_amazing https://old.reddit.com/r/ClaudeAI/comments/1gq9vpx/saw_the_o... [1] https://old.reddit.com/r/LocalLLaMA/comments/1grahpc/gemini_exp_1114_now_ranks_joint_1_overall_on https://old.reddit.com/r/LocalLLaMA/comments/1grahpc/gemini_...
- helloplanets 2y agoAgreed, it's clearly a data poisoning attack. It's a pretty specific portion of the dataset the user is in after so many tokens have been sent back and forth. Could be some strange Unicode characters in there so it's snapped into the infected portion quicker, could be the hundredth time this user is doing some variation of this same chat to get the desired result, etc. It is weird that Gemini's filters wouldn't catch that reply as malicious, though.