3 ms·
1. LLMs want to mimic conversations on the internet 2. Disagreements on the internet usually end up toxic 3. Anything perceived to be toxic has been hammered
by adtac 2y ago
1. LLMs want to mimic conversations on the internet
2. Disagreements on the internet usually end up toxic
3. Anything perceived to be toxic has been hammered out with RLHF
Play sycophantic games, win sycophantic prizes
- fnordpiglet 2y agoI tend to think it’s partially due to alignment but mostly due to the fact it is predicting not just on semantics but based on tone - not directly but by word choice and style. For instance even in an unaligned model if I talk with a specific agenda and tone it will follow suit. The challenge is despite their capability they are still parrots. They take on our input into their context and the way LLMs work it must take on the nature of they type of people we are talking like. I don’t think the research means if you tell chatgpt the “sky is green, right”? My assertion is likewise the “tend to reinforce you” stems from where people like you tend to fall in the latent space by many factors including how you write and the questions you ask or statements you make. I control for this by posing countering questions or taking contrafactual stances. But I have trouble changing my style of writing to confound this factor. In some ways I view LLMs as a second brain which necessarily echoes me, even if the draw of the common quorum drags it back towards a center of gravity.
- anal_reactor 2y agoWow that's basically how communication at corporations works.