3 ms·
Language models have always had an issue with negatives. A negative like do “not” xyz is just not encoded the same as spelling out what you want vs what you do
by sroussey 20d ago
Language models have always had an issue with negatives.
A negative like do “not” xyz is just not encoded the same as spelling out what you want vs what you don’t want.
Harder to write though.
- sigmoid10 20d agoI would say in this case abliteration is the likely culprit. To uncensor a model this way, you literally deactivate the parts that would enact refusals. As in things it was told not to do. But the real process is more like brain surgery performed by a alchemist according to an ancient religious book where noone involved really understands what is actually happening in the model.
- AndyNemmity 20d agoExactly, I wrote a blog post in what feels like a long time ago on this topic. https://vexjoy.com/posts/positive-framing-agents-skills/ https://vexjoy.com/posts/positive-framing-agents-skills/
- PotatoPrime 20d agoInteresting read, thanks for re-sharing! I noticed your joy-check link 404's now... I tried poking around your /skills/ folder but didn't find it easily. Should you still have that available I'd love to check it out. edit: Found it if others are looking: https://github.com/notque/vexjoy-agent/blob/main/skills/code-quality/joy-check/SKILL.md https://github.com/notque/vexjoy-agent/blob/main/skills/code...
- AndyNemmity 20d agoOh, thanks for letting me know. I need to fix that.