3 ms·
I sort of wonder whether having a private language-style can't be used to avoid random prompt injections and things of that sort. Maybe they've screwed up, but
by impossiblefork 1mo ago
I sort of wonder whether having a private language-style can't be used to avoid random prompt injections and things of that sort.
Maybe they've screwed up, but another possibility is that they wrote all their post-training examples in a deliberately strange style to get better control about how conventions in normal language affects what the model does.
- impossiblefork 1mo agoSorry, I seem to have messed up this comment, I didn't mean prompt injections, I meant leakage of human thought into the model. Basically the idea is that if I train a model that speaks like a human, it will follow human biases. If I train a model that talks like Claude, it will follow the bias of my Claude-style examples.