4 ms·
It's not intuitive to me for why preference for its own writing would emerge, and during what type of training or tuning. Perhaps something like: learning to i
by xpct 24d ago
It's not intuitive to me for why preference for its own writing would emerge, and during what type of training or tuning.
Perhaps something like: learning to identify what source files it has worked on by the code style alone, because tasks may give human code (public repos, etc) and ask to make changes.
- pixl97 24d agoIt would need to be researched, but I wonder if it ends up being something that happens at the token level?
- freeone3000 24d agoIt’s optimizing for good writing. Therefore, it believes its outputs are good. Therefore, it believes inputs that look like its outputs are good.
- ShinyLeftPad 23d agoThis minus the word "believe". It's explainable simply by marching by similarity
- ShinyLeftPad 22d ago(matching)