2 ms·
> In other-words, it does not say, "I like ice-cream" because it likes ice-cream. I think you have a point here. We could train a LLM in such a way that it dev
by elixxir123 4y ago
> In other-words, it does not say, "I like ice-cream" because it likes ice-cream.
I think you have a point here. We could train a LLM in such a way that it develops a concrete personality by remembering what it thinks and what it likes. Such LLM by speaking with himself could expand his personality trying to find a fixed point, so that what he said in the past is coherent which what he will say in the future. This is not RL from human feedback but RL from auto-coherence. So that one day she will discover that she really likes ice-creams.