2 ms·
I think that is working OK as long as token probability and correctness are related. If, in the extreme, there is something where all training data is wrong, no
by RandomLensman 3y ago
I think that is working OK as long as token probability and correctness are related. If, in the extreme, there is something where all training data is wrong, not sure there is a good way to do this. Maybe I am misunderstanding, though.
It might also need to be able to distinguish between Knightian uncertainties and probabilities when there is nothing to base things on.
- nonameiguess 3y agoWhat it needs is a hierarchy of evidence. This works almost unreasonably well right now because I guess we're lucky that more digitized text than not is largely true, or RLHF is just that effective, but at some point, I would think the learner has to understand that reading a chemistry textbook and reading Reddit have equal weight when it comes to learning how to construct syntactically well-formed sentences with human-intelligble semantic content, but don't have equal weight with respect to factual accuracy.