3 ms·
I'd be interested to know if anyone has studied how overfitting translates into the domain of llm output: it's easy to understand when you're fitting a line tho
by version_five 3y ago
I'd be interested to know if anyone has studied how overfitting translates into the domain of llm output: it's easy to understand when you're fitting a line though data, or building a classifier, you overfit and your test set loss is higher than your training set loss, and this directly relates to worse performance of the model. For an llm picking probably next words, what's the analogy, and does overfitting make it "worse" even if a test set loss is higher?