3 ms·
You could say this explicitly shows how little it has improved.
by mattigames 2y ago
You could say this explicitly shows how little it has improved.
- eru 2y agoI don't think so; it only shows a lack of polish: you can avoid one word repetition like this with some fairly straightforward software engineering (by tweaking the sampling from the networks probability distribution over tokens), it has not much to do with progress or lack thereof in LLMs more generally. However, I am somewhat surprised that whatever they are doing to avoid repetition seems to be crafted again for each model (or at least for 4.5) instead of using the same system that successfully avoids repetition for their more established models. To say things another way: if 4.5 had successfully avoided repetition, I wouldn't have taken it as a sign of progress, so I'm not taking the opposite as a big sign of lack of progress.