3 ms·
The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end
by gdiamos 27d ago
The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end