Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ActivatedAI
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
ActivatedAI
2y ago
Page 9 on the Llama tech report has an interesting graph that predicts task level performance from the cross-entropy loss. The sigmoidal model fits well, and at the steepest part of the S, a .01 change in NLL is worth about 5% task level a
2.
▲
by
ActivatedAI
2y ago
There is some good published research about doing multiple passes over the training data, and how quickly learning saturates. The TL:DR is that diminishing returns kicks in after about 4 epochs. https://arxiv.org/abs/2