4 ms·
> I think what they're getting at is that for a given unit of compute, this method achieves 125% performance. This is not what they're getting at; I explained
by dvt 7mo ago
> I think what they're getting at is that for a given unit of compute, this method achieves 125% performance.
This is not what they're getting at; I explained exactly what they're getting at. I mean, your equivalence of "loss" (what authors actually measured) and "performance" is just bizarre. We use benchmarks to measure performance, and the numbers there were like 1-5% better (apart from the GPQA-Diamond outlier).
Do people even read these papers?
- jszymborski 7mo ago> Do people even read these papers? Overwhelmingly, no. You may have mistaken this for a lab's reading group, but most people here just skim the README, maybe read the abstract or figures. Expecting them to do more is uh... a bit strange? But also you can forgive people for equating loss with performance, which are admittedly different but related ideas.