3 ms·
> Rule of thumb is that you need ~20 tokens per parameter. That rule of thumb is wrong. The chinchilla paper has it anywhere between 1 and 100 tokens per param
by causalmodels 3y ago
> Rule of thumb is that you need ~20 tokens per parameter.
That rule of thumb is wrong. The chinchilla paper has it anywhere between 1 and 100 tokens per parameter.