5 ms·
I think if you start taking cost into account for tasks which models have saturated a long time ago, the bitter lesson stops being applicable. If your task is
by OtherShrezzing 1mo ago
I think if you start taking cost into account for tasks which models have saturated a long time ago, the bitter lesson stops being applicable.
If your task is "write a function to sort a list of 1mn entries alphabetically", a small coding model will exactly match the performance of a frontier model (it'll 1-shot it, with 100% accuracy). So the only differentiator is cost, and the frontier model will cost orders of magnitude more than the specialised model.
For most intelligence based tasks, you don't (and never have) needed the tool which "performs best at all tasks". You need the cheapest one which performs adequately for your immediate task.
This doesn't mean the bitter lesson is incorrect. At the frontier, it's still correct. It means that it's not applicable at all to lots of tasks.