3 ms·
This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the exp
by heresalexandria 2mo ago
This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the expanding generalized capabilities of the SoTA, which makes them effectively meaningless.
If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models are not?
To me, a better signal of capability would be similarly performing novel work at the same or better level, which they presently are not. I say this as someone who very much looks forward to open models being more capable, but to deny the gap is misguided hopeful hype.
- ChanderG 2mo agoWhy? Why is the premise that Fine-tuned models should be geared towards new discoveries? The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.
- heresalexandria 2mo agoI'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level. If the appropriate qualifiers were added to say "cheap model X does better at test Y than expensive model Z when we fine tune X to take Y test of existing knowledge" then it would be a more accurate statement, but naturally less impressive.
- echelon 2mo ago> I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level. What if you have to do the task a billion times? Which model will you choose?
- ozim 2mo agoMaybe because people who are target audience don’t need to have it spelled out like that? People who are not really into it, don’t care.
- antupis 2mo agoSpeed play also you can get much faster responses with 9b model.
- nine_k 2mo agoHuge SOTA models are like a floodlight. They elucidate a huge area at once. A fine-tuned small model is like a laser pointer. It only illuminates a tiny specific spot. But it can illuminate it as brightly as the huge floodlight, for a tiny fraction of cost.
- skybrian 2mo agoThis is about saving money by using the right tool for the job. If you have a system that does a lot of mundane, repetitive work, you don't need a frontier model to do it. It doesn't mean frontier models aren't good at harder tasks.
- heresalexandria 2mo agoThat's fair and I agree with this framing.
- iwontberude 2mo agoThat was always the framing. I don’t understand your pedantry.
- hahahaa 2mo agoDepends on use case. That email classifier for legal emails: cheaper at scale as a small tuned model. Let alone better for the planet. Frontier model may have done that tuning!