4 ms·
And why won’t the frontier models continue to become better? The open models are getting better but so are the frontier models. The frontier models might remain
by jostmey 1mo ago
And why won’t the frontier models continue to become better? The open models are getting better but so are the frontier models. The frontier models might remain in a constant race to remain ahead
- ForHackernews 1mo agoThey are running out of novel, clean training data and compute. There is probably a limit to how much improvement can be squeezed out of LLMs. Recent improvements have been more about orchestration and "reasoning" loops (i.e. iteratively feeding context back through the model).
- svachalek 1mo agoFor base models they really must be running out of new training materials. It's more about size and architecture. But they still seem to be getting big strides out of improving the post training. They keep dropping point upgrades in under 8 weeks lately, which is an insane pace for product release. At some point it's going to slow down but we're not close yet imo.
- HappyPanacea 1mo agoDiminishing returns on both intelligence and training, mostly
- sweetjuly 1mo agoI don't think it's a matter of "staying ahead"; the proprietary frontier models are better, but the trouble for them is that open weight models are good enough in increasingly many cases. This raises the floor on the frontier companies and cuts their total addressable market by commodifying the easier LLM tasks. This is really the central argument of tfa :)
- cman1444 1mo agoThe big question is if there are larger economic gains to be made from ever greater intelligence. It seems to me like that may be the case for only a select few hard problems, while the vast majority of tasks approach their economic ceiling asymptotically with intelligence.
- euroderf 1mo agoYou suggest that the market segment for frontier models is necessarily shrinking. But what if it's just a failure of imagination on the part of users ?
- hadlock 1mo agoI suspect due to model distillation, their need to stay on top for IPO value is more important than absolutely crushing their competition, who will simply harvest their output and distill it for training data on a ~3 month delay. Also they are probably hitting a wall on increase in intelligence vs training time.