3 ms·
I think it is more likely that a smaller model (<400B) with similar intelligence gets developed long before the hardware to serve a 2.4T model gets cheaper than
by ak_t 2mo ago
I think it is more likely that a smaller model (<400B) with similar intelligence gets developed long before the hardware to serve a 2.4T model gets cheaper than 10k.
- edg5000 2mo agoAh, I hadn't thought of that. To what degree have we seen this already? What would you consider as the biggest jump in intelligence per weight?
- ak_t 2mo agoIt's been pretty consistent, the smaller models (hundreds of billions of params) usually catch up in 6 months or so to their frontier counterparts, at least on benchmarks.