4 ms·
> The number of folks with the technical chops to really push on the forefront of AI is just not that big. Utter bullshit. We are still scaling transformer arc
by nickysielicki 22d ago
> The number of folks with the technical chops to really push on the forefront of AI is just not that big.
Utter bullshit. We are still scaling transformer architectures initially introduced ten years ago. Ten years of phds trying to improve on what Google produced ten years ago, mostly failing.
What has actually changed and can explain the progress we’ve seen? Do you really think it’s just a collection of better and badder RL gyms? Get real!
The main thing driving progress in ai is massive amounts of capital investment and incremental improvements in hardware (mostly memory bandwidth/capacity) and computer networking (we got better at collectives). The AI labs, ironically, have nothing to do with it. They’re the vessel for capital. The people actually driving things forward mainly work at nvidia.
The reason the models are better today than they were 3 years ago is almost exclusively due to better hardware and infrastructure software. Not better data, not better model architecture. The evidence of this is pretty easy to feel: the reason opus 5 doesn’t feel much more capable than opus 4.6 did is because they run on the same hardware generation. The reason opus 4.6 felt much more capable than anything before it is because it coincided with the scale out of a new hardware generation.
- nickysielicki 22d agoThis also explains why Gemini is lagging, they don’t have an nvidia training cluster.