4 ms·
In Jan, when deepseek launched, Dario Amodei had to disclose they spent about $10M to train the last generation of models (his arguments was deepseek was on the
by ankit219 1y ago
In Jan, when deepseek launched, Dario Amodei had to disclose they spent about $10M to train the last generation of models (his arguments was deepseek was on the curve, not breaking it).
They earned $250M in May based on ARR, and about $400M in july. Model training is going to be amortized over multiple years anyway. I am not privy to how much they spent, not going to comment on that. GM was public news, and hence I got that.
Re Zitron's analysis, I don't find them to be reliable or compelling.
- kgwgk 1y ago> Model training is going to be amortized over multiple years anyway. Claude 4 launch was not even fifteen months after the launch of Claude 3 (which is discontinued). The “multiple” is 1.2 - I wouldn’t call that “multiple years”.
- singron 1y agoIt doesn't make sense to amortize model training over multiple years since they train multiple models per year (e.g. Claude 3.5, 3.7, and 4 were released within 12 months). Or you can, but then you have to overlap amortization schedules multiple times over. E.g. if they amortized over 24 months, then they would still be amortizing Claude 2.1, 3, 3.5, 3.7, 4, and 4.1.