3 ms·
405b even at low quants would have very low tokens generation speed, so even if you got the 192GB it would probably not be a good experience. I think 405b is th
by tarruda 2y ago
405b even at low quants would have very low tokens generation speed, so even if you got the 192GB it would probably not be a good experience. I think 405b is the kind of model that only makes sense to run in clusters of A100/H100.
IMO it is not worth it, 70b models at q8 are already pretty darn good, and 128gb is more than enough for those.
- a_wild_dandan 2y agoExactly! Have you tried the Phi models? To me, they indicate that we can get much more efficient models. In a few years, 70b on gold standard synthetic data + RL might run circles around SotA. It's such an exciting time to be alive.