3 ms·
I think people talking about a 100T GPT didn't mean a dense transformer but some sort of extreme Mixture-of-Experts which is much more amenable to low-resource
by airgapstopgap 3y ago
I think people talking about a 100T GPT didn't mean a dense transformer but some sort of extreme Mixture-of-Experts which is much more amenable to low-resource setups and complicates this discussion.
In any case, it's almost certainly not bigger than 1T, even if it's not a dense transformer (PaLM-2 is and makes do with 340B, but it isn't exactly on par).