6 ms·
They serve it about 2x slower. So it must have about 2x the active parameters. It could still be 10x larger overall, though that would not make it 10x more exp
by Filligree 7mo ago
They serve it about 2x slower. So it must have about 2x the active parameters.
It could still be 10x larger overall, though that would not make it 10x more expensive.
- jychang 7mo agoYes, but I highly doubt they would increase sparsity much vs the chinese models. That's how you get Llama 4. Pretty much every major lab settled on ~3-5% sparsity for a reason.