9 ms·
As tome mentioned we don’t quantize, all activations are FP16 And here are some independent benchmarks https://artificialanalysis.ai/models/llama-2-chat-70b ht
by bsima 3y ago
As tome mentioned we don’t quantize, all activations are FP16
And here are some independent benchmarks https://artificialanalysis.ai/models/llama-2-chat-70b https://artificialanalysis.ai/models/llama-2-chat-70b
- xvector 3y agoJesus Christ, these speeds with FP16? That is simply insane.
- throwawaymaths 3y agoAsk how much hardware is behind it.
- modeless 3y agoAll that matters is the cost. Their price is cheap, so the real question is whether they are subsidizing the cost to achieve that price or not.
- noir_lord 3y ago> All that matters is the cost. Not really, sustainability matters, if they are the only game in town, you want to know that game isn't going to end suddenly when their runway turns into a brick wall.
- modeless 3y agoCost, not price.
- throwawaymaths 3y agoThe point of asking how much hardware is to estimate the cost? (Both capital and operational, i.e. power)