3 ms·
I would guess that a) they assume their users will have a lot of GPU ram, b) actual running costs will depend on what inference engine/framework you're using (f
by wsgeorge 3y ago
I would guess that a) they assume their users will have a lot of GPU ram, b) actual running costs will depend on what inference engine/framework you're using (for example, some GGUF quants are very cheap).