4 ms·
In practice, it's fine to stick with "just" 8k or 16k or 32k. If you're working with data of over 128k tokens I'd personally not recommend using an open model a
by azeirah 2y ago
In practice, it's fine to stick with "just" 8k or 16k or 32k. If you're working with data of over 128k tokens I'd personally not recommend using an open model anyway unless you know what you're doing. The models are kinda there, but the hardware mostly isn't.
This is only realistic right now for people with those unified memory MacBook or for enthusiasts with Epyc servers or a very high end workstation built for inference.
Anything above that I don't consider "consumer" inference