4 ms·
This only decreases memory cost of input context window, not the memory cost to load and run the models.
by ComputerGuru 2y ago
This only decreases memory cost of input context window, not the memory cost to load and run the models.
- freehorse 2y agoContext window requires ram too.
- ynniv 2y ago"Only"
- keyle 2y agoI agree with you though, the title is misleading.
- BoorishBears 2y agoTitle is perfect. Their typical audience probably understands "memory" better than "context window", but then if you've actually deployed these systems it's not difficult to go the other way, from "memory" to "context window" since the context window specifically is known to take additional VRAM over the model itself
- solarkraft 2y agoAnd that’s what matters the most! To me, at small model sizes (1-8B), anyway. A few thousans tokens already bog my RAM down quite a lot and I’d love to have more - I’d go as far as saying that context greatly determines LLM capability at this point.
- danielbln 2y agoYes, pretraining and post-training is nice and important, but in-context learning turns LLMs from toys into tools.
- erikmalkavian 2y ago[dead]