7 ms·
If anyone is interested in running this at home, please follow the llama-int8 project [1]. LLM.int8() is a recent development allowing LLMs to run in half the m
by v64 4y ago
If anyone is interested in running this at home, please follow the llama-int8 project [1]. LLM.int8() is a recent development allowing LLMs to run in half the memory without loss of performance [2]. Note that at the end of [2]'s abstract, the authors state "This result makes such models much more accessible, for example making it possible to use OPT-175B/BLOOM on a single server with consumer GPUs. We open-source our software." I'm very thankful we have researchers like this further democratizing access to this data and prying it out of the hands of the gatekeepers who wish to monetize it.
[1] https://github.com/tloen/llama-int8 https://github.com/tloen/llama-int8
[2] https://arxiv.org/abs/2208.07339 https://arxiv.org/abs/2208.07339
- rnosov 4y agoHmmm, the Github repo suggests that you might be able to run the 65B model on a single A100 80gb card. At the moment, the spot price on Google cloud for this card is $1.25/hour which makes it not so crazy expensive...
- nabla9 4y ago$1.25/hour is roughly a year of GPU time until it exceeds the price of A100 80GB card.
- metadat 4y agoI think OP meant that $1.25/hr makes this accessible for people try it out themselves cost effectively, without having to spend thousands or tens of thousands up front to obtain a capable hardware rig. Obviously $1.25/hr 24/7 does add up quickly, after one month the bill would come to $900.
- deleted 4y ago[deleted]
- downvotetruth 4y agoEagerly awaiting the int8 vs 4 benchmarks. Also, it can run on CPU https://github.com/markasoftware/llama-cpu https://github.com/markasoftware/llama-cpu So, an int8 patch could allow the 65B to run on a standard 128 GB setup assuming the 65B model's cache bursts fit, which if I were to speculate is why the released models stop @ 65B & meta likely already has larger unreleased internal ones.
- v64 4y agoearly int4 experiments seem to indicate it's possible but you do lose performance, see this thread https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_is_it_possible_to_run_metas_llama_65b_model_on/ https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_... edit: to clarify, it may be possible to get this loss back and there is reason to be optimistic
- CuriouslyC 4y agoProbably the best method is to just train it on int4 in the first place. Fine tuning after quantization would definitely help though.
- sp332 4y agoIsn't that backwards? You need fairly good resolution during training or your gradients will be pointing all over the place. Once you've found a good minimum point, moving a little away from it with reduced precision is probably OK.
- brookst 4y agoI have no idea what the right answer is, but I think the argument for int4 training is that the loss measurements would take the lower resolution of the model as a whole into account. Is it better to have billions of high resolution parameters and quantize them at the end, or to train low resolution parameters where the training algorithms see the lower resolution? It’s beyond me, but I’d love to know.
- swyx 4y agowhy is it that these models tend to be released as float16 and converting to int8 is left to the reader? is there something special about training that defaults you to float16?
- sillysaurusx 4y agoThey were trained in fp16, and researchers tend to release whatever format they trained. It’s hard enough to do a large release that it’s best not to try to have too many goals, for the same reason most software projects try not to do too much lest their schedule slip. Still, I’m a little sad they didn’t release the optimizer weights. It would’ve given us so much valuable info about the dataset, among other benefits.
- dspillett 4y agoPrecision, aiming those names refer to standard binary numeric types. IEEE754 16-bit floats carry 11 significant digits with absolute precision so by coverting to 8-bit integers you lose some of that. Depending on the distribution of the values in those floats you could be loosing a lot more detail then this would imply, which is the reason we use floating point numbers for anything in the first place (rather than using an int16 where you have greater precision at you maximum scale but much less at lower scales). So if the model is computed using float16s, distribute as-is and let the end user choose to user it like that or compromise for faster processing of there system can deal with many billions of int8s more effectively.
- dspillett 4y ago("aiming" should have been "assuming" in that second word – noticed far too late to correct, I really should stop using my phone's slide keyboard, either it or I or both are getting far less reliable)
- charcircuit 4y agoQuantization and other optimizations are more for productionizing models. You start with something accurate and then you start making tradeoffs to get the inference time to fit into your compute, memory, and time budgets.
- nextaccountic 4y agoIf the model weights are stored as int8, does this mean that the floating point capacity of the GPU is wasted? Or the int8 is converted to float in the GPU?
- woodson 4y agoWell, tensor cores support int8 instructions (at least from Turing onwards), so the hardware is being used, if that’s your concern.
- causality0 4y agoI feel like we're less than a decade away from being able to hook LLMs into gaming. How incredible would it be to have NPCs driven by LLM?
- ZunarJ5 4y agoThere are already several plugins for Unreal Engine. I am going to assume the same for Unity. https://www.youtube.com/watch?v=i-Aw32rgM-w&ab_channel=KellanMythen https://www.youtube.com/watch?v=i-Aw32rgM-w&ab_channel=Kella...
- visarga 4y agoWe'll soon have LLMs in operating systems, LLMs in browsers and you are right, probably also in games. LLMs will be the platform on which we build almost everything.
- jesusofnazarath 4y ago[dead]
- SloopJon 4y agoThere was an Ask HN post about that idea a couple of months ago: https://news.ycombinator.com/item?id=34478503 https://news.ycombinator.com/item?id=34478503 I have long wished for less linear stories in video games, where branching narrative (a la Choose Your Own Adventure) is one possible way to give the player agency. The problem is, true branches are expensive, because you end up writing a bunch of content the player never experiences. I see a lot of potential, but it's going to take a different kind of craftsmanship, and likely many iterations, to realize something more than a novelty.
- causality0 4y agoI much prefer handcrafted stories and quests. Characters that respond dynamically to the story and the player's actions, however, is quite tantalizing.