7 ms·
Show HN: I made a GPU VRAM calculator for transformer-based models
- cchance 3y agoVery nice, would be cool to have a little i next to each spot to explain what each thing is for newer users (batch size, etc)
- a2128 3y agoI noticed the default parameter count value is 1.418 billion but if you erase it you can't actually enter it back because you can't type a decimal point in the input area. Also, you can't enter parameter counts smaller than 1 billion
- sp332 3y agoIt works if you type the digits first and then insert the decimal point after.
- ilaksh 3y agoDoes this have an option for quantization levels? Don't think I saw it.
- ComputerGuru 3y agoI second the request for quantization, eg for exl2.
- furiousteabag 3y agoThere is no option to select quantized version yet. Will work on that!
- roseway4 3y agoWhile not as pretty (and mobile-friendly) as the original link, the calculators below support modeling LoRA-based training, alongside full finetuning. https://huggingface.co/spaces/Vokturz/can-it-run-llm https://huggingface.co/spaces/Vokturz/can-it-run-llm https://rahulschand.github.io/gpu_poor/ https://rahulschand.github.io/gpu_poor/
- ComputerGuru 3y agoThey seem to be broken when I try any HF ids besides what came preconfigured. e.g. just tried brucethemoose/Yi-34B-200K-DARE-merge-v5-3.1bpw-exl2-fiction or LoneStriker/shisa-7b-v1-3.0bpw-h6-exl2
- icelancer 3y agoSecond link hasn't been working for awhile.
- 3abiton 3y agoBeen looking for something like thos for a while! I googled a lot, and this link never popped up. I feel google search is regressing.
- deleted 3y ago[deleted]
- a_wild_dandan 3y agoAre people still rawdoggin' 16-bit models? I almost exclusively use 5-bit inference quants (or 8-bit natives like Yi-34b) on my MacBook Pro. Tiny accuracy loss, runs fast, and leave plenty of (V)RAM on the table. Mixtral 8x7 is my new daily driver, and only takes like 40GB to run! I wonder if I could run two of them talking to each other...
- rubatuga 3y agoPure 16bit is horrible for training, sorry.
- rdedev 3y agoDoesn't using bf16 alleviate the problem? At least I've had success training a Bert like model from scratch
- shikon7 3y agoI wonder about that too. With the small precision, parameter updates might be too small to have an effect (is it possible to use some sort of probabilistic update in that case?) Unfortunately, I haven’t found any resources describing the feasibility of full fp16 or bf16 training.
- rdedev 3y agoAh my bad. I am using mixed precision training in the my previous comment. You might find this paper interesting: https://arxiv.org/pdf/2010.06192.pdf https://arxiv.org/pdf/2010.06192.pdf
- furiousteabag 3y agoYou are correct, training sorely in fp16/bf16 can lead to imprecise weight updates or even gradients turning to zero. Because of that, mixed precision is used. In mixed precision training, we keep a copy of the weights in fp32 (master model) and the training loop looks like this: compute the output with the fp16 model, then the loss -> back-propagate the gradients in half-precision -> copy the gradients in fp32 precision -> do the update on the master model (in fp32 precision) -> copy the master model in the fp16 model. We also do loss scaling which means multiplying the output of the loss function by some scalar number before backprop (necessary in fp16 but not required in bf16). Check out the fastai docs for more details: https://docs.fast.ai/callback.fp16.html https://docs.fast.ai/callback.fp16.html
- samspenc 3y agoConsumer grade GPUs like NVidia's 3090 and 4090 max out at 24 GB VRAM, and those cost $1000-2000 each. You can get higher VRAM but need enterprise GPUs which are in the five figures, easily starting at $30K a pop. Per this calculator, for training, only gpt2-large and gpt2-medium would work with those two top-of-the-line GPUs. For inference it's certainly a bit better, only the Llama-2-70b-hf and Llama-2-13b-hf don't fit in that much VRAM, all the other models do.
- alexhutcheson 3y agoNvidia’s workstation cards are available with more RAM than the consumer cards, at a lower price than the datacenter cards. RTX 6000 Ada has 48 GB VRAM and retails for $6800, and RTX 5000 Ada has 32 GB VRAM and retails for $4000[1]. Very large models have to be distributed across multiple GPUs though, even if you’re using datacenter chips like H100s. [1] https://store.nvidia.com/en-us/nvidia-rtx/store/ https://store.nvidia.com/en-us/nvidia-rtx/store/
- slabity 3y agoOther than power consumption, is there any reason to prefer a single workstation card over multiple consumer cards then? A single $6800 RTX 6000 Ada with 48GB of VRAM vs 6x 7900XTX with a combined total of 144GB of VRAM honestly makes this seem like a no brainer to me.
- ttt3ts 3y agoYou have to pass the context between GPUs for large models that don't fit in VRAM. Often ends up slower. Also, tooling around AMD GPUs is still poor in comparison.
- alexhutcheson 3y agoYou can only fit 1-2 graphics cards in a “normal” ATX case (each card takes 2-3 “slots”). If you want 4 cards on one machine, you need a bigger/more expensive motherboard, case, PSU, etc. I haven’t personally seen anyone put 6 cards in a workstation.
- _giorgio_ 3y agoWhat is the usual way to do it inside the python file that defined the model?
- thatguysaguy 3y agoThis only lists first moments, but Adam stores estimates of first and second moments.
- furiousteabag 3y agoBy default, SGD w momentum is enabled as optimizer. You may try selecting Adam and it will list second moments as well.
- twayt 3y agoThis is actually pretty useful
- lgkk 3y agoOn mobile, iOS specifically in safari, your drop downs are hard to use. I’m not able to dismiss the keyboard. Is that an issue on my end?