3 ms·
Demo on actual 4090 with flux schnell for next few hours: https://5jkdpo3rnipsem-3000.proxy.runpod.net/ https://5jkdpo3rnipsem-3000.proxy.runpod.net/ Its basic
by mesmertech 2y ago
Demo on actual 4090 with flux schnell for next few hours: https://5jkdpo3rnipsem-3000.proxy.runpod.net/ https://5jkdpo3rnipsem-3000.proxy.runpod.net/
Its basically H100 speeds with 4090, 4.80it/s. 1.1 sec for flux schenll(4 steps) and 5.5 seconds for flux dev(25 steps). Compared to normal speeds(comfyui fp8 with "--fast" optimization") which is 3 seconds for schnell and 11.5 seconds for dev
- yakorevivan 2y agoHey, can you share the inference code please? Thanks..
- superkuh 2y agohttps://github.com/mit-han-lab/nunchaku https://github.com/mit-han-lab/nunchaku
- oneshtein 2y agoCannot compile it locally on Fedora 40: nunchaku/third_party/spdlog/include/spdlog/common.h(144): error: namespace "std" has no member "function" using err_handler = std::function<void(const std::string &err_msg)>; ^
- mesmertech 2y agoYea its a pain, I'm trying to make an api endpoint for a website I own, and working on a docker image. This is what I have for now that "just" works: the conda always yes thing makes sure that you can just paste the script and it all works instead of having to press "y" for each install. Also if you don't feel like installing a wheel from random person on the internet, replace that step with "pip install -e ." as the repo suggests. I compiled that one with cuda 12.4 cause that was the part takes the most time and is what most often seems to be breaking. Also I'm not sure if this will work on Fedora, I tried this on a runpod machine with 4090(apparently it only works on few cards, 3090, 4090, a100 etc) with Cuda 12.4 on host machine and "runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04" this image as base. EDIT: using pastebin instead as HN doesn't seem to jive with code blocks: https://pastebin.com/zK1z0UdM https://pastebin.com/zK1z0UdM
- oneshtein 2y agoAlmost working: [2024-11-09 19:33:55.214] [info] Initializing QuantizedFluxModel [2024-11-09 19:33:55.359] [info] Loading weights from ~/.cache/huggingface/hub/models--mit-han-lab--svdquant-models/snapshots/d2a46e82a378ec70e3329a2219ac4331a444a999/svdq-int4-flux.1-schnell.safetensors [2024-11-09 19:34:01.432] [warning] Unable to pin memory: invalid argument [2024-11-09 19:34:02.143] [info] Done. terminate called after throwing an instance of 'CUDAError' what(): CUDA error: pointer does not correspond to a registered memory region (at /nunchaku/src/Serialization.cpp:32)
- mesmertech 2y agoprolly make sure your host machine cuda is also 12.4 and if not, update the other cuda versions I have on the pastebin to the one you have. I don't think it works with cuda 11.8 tho, remember trying it once but yea, can't help you outside of runpod, I haven't even tried this on my home PCs yet. for my usecase of serverless API, it seems to work
- bufferoverflow 2y agoDamn, it runs very fast.
- AzN1337c0d3r 2y agoIt's worth noting this is laptop 4090 GPU which is more like in the range of desktop 4070 performance.
- mesmertech 2y agoThis specific link I shared is the quant running on a 4090 I rented on runpod, I have no affiliation with the repo itself
- qeternity 2y agoThe compute differential between an H100 and a 4090 is not huge. The main single GPU benefits are larger memory (and thus memory bandwidth) and native fp8. But these matter less for diffusion models.
- mesmertech 2y agoThats what I thought as well, but FP8 is much faster on h100, like 2x-3x. You can check it/s here: https://github.com/aredden/flux-fp8-api https://github.com/aredden/flux-fp8-api Its why fal, replicate, pretty much all big diffusion api providers use h100 tldr; 4090 is max 3.51 it/s even with all the current optimizations. h100 is 11.5it/s with all optimizations, and even without its 6.1 it/s
- boroboro4 2y agoProviders use h100 because using 4090 in DCs is grey area, since Nvidia doesn't permit it. Paper discussing here is using 4 bit compute, which is 4x on 4090 in comparison with bf16 compute, while h100 doesn't have this at all (i.e. best you can get is 2x compute with fp8). So this paper will even out difference between those two to some extent. If to judge by theoretical numbers - H100 has 1979 TFLOPs fp8 compute, and 4090 has 1321 TOPS. Which puts it around ~65% of performance. Given the price of it ~$2K compared to H100s ~$30K this seems like a very good deal. But again, no 4090 in DCs.