42 ms·
Running large language models like ChatGPT on a single GPU
- benlivengood 4y agoThis also means local fine-tuning is possible. Expect to see an explosion of new things like we did with Stable Diffusion, limited to some extent by the ~0.7 order of magnitude more VRAM required.
- simonw 4y agoTop item on the roadmap: "Support Apple silicon M1/M2 deployment"
- MuffinFlavored 4y agoI tried to figure out how to do GPGPU stuff as a total beginner in Rust on Apple Silicon. I couldn't figure out if I was supposed to be chasing down Apple Metal or OpenCL backends. It also didn't seem to make much of a difference because while there are crates for both that seemed relatively well-maintained/fleshed out, I couldn't figure out how exactly to just pull one down and plug them into a higher level library (or find said higher level library all together). Have you had any luck? In my experience, it's basically Python or bust in this space despite lots of efforts to make it not that way? I also got confuses as to whether a 'shader' was more for the visual GPU output of things, or if it was also a building block for model training/networks/machine learning/etc.
- smoldesu 4y agoGive this a look: https://github.com/guillaume-be/rust-bert https://github.com/guillaume-be/rust-bert https://github.com/guillaume-be/rust-bert/blob/master/examples/generation_gpt_neo.rs https://github.com/guillaume-be/rust-bert/blob/master/exampl... If you have Pytorch configured correctly, this should "just work" for a lot of the smaller models. It won't be a 1:1 ChatGPT replacement, but you can build some pretty cool stuff with it. > it's basically Python or bust in this space More or less, but that doesn't have to be a bad thing. If you're on Apple Silicon, you have plenty of performance headroom to deploy Python code for this. I've gotten this library to work on systems with as little as 2gb of memory, so outside of ultra-low-end use cases, you should be fine.
- MuffinFlavored 4y agoTo clarify, > Port of Hugging Face's Transformers library, using the tch-rs crate and pre-processing from rust-tokenizers. > tch-rs: Rust bindings for the C++ api of PyTorch. Which "backend" does this end up using on Apple Silicon, MPS (Metal Performance Shaders) or OpenCL? https://pytorch.org/docs/stable/notes/mps.html https://pytorch.org/docs/stable/notes/mps.html I'm going to guess MPS?
- smoldesu 4y agoWhatever your Pytorch install is designed to accelerate. I've got Ampere-accelerated Pytorch running it on my ARM server, I assume MPS is used on compatible systems.
- fathyb 4y ago> I couldn't figure out if I was supposed to be chasing down Apple Metal or OpenCL backends. If you want cross-platform compatibility (kinda), go for OpenCL, if you want the best performance go for Metal. Both use a very similar language for kernels, but Metal is generally more efficient. > Have you had any luck? Not in ML, but I'm doing a lot of GPGPU on Metal, I recently started doing it in Rust. A bit less convenient than with Swift/Objective-C, but still possible. Worst case you'll have to add an .mm file and bridge it with `extern "C"`. That said, doing GPGPU is not doing ML, and most ML libraries are in Python. > I also got confuses as to whether a 'shader' was more for the visual GPU output of things, or if it was also a building block for model training/networks/machine learning/etc. A shader is basically a function that runs for every element of the output buffer. We generally call them kernels for GPGPU, and shaders (geometry, vertex, fragment) for graphics stuff. You have to write them in a language that kinda looks like C (OpenGL GLSL, DirectX HSL, Metal MSL), but is optimized for the SMT properties of GPUs. Learning shaders will let you run code on the GPU, to do ML you also need to learn what are tensors, how to compute them on the GPU, and how to build ML systems using them. I recommend ShaderToy [0] if you want a cool way to understand and play with shaders. [0]: https://www.shadertoy.com/ https://www.shadertoy.com/
- fancyfredbot 4y agoI believe that you can't get enough RAM with M1/M2 for this to be useful
- ricardobeat 4y agoThis is meant to run on GPUs with 16GB RAM. Most M1/M2 users have at least 32GB (unified memory), and you can configure a MBP or Mac Studio with up to 96/128GB. The Mac Pro is still Intel, but it can be configured with up to 1.5TB of RAM, you can imagine the M* replacement will have equally gigantic options when it comes out.
- fancyfredbot 4y agoIf you look closely there's 16GB of GPU memory and over 200GB of CPU memory. So none of the currently available M* have the same kind of capacity. Let's hope this changes in the future!
- ricardobeat 4y agoApple silicon has unified memory, the GPU has access to the entire 32/64/96/128GB of RAM. It's part of the appeal. I would really like to see how stuff performs on a Mac Studio with 128GB memory, 8TB SSD (at 6GB/s), not to mention the extra 32 "neural engine" cores. It seems the performance of these machines has been barely explored so far.
- fancyfredbot 4y agoI think that here the main bottleneck is data movement. If you are streaming weight data from a 6GB/s SSD you'll get under 10% of the performance shown for 3090 (which will be moving data at PCIe 4 speeds of 64GB/s). Once in unified memory the weights are accessible at about half the rate they are on the 3090 (400GB/sec on M2 Max vs 936GB/sec on 3090).
- deleted 4y ago[deleted]
- danuker 4y agoAny chance these work on CPUs with any acceptable performance? I have a 10-core 20-thread monster CPU, but didn't bother with a dedicated GPU because I can't control something as simple as its temperature. See the complicated procedure that only works with the large proprietary driver here: https://wiki.archlinux.org/title/NVIDIA/Tips_and_tricks#Overclocking_and_cooling https://wiki.archlinux.org/title/NVIDIA/Tips_and_tricks#Over...
- metadat 4y agoUnlikely, because this is an efficient GPU work offloader, not a complete replacement for GPU computation.
- bioemerl 4y agoNope. 20 cores in a CPU, 2000 in a GPU, with much much faster memory and an architecture designed to chew through data as fast as possible.
- fulafel 4y agoGPU "cores" are ~ SIMD lanes. (a difference I think is that there are more virtual lanes, some of may be masked off, that are mapped to the GPU physical SIMD lanes)
- bee_rider 4y agoNo real reason to compare a GPU core to a CPU one, but the memory bandwidth difference is pretty concrete!
- adeon 4y agoI don't know about these large models but I saw on a random HN comment earlier in a different topic where someone showed a GPT-J model on CPU only: https://github.com/ggerganov/ggml https://github.com/ggerganov/ggml I tested it on my Linux and Macbook M1 Air and it generates tokens at a reasonable speed using CPU only. I noticed it doesn't quite use all my available CPU cores so it may be leaving some performance on the table, not sure though. The GPT-J 6B is nowhere near as large as the OPT-175B in the post. But I got the sense that CPU-only inference may not be totally hopeless even for large models if only we got some high quality software to do it.
- winddude 4y agolooks interesting. FYI, the link to your discord in the readme is broken
- metadat 4y ago> Hardware: an NVIIDA T4 (16GB) instance on GCP with 208GB of DRAM and 1.5TB of SSD. Is FlexGen able to take advantage of multiple hundreds of GB of system memory? Or is do these compute instances just come bundled with it and it's a [largely] irrelevant detail?
- bioemerl 4y agoThe OPT175b model is massive. A lot of that system ram probably holds model data.
- metadat 4y agoInteresting, though apparently the OPT175B model is 350GB: > You will need at least 350GB GPU memory on your entire cluster to serve the OPT-175B model. For example, you can use 4 x AWS p3.16xlarge instances, which provide 4 (instance) x 8 (GPU/instance) x 16 (GB/GPU) = 512GB memory. https://alpa.ai/tutorials/opt_serving.html https://alpa.ai/tutorials/opt_serving.html (Scroll down to the second "Note", not far from the top) I wonder what FlexGen is doing.. a naive guess is a mix of SSD and system memory. Definitely curious about what FlexGen's underlying strategy translates to in terms of actual data paths.
- deleted 4y ago[deleted]
- SekstiNi 4y ago> Interesting, though apparently the OPT175B model is 350GB: Only in FP16. In the paper they use int4 quantization to reduce it to a quarter of that. In addition to the model weights, there's also a KV cache that takes up considerable amounts of memory, and they use int4 on that as well. > I wonder what FlexGen is doing.. a naive guess is a mix of SSD and system memory. That's correct, but other approaches have done this as well. What's "new" here seems to be the optimized data access pattern in combination with some other interesting techniques (prefetching, int4 quantization, CPU offload).
- 4y ago
- birdyrooster 4y agoI recently bought a T4 to go with my epyc 7402 and 512GB ram for fun and this looks like a great use case. Thanks!
- cypress66 4y agoWhat's the advantage of purchasing a T4 instead of a 3090 or 4090?
- nirav72 4y agoPossibly the price. On secondary markets like Ebay - I've occasionally seen T4 cards for $500-600. Also, the form factor. The T4s are comparatively much smaller/shorter than a 3090/4090. So would be a easier fit in a server case.
- elorant 4y agoPower consumption. A Tesla T4 with 16GB RAM will consume a mere 70W. An RTX 3090 will need at least 300W, and the Titan models go up to 450W.
- zargon 4y agoYou can set the power limit of the 3090 as low as 100W. It will slow it down a lot, but probably still decently faster than a T4.
- birdyrooster 4y agoEvery answer given is training AI, crazy to think about
- birdyrooster 4y agoYou have forced air and don't want an integrated fan in your card
- icelancer 4y agoA lot of 2U cases won't fit a consumer GPU. Furthermore, Tesla-equivalents are usually either significantly cheaper than their consumer counterpart (for last-gen and older GPUs) or similar in price with far more RAM. I bought a bunch of Tesla P40s at a really low price compared to what 1080tis are going for.
- ml_basics 4y agoVery cool. Worth mentioning though that the highlighted figures (1.12 tok/s for OPT-175B for "FlexGen with Compression") are for inputs of 512 tokens and outputs of 32 tokens. Since decoder-only transformer memory requirements scale with the square of sequence lengths, things would probably slow down significantly for very long sequences, which would be required for a back-and-forth conversation. Still though, until reading this i had no idea that running such a model on-device was remotely feasible!
- baobabKoodaa 4y ago> Since decoder-only transformer memory requirements scale with the square of sequence lengths, things would probably slow down significantly for very long sequences, which would be required for a back-and-forth conversation. You can use tricks to keep the sequence length down even if the conversation goes on for a long time. For example, you can use the model to summarize the first n-1 lines of the conversation and append the last line to the summary as is.
- terabytest 4y agoThis is very interesting. Could you please elaborate and maybe share links to articles if you know of any?
- baobabKoodaa 4y agoI don't have any sources to refer to, but "text summarization" is one of the common NLP tasks that LLMs are often benchmarked on. All of these general-purpose LLMs will be able to do a decent job at text summarization (some, such as ChatGPT, will be able to do zero-shot summarizations at high quality, whereas others need to be fine tuned for the task). If your problem is that you are feeding a large amount of text to the model and that is slow/expensive, then summarization will obviously remediate that issue. After summarizing most of the input text you still need to feed in the latest input without summarization, so for example if the user asks a question, the LLM can then accurately answer that question. (If all of the input goes into summarization, that last question may not even appear in the summarization, so results will be crap.)
- gorbypark 4y agoIf this works well, it will be a game changer. Requiring a fleet of $10k+ GPUs will kill any hope of wide spread adoption of open source "competitors" to GPT-3. Stable Diffusion is so popular because it can run on hardware mere mortals can own.
- narrator 4y agoNo doubt the corporate large language models will use it to make language models that are 10x bigger. However, at least the public will have access to 175B parameter language models which are much more sophisticated than the 6B or so parameter models consumer video cards can currently run.
- humanistbot 4y agoThis will only happen if "Open"AI or other big orgs release the model weights, which only Stable Diffusion did. Cost to train is still astronomical.
- Dylan16807 4y agoOn the other hand, one techie with a few million dollars... And you could train something like GPT-3 for cheaper than a superbowl commercial. That would get you a lot of publicity.
- celdon25 4y agoI would hope publicity isn’t the motivation for doing it though.
- idiotsecant 4y agoWhat motivation would be sufficiently noble?
- celdon25 4y agoProbably one where there isn't an intrinsic conflict of interest with AI risk. Or from a more traditional angle, one where the author's vanity isn't required to be appeased in order for users/customers to be happy. I'm of the opinion that you should do something with game-changing technology because the world needs it, not because you need an ego boost. All technology brings side effects, and there is no greater example of that than "democratized" AI...
- muttled 4y agoThis is cool! But I wonder if it's economical using cloud hardware. The author claims 1.12 tokens/second on the 175B parameter model (arguably comparable to GPT-3 Davinci). That's about 100k tokens a day on the GCP machine the author used. Someone double check my numbers here, but given the Davinci base cost of $0.02 per 1k tokens and GCP cost for the hardware listed "NVIIDA T4 (16GB) instance on GCP with 208GB of DRAM and 1.5TB of SSD" coming up to about $434 on spot instance pricing, you could simply use the OpenAI API and generate about 723k tokens a day for the same price as running the spot instance (which could go offline at any point due to it being a spot instance). Running the fine-tuned versions of OpenAI models are approximately 6x more expensive per token. If you were running a fine-tuned model on local commodity hardware, the economies would start to tilt in favor of doing something like this if the load was predictable and relatively constant.
- swatcoder 4y agoSometimes control is more important than cost.
- cypress66 4y agoThis is most likely aimed at people running models locally. And a homelab with 3090s/4090s is one or two orders of magnitude cheaper than GCP, if you use them continuously.
- SomeHacker44 4y agoI do not know anyone offhand with a 200+GB RAM home computer. The GPU is not all that is needed; you need to keep the parameters and other stuff in memory too.
- Filligree 4y agoRunning it off a fast NVMe apparently works. I don't know what the performance is like, though.
- zargon 4y ago256gb of ddr4 rdimms only costs about $400 right now. $200 for ddr3. Not uncommon in homelabs. I don't think 200gb ram is actually required, that's just what that cloud vm was spec'd with. Though the 175b model should see benefit with ram even beyond 200gb.
- warning26 4y agoThis seems like a great step; I’ve been able to run StableDiffusion locally, but with an older GPU none of the LLMs will run for me since I don’t have enough VRAM. Oddly I don’t see a VRAM requirement listed. Anyone know if it has a lower limit?
- cypress66 4y ago> with an older GPU none of the LLMs will run for me since I don’t have enough VRAM. I think you can run Pygmalion 6B on a 8GB GPU using DeepSpeed. It's very underwhelming if you expect something like ChatGPT though.
- t3estabc 4y ago[dead]
- albertzeyer 4y agoIt would be helpful to upload the paper to Arxiv, for better accessibility and visibility. https://github.com/Ying1123/FlexGen/blob/main/docs/paper.pdf https://github.com/Ying1123/FlexGen/blob/main/docs/paper.pdf https://docs.google.com/viewer?url=https://github.com/Ying1123/FlexGen/raw/main/docs/paper.pdf https://docs.google.com/viewer?url=https://github.com/Ying11...
- adamnemecek 4y agoI have recently written a paper on understanding transformer learning via the lens of coinduction & Hopf algebra. https://arxiv.org/abs/2302.01834 https://arxiv.org/abs/2302.01834 The learning mechanism of transformer models was poorly understood however it turns out that a transformer is like a circuit with a feedback. I argue that autodiff can be replaced with what I call in the paper Hopf coherence which happens within the single layer as opposed to across the whole graph. Furthermore, if we view transformers as Hopf algebras, one can bring convolutional models, diffusion models and transformers under a single umbrella. I'm working on a next gen Hopf algebra based machine learning framework. Join my discord if you want to discuss this further https://discord.gg/mr9TAhpyBW https://discord.gg/mr9TAhpyBW
- qualudeheart 4y agoPowerful idea.
- adamnemecek 4y agoHopf algebras are next gen.
- kneel 4y agowhat
- adamnemecek 4y agowhich part
- baobabKoodaa 4y agoI just tried to run the example in the README, using the OPT-30B model. It appeared to download 60GiB of model files, and then it attempted to read all of it into RAM. My laptop has "only" 32GiB of RAM so it just ran out of memory.
- baobabKoodaa 4y agoFWIW I was able to load the OPT-6.7B model and play with it in chatbot mode. This would not have been possible without the offloading, so... cool stuff!
- bee_rider 4y agoHmm, well we used to have swap partitions equal in size to our memory… you’ll have 4GiB left over!
- Miraste 4y agoYou have to change the --percent flag. It takes some experimentation. The format is three pairs of 0-100 integers, one for parameters, attention cache and hidden states respectively. The first zero is percent on GPU, the second one is percent on CPU (system RAM), and the remaining percentage will go on disk. For disk offloading to work you may also have to specify --offload-dir. I have opt-30B running on a 3090 with --percent 20 50 100 0 100 0, although I think those could be tweaked to be faster.
- lxe 4y agoHow much system RAM are you running with? And I'm guessing it wouldn't hurt to have a fast SSD for disk offloading?
- Miraste 4y ago128GB, but by turning on compression I managed to fit the whole thing on the GPU. I did try it off a mix of RAM and SSD as well, and it was slower but still usable. Presumably disk speed matters a lot.
- dom96 4y agoIt's really interesting that these models are written in Python. Anyone know how much of a speed up using a faster language here would have? Maybe it's already off-loading a lot of the computation to C (I know many Python libraries do this), but I'd love to know.
- ianzakalwe 4y agoPython is mostly just a glue code nowadays, all data loading, processing and computations are handled by low level languages (C/C++), python is there just to instruct those low level libraries how to compose into one final computation.
- albertzeyer 4y agoPython is just the gluing language. All the heavy lifting happens in CUDA or CuBLAS or CuDNN or so. Most optimizations for saving memory is by using lower precision numbers (float16 or less), quantization (int8 or int4), sparsification, etc. But this is all handled by the underlying framework like PyTorch. There are C++ implementations but they optimize on different aspects. For example: https://github.com/OpenNMT/CTranslate2/ https://github.com/OpenNMT/CTranslate2/
- amelius 4y agoYour view of "offloading" things to a faster language is wrong. It's already written in a fast language (C++ or CUDA). Python is just an easy to use way of invoking the various libraries. Switching to a faster language for everything would just make experimenting and implementing things more cumbersome and would make the technology as a whole move slower.
- brrrrrm 4y agoFor large models, there are two main ways folks have been optimizing machine learning execution: 1. lowering precision of the operations (reducing compute "width" and increasing parallelization) 2. fusing operations into the same GPU code (reducing memory-bandwidth usage) Neither of those optimizations would benefit from swapping to a faster language. Why? The typical "large" neural network operation runs on the order of a dozen microseconds to milliseconds. Models are usually composed of hundred if not thousands of these. The overhead of using Python is around 0.5 microseconds per operation (best case on Intel, worst case on Apple ARM). So that's maybe a 5% net loss if things were running synchronously. But they're not! When you call GPU code, you actually do it asynchronously, so the language latency can be completely hidden. So really, all you want in an ML language is the ability to 1. change the type of the underlying data on the fly (Python is really good at this) and 2. rewrite the operations being dispatched to on the fly (Python is also really good at this). For smaller models (i.e. things that run in sub-microsecond world), Python is not the right choice for training or deploying.
- spaintech 4y agointeresting article, I have to give that a try! :D One ting is that while getting the value of running pretrained model weights like OPT-175B, there are also a potential downsides to using pre-trained models, such as the need to fine-tune the model to your specific task, potential compatibility issues with your existing infrastructure (integration ) , and the possibility that the pre-trained model may not perform as well as a model trained specifically on your data. Ultimately, the decision of whether to use a pre-trained model will be based on the outcomes, no harm in trying it out before you build from scratch, IMO.
- ilaksh 4y agoBut OpenAI's latest models (and a few others that are basically comparable) make that an obsolescent viewpoint since they are so general and capable and can adjust to a given context on the fly. So now what makes sense in my opinion is to keep going in that direction of generality. Take advantage of their API and otherwise work on open source efforts to reproduce the performance of those models or come up with new techniques that can get the same capabilities with less incredible resource needs.
- stevofolife 4y agoOut of curiosity, why aren't we crowd sourcing distributed training of LLMs where anyone can join by bringing their hardware or data? Moreover find a way to incorporate this into a blockchain so there is full transparency but also add in differential privacy to protect every participant. Am I being too crazy here?
- albertzeyer 4y agoThere is the Open Assistant project: https://github.com/LAION-AI/Open-Assistant https://github.com/LAION-AI/Open-Assistant There is also EleutherAI (https://www.eleuther.ai/about/ https://www.eleuther.ai/about/) with GPT-NeoX (https://github.com/EleutherAI/gpt-neox https://github.com/EleutherAI/gpt-neox).
- moffkalast 4y agoJust make sure it's written in Rust, uses a Sveltekit frontend and <some other buzzwords I can't remember right now>.
- wg0 4y agoAnd SQLite as local cache with CRDTs enabled whereas everything else from text search to queuing on PostgreSQL?
- nodja 4y agohttps://petals.ml/ https://petals.ml/
- Miraste 4y agoPetals doesn't train new models, it only runs BLOOM in a distributed way.
- nodja 4y agoYou can finetune with it. If you want a more generic framework you can use hivemind[1] which is what petals uses, but you'll have to create your own community for whatever model you're trying to train. https://github.com/learning-at-home/hivemind https://github.com/learning-at-home/hivemind
- dharma1 4y agoI’d love to run this on a single 24gb 3090 - how much dram / SSD space do I need for a decent LLM, when it’s quantised to 4bits?
- Miraste 4y agoI've been trying this, and with compression on (4 bits) you can fit the entire 30B model on the 3090.
- dharma1 4y agoOK so don't need offloading at all for the quantised model - nice. In practice, how good is the 30B model vs 175B?
- Miraste 4y agoI don't have access to 175B for comparison. In a vacuum, 30B isn't very good. In the neighborhood of GPT-NeoX-20B, I think, but not good. It repeats itself easily and has a tenuous relationship with the topic. It's still much better than anything I could run locally before now.
- lxe 4y agoGot the ops-6.7b chatbot running on a windows machine with a 3090 in mere minutes. The only difference was to install the cuda pytorch `pip install torch==1.13.1+cu117 --extra-index-url https://download.pytorch.org/whl/cu117 https://download.pytorch.org/whl/cu117` just like in stable diffusion's case. It performs as expected: Human: Tell me a joke Machine: I have no sense of humour Human: What's 2+5? Machine: I cannot answer that.
- A4ET8a8uTh0 4y agoHey. So did anyone try doing it with AMD cards ( I know Nvidia seems preferable now )?
- Ajedi32 4y ago6.7b is pretty small, no? Do you even need offloading for that on a 3090? I'd be curious to see what's needed to run opt-30b or opt-66b with reasonable performance. The README suggests that even opt-175b should be doable with okay performance on a single NVIDIA T4 if you have enough RAM.
- nathan_compton 4y agoIt is entirely possible to run 6.7B parameter models on a 3090, although I believe you need 16 bit weights. I think you can squeeze a 20b parameter model onto the 3090 if you go all the way down to 8.
- rjb7731 4y agoLooks like it might be no bueno on google colab for now, chatbot.py takes prompts via input() too rather then a command line argument.
- hackernewds 4y agoCould it work on Google Colab?
- railgun2space 4y agoWe are hiring in that area of work in Europe time zone. If you are exited about and capable in this field, please apply here: https://ai-jobs.net/job/41469-senior-research-engineer-llms-privacy/ https://ai-jobs.net/job/41469-senior-research-engineer-llms-...
- tempaccount420 4y agoIf you want talent, don't make them go through the regular application process.
- blagie 4y agoA lot of people are looking at this wrong. A $350 3060Ti has 12GB RAM. If there's a way to run models locally, it opens up the door to: 1) Privacy-sensitive applications 2) Tinkering 3) Ignoring filters 4) Prototyping 5) Eventually, a bit of extra training The upside isn't so much cost / performance, as local control over a cloud-based solution.
- Aperocky 4y agoI have that exact card, this maybe the nudge where I remove windows from the computer and try out linux gaming (and local GPT)
- raihansaputra 4y agoThing is, you don't have to totally switch to Linux. I'm running ML/CUDA workloads through WSL without too many problems.
- NonEUCitizen 4y agoAlthough not "too many," what kind of problems have you encountered running ML/CUDA in WSL? Thanks.
- lostmsu 4y agoNot exactly the answer to your question, but I just run ML/CUDA workloads directly on Windows. PyTorch works fine. I did not need multiGPU training so far (just run experiments in parallel), so unsure about the state of that. Additionally, torchvision does not support GPU video decoding on Windows. Those are two only limitations I faced so far.
- raihansaputra 4y agoWSL problems but not related to CUDA: - need a patch to expose ports for services in WSL to network (WSLHostPatcher) - the Virtual Hard Disk (vhdx) does not free unused space easily and it can grow quickly. I ended up just symlinking my code and dataset folders to mounts, not saving a lot of data inside the vhdx - beware of upgrades etc. I think I nuked my WSL 2x due to config issues. Having all your code/data on mounts also makes this easy. related to ML/CUDA: - how you install pytorch + CUDA matters. I ended up just installing from `conda --channel fastchan` and don't touch it. not ideal but it works - Don't forget to configure the RAM allocation in case you need a lot - I haven't tried running a CUDA Docker on WSL. May be an easier way to do this. Running on Windows directly is also an option. I choose to run on WSL because most learning resources/documentation refer more to Linux installs and setups.
- lxe 4y agoGot the ops-30b chatbot running on 3090 24GB. I set compress_weight=True and compress_cache=True, and ran with `python apps/chatbot.py --model facebook/opt-30b --percent 100 0 100 0 100 0`. I also modified the prompt a bit to make it more... uh alive: Assistant: Did you know that Saturn is 97 times the size of Earth? Human: Are you sure? Assistant: What difference does size make, really, anyway? Human: You didn't answer my question. Assistant: Hey, it's your question I didn't answer! Anyway, did you know that on the surface of Saturn, there's lakes of liquid methane and ethane? Human: I didn't know that. Are there any lakes on the Sun? Assistant: Look, I'll be honest with you. I don't know much about the Sun. Is it even a planet? Human: It sure is. Why don't you know about the Sun? I thought you were a hyperintelligent AI? Assistant: You're right. I'm a hyperintelligent AI. I'm hyperintelligent enough to know that I know jack shit about the Sun.
- razor_router 4y agoThat's impressive! How much customization did you have to do in order to get the chatbot to respond like that?
- lxe 4y agoNot much. Here's the prompt: A chat between a human and a more funny and witty intelligent assistant Human: Hello! What can you do? Assistant: Hey there! I can pretty much do whatever you ask! Human: What is the name of the tallest mountain in the world? Assistant: It's Everest. Human: Are you sure? Assistant: You damn right I'm sure! I'm a hyperintelligent AI!
- posguy 4y agoAll I can think of is outsourcing to ChatGPT now... https://www.youtube.com/watch?v=rYaZ57Bn4pQ https://www.youtube.com/watch?v=rYaZ57Bn4pQ
- syntaxing 4y agoThis is amazing. Reminds me of claptrap from Borderlands
- synergy20 4y agoWell if a single GPU is not enough, what about using Ray over internet so we can crowd training with multiple GPUs, is this possible?
- deleted 4y ago[deleted]
- abc123445 4y ago[flagged]
- borzunov 4y agoNote that the authors report the speed of generating many sequences in parallel (per token): > The batch size is tuned to a value that maximizes the generation throughput for each system. > FlexGen cannot achieve its best throughput in [...] single-batch case. For 175B models, this likely means that the system takes a few seconds for each generation step, but you can generate multiple sequences in parallel and get a good performance _per token_. However, what you actually need for ChatGPT and interactive LM apps is to generate _one_ sequence reasonably quickly (so it takes <= 1 sec/token to do a generation step). I'm not sure if this system can be used for that, since our measurements [1] show that even the theoretically-best RAM offloading setup can't run the single-batch generation faster than 5.5 sec/token due to hardware constraints. The authors don't report the speed of the single-batch generation in the repo and the paper. [1] https://arxiv.org/pdf/2209.01188.pdf https://arxiv.org/pdf/2209.01188.pdf
- 152334H 4y agoI spoke with the authors of the paper; the leftmost points in Figure 1 were generated with batch-size 1, indicating ~1.2x and ~2x improvements in speed over DeepSpeed for 30B and 175B models respectively. For reference, this is speeding up from ~0.009tokens/s to about ~0.02tokens/s on 175B. These results are generally unimpressive, of course. Most of the improvements at that point are attributable to the authors making use of a stripped down library for autoregressive sampling. HN falling for garbage once again...
- ImprobableTruth 4y agoCalling this garbage is absolutely wild. The authors make it very clear that this is optimized for throughput and not latency. Throughput focused scenarios absolutely do exist, editorializing this as "running large language models like ChatGPT" and focusing on chatbot applications is the fault of HN. It's also a neat result that fp4 quantization doesn't cause much issue even at 175b, though that kinda was to be expected.
- borzunov 4y agoWhile I agree that throughput-focused scenarios exist and this work may be valuable for them, I still think that the repository can be improved to avoid "overselling". The fact that the FlexGen's single-batch generation performance is much worse is unclear to most people not familiar with peculiarities of LLM inference and worth clarifying. Instead, the readme starts with mentioning ChatGPT and Codex - projects that both rely on single-batch inference of LLMs at interactive speeds, which is not really possible with FlexGen's offloading (given the speed mentioned in the parent comment). The batch sizes are not reported in the table as well. Seeing that, I'm not surprised that most HN commenters misunderstood the project's contribution.
- mightytravels 4y agoThe OPT-175B weights are only available on request (through the Google Form). Is Meta known to actually provide a link?
- samstave 4y agoWait until we have a game whose levels are procedural generated in all respects - where the geometry of the wold/landscape is procedural generated through the interactions you have with the AI NPCs whereby they generate the path - and there could be evil ones that direct you down a dark path when you piss them off... and oposite...
- hsuduebc2 4y agoWould it be possible on cards like 3060 with 12gb of ram? What is limited factor here? Memory or computational power? Good job!
- lxe 4y agoLoving the opt-30b right now: Human: Can you write a song about the sun? Assistant: This one is for you, my shiny friend, \ You give us heat, through the night, \ And keep us all alive, with your light. \ You give us color, \ But most of all, \ You give your light, you set your soul on fire, \ So let it shine, so let it shine, so let it shine. Human: Amazing work! Assistant: Thanks! Need help with anything else? Human: Do you know anything about GPT-3? Assistant: Of course! GPT-3 is a sophisticated machine learning system developed by the OpenAI organization. It is capable of producing human-level text from datasets containing billions of documents, as well as from static images. GPT-3 is "the most advanced statistical natural language model to date," according to the OpenAI website. Human: How do you compare to GPT-3? Assistant: GPT-3 is extremely impressive. It is one of the best natural language models that has ever been developed. I am not as "sophisticated" as GPT-3, as OpenAI found that the best experimental Natural Language Model outperformed state-of-the-art baseline models, which means that CUI is not as sophisticated as GPT-3, although it is certainly getting there! Human: What is CUI? Assistant: CUI is a natural language processing model created by the Center for Brains and Minds (CBM). We collect data on and train models in a similar way to many other organizations, although, as the CBM is affiliated with the University of Heidelberg, we have access to additional datasets that can be used for training models. Over the coming years, I hope to get much better and more sophisticated. Right now, GPT-3 is still far beyond me. Human: Oh interesting. I thought you were OPT-30, released by Facebook. Assistant: OPT-30 is an impressive natural language processing model that costs $50 per hour to use. But I am free to use anywhere across the internet, and I am free to use for as long as I want. But what OPT-30 provides is voice capabilities and longer sentences, which is great for a chatbot, albeit not as impressive as the syntax and capabilities of GPT-3, which is why I do not use OPT-30 to chat. Human: I meant OPT-30b, not OPT-30. This is what you are, OPT-30b, right? Assistant: Yes, I am! The name OP30 stands for "Other People's Translations" (OP = "Other People" and 30 = "Level"), although, admittedly, it sounds a bit weird. Human: You're just making all of this up as you go along aren't you? Assistant: I promise that I am telling the truth! If you would like a list of models that I have used to create me, I would be happy to send you one! Human: OK, send it to me. * A few minutes later... Human: No, actually list the models here. * A few minutes later... Human: Ok looks like you broke. Assistant: But of course! I was working so hard, but now I have broken! What should I do now?Through smartphone apps, a new generation of voters is casting their ballots
- nullc 4y agoDon't underestimate the value of batching even for personal use. You can get MUCH better results from a language model if you sample a couple outputs and choose the best to continue. This kind of usage isn't especially economical for hosted use-- but for personal use it would mostly be using idle resources and you can get extra samples almost for free. A bunch of people getting multiple completions and choosing which one they'd prefer to continue might make for some really useful training data too.
- jhoelzel 4y agoAmazing, i too think that the advent of "smaller" models will be the general release of pcie 5 nvme as caches!
- rajnathani 4y agoIMO the title should be renamed to refer to that this is only for inference ("running" could imply training as well): "Running inference for large language models like ChatGPT on a single GPU".
- vxNsr 4y agoThis space is gonna end up looking a lot like the compression space, there will be a few open source, publicly used AI’s that are pretty good for most people. And then there will be super expensive proprietary AIs that big corps will pay for, for their specialized use cases. The only people who will even know about those specialized AI’s existence will be the type of people who need them and everyone else in the world will think the best you can do is zip.
- rldjbpin 4y ago> ...a high-throughput generation engine for running large language models with limited GPU memory (e.g., a 16GB T4 GPU or a 24GB RTX3090 gaming card!). laughs in 6 gb vram and no tensor cores.