20 ms·
Run Llama 13B with a 6GB graphics card
- anshumankmr 3y agoHow long before it runs on a 4 gig card?
- rain1 3y agoYou can offload only 10 layers or so if you want to run on a 4GB card
- tarr11 3y agoWhat is the state of the art on evaluating the accuracy of these models? Is there some equivalent to an “end to end test”? It feels somewhat recursive since the input and output are natural language and so you would need another LLM to evaluate whether the model answered a prompt correctly.
- klysm 3y agoIt’s going to be very difficult to come up with any rigorous structure for automatically assessing the outputs of these models. They’re built using effectively human grading of the answers
- RockyMcNuts 3y agohmmh, if we have the reinforcement learning part of reinforcement learning with human feedback, isn't that a model that takes a question/answer pair and rates the quality of the answer? it's sort of grading itself, it's like a training loss but it still tells us something?
- sroussey 3y agoLlama cpp and others use perplexity: https://huggingface.co/docs/transformers/perplexity https://huggingface.co/docs/transformers/perplexity
- tikkun 3y agohttps://chat.lmsys.org/?arena https://chat.lmsys.org/?arena (Click 'leaderboard')
- dclowd9901 3y agoHas anyone tried running encryption algorithms through these models? I wonder if it could be trained to decrypt.
- Hendrikto 3y agoThat would be very surprising, given that any widely used cryptographic encryption algorithm has been EXTENSIVELY cryptanalyzed. ML models are essentially trained to recognize patterns. Encryption algorithms are explicitly designed to resist that kind of analysis. LLMs are not magic.
- dclowd9901 3y agoAll of what you said is true, for us. I know LLMs aren’t magic (lord knows I actually kind of understand the principles of how they operate), but they have a much greater computational and relational bandwidth than we’ve ever had access to before. So I’m curious if that can break down what otherwise appears to be complete obfuscation. Otherwise, we’re saying that encryption is somehow magic in a way that LLMs cannot possibly be.
- NegativeK 3y ago> Otherwise, we’re saying that encryption is somehow magic in a way that LLMs cannot possibly be. I don't see why that's an unreasonable claim. I mean, encryption isn't magic, but it is a drastically different process.
- nl 3y ago> So I’m curious if that can break down what otherwise appears to be complete obfuscation. This seems to be a complete misunderstanding of what encryption is. Obfuscation generally means muddling things around in ways that can be reconstructed. It's entirely possible a (custom - because you'd need custom tokenization) LLM could deobfuscate things. Encryption OTOH means using a piece of information that isn't present. Weak encryption gets broken because that missing information can be guessed or recovered easily. But this isn't the case for correctly implemented strong encryption. The missing information cannot be recovered by any non-quantum process in a reasonable timeframe. There are exceptions - newly developed mathematical techniques can sometimes make recovering that information quicker. But in general math is the weakest point of LLMs, so it seems an unlikely place for them to excel.
- syntaxing 3y agoThis update is pretty exciting, I’m gonna try running a large model (65B) with a 3090. I have ran a ton of local LLM but the hardest part is finding out the prompt structure. I wish there is some sort of centralized data base that explains it.
- rain1 3y agoTell us how it goes! Try different numbers of layers if needed. A good place to dig for prompt structures may be the 'text-generation-webui' commit log. For example https://github.com/oobabooga/text-generation-webui/commit/334486f527bc97f61eb3264def4e03a0dab9b369 https://github.com/oobabooga/text-generation-webui/commit/33...
- int_19h 3y agoI tried llama-65b on a system with RTX 4090 + 64Gb of DDR5 system RAM. I can push up to 45 layers (out of 80) to the GPU, and the overall performance is ~800ms / token, which is "good enough" for real-time chat.
- guardiangod 3y agoI got the alpaca 65B GGML model to run on my 64GB ram laptop. No GPU required if you can tolerate the 1 token per 3 seconds rate.
- syntaxing 3y agoSupposedly the new update with GPU offloading will bring that up to 10 tokens per second! 1 token per second is painfully slow, that’s about 30s for a sentence.
- int_19h 3y ago10 tokens / second is what you get running llama-30b entirely on the GPU. A 65b model will be slower than that since there's more compute involved.
- holoduke 3y agoWhy does AMD or Intel not release a medium performant GPU with minimum 128gb of memory for a good consumer price. These models require lots of memory to 'single' pass an operation. Throughput could be bit slower. A 1080 Nvidia with 256gb of memory would run all these models fast right? Or am I forgetting something here.
- hackernudes 3y agoI don't think there was a market for it before LLMs. Still might not be (especially if they don't want to cannibalize data center products). Also, they might have hardware constraints. I wouldn't be that surprised if we see some high ram consumer GPUs in the future, though. It won't work out unless it becomes common to run LLMs locally. Kind of a chicken-and-egg problem so I hope they try it!
- the8472 3y ago> I don't think there was a market for it before LLMs. At $work CGI assets sometimes grow pretty big and throwing more VRAM at the problem would be easier than optimizing the scenes in the middle of the workflow. They can be optimized, but that often makes it less ergonomic to work with them. Perhaps asset-streaming (nanite&co) will make this less of an issue, but that's also fairly new. Do LLM implementations already stream the weights layer by layer or in whichever order they're doing the evaluation or is PCIe bandwidth too limited for that?
- rahimnathwani 3y agoOn my system, using `-ngl 22` (running 22 layers on the GPU) cuts wall clock time by ~60%. My system: GPU: NVidia RTX 2070S (8GB VRAM) CPU: AMD Ryzen 5 3600 (16GB VRAM) Here's the performance difference I see: CPU only (./main -t 12) llama_print_timings: load time = 15459.43 ms llama_print_timings: sample time = 23.64 ms / 38 runs ( 0.62 ms per token) llama_print_timings: prompt eval time = 9338.10 ms / 356 tokens ( 26.23 ms per token) llama_print_timings: eval time = 31700.73 ms / 37 runs ( 856.78 ms per token) llama_print_timings: total time = 47192.68 ms GPU (./main -t 12 -ngl 22) llama_print_timings: load time = 10285.15 ms llama_print_timings: sample time = 21.60 ms / 35 runs ( 0.62 ms per token) llama_print_timings: prompt eval time = 3889.65 ms / 356 tokens ( 10.93 ms per token) llama_print_timings: eval time = 8126.90 ms / 34 runs ( 239.03 ms per token) llama_print_timings: total time = 18441.22 ms
- rain1 3y agoThat is a crazy speedup!!
- GordonS 3y agoIs it really? Going from CPU to GPU, I would have expected a much better improvement.
- qwertox 3y agoI feel the same. For example some stats from Whisper [0] (audio transcoding, 30 seconds) show the following for the medium model (see other models in the link): --- GPU medium fp32 Linear 1.7s CPU medium fp32 nn.Linear 60.7s CPU medium qint8 (quant) nn.Linear 23.1s --- So the same model runs 35.7 times faster on GPU, and compared to an "optimized" model still 13.6. I was expecting around an order or magnitude of improvement. Then again, I do not know if in the case of this article the entire model was in the GPU, or just a fraction of it (22 layers) and the remainder on CPU, which might explain the result. Apparently that's the case, but I don't know much about this stuff. [0] https://github.com/MiscellaneousStuff/openai-whisper-cpu https://github.com/MiscellaneousStuff/openai-whisper-cpu
- dinobones 3y agoWhat is HN’s fascination with these toy models that produce low quality, completely unusable output? Is there a use case for them I’m missing? Additionally, don’t they all have fairly restrictive licenses?
- tbalsam 3y agoI never thought I'd see the day when a 13B model was casually referred to in a comments section as a "toy model".
- andrewmcwatters 3y agoStart using it for tasks and you'll find limitations very quickly. Even ChatGPT excels at some tasks and fails miserably at others.
- tbalsam 3y agoOh, I've been using language models before a lot (or at least some significant chunk) of HN knew the word LLM, I think. I remember when going from 6B to 13B was crazy good. We've just normalized our standards to the latest models in the era. They do have their shortcomings but can be quite useful as well, especially the LLama class ones. They're definitely not GPT-4 or Claude+, for sure, for sure.
- az226 3y agoCompared to GPT2 it’s on par. Compared to GPT3, 3.5, or 4, it’s a toy. GPT2 is 4 years old, and in terms of LLMs, that’s several life times ago. In 5-10 years, GPT3 will be viewed as a toy. Note, “progress” will unlikely be as fast as it has been going forward.
- tbalsam 3y agoGPT-2's largest model was 1.5B params, LLama-65B was similar to the largest GPT3 in benchmark performance but that model was expensive in the API, a number of the people would use the cheaper one(s) instead IIRC. So this is similar to a mid tier GPT3 class model. Basically, there's not much reason to Pooh-Pooh it. It may not perform quite as well, but I find it to be useful for the things it's useful for.
- ranger_danger 3y agoWhy can't these models run on the GPU while also using CPU RAM for the storage? That way people will performant-but-memory-starved GPUs can still utilize the better performance of the GPU calculation while also having enough RAM to store the model? I know it is possible to provide system RAM-backed GPU objects.
- ACV001 3y agoThe future is this - these models will be able to run on smaller and smaller hardware eventually being able to run on your phone, watch or embedded devices. The revolution is here and is inevitable. Similar to how computers evolved. We are still lucky that these models have no consciousness, still. Once they gain consciousness, that will mark the appearance of a new species (superior to us if anything). Also, luckily, they have no physical bodies and cannot replicate, so far...
- canadianfella 3y ago[dead]
- olabyne 3y agoThe phone part is already there ! https://mlc.ai/mlc-llm/ https://mlc.ai/mlc-llm/ (granted, this is only a 7b-model running with 4bits)
- alg_fun 3y agowouldn't i be faster to use ram as a swap for vram?
- sroussey 3y agoI wish this used the webgpu c++ library instead, then it could be used in any GPU hardware.
- qwertox 3y agoIf I really want to do some playing around in this area, would it be good to get a RTX 4000 SFF which has 20 GB of VRAM but is a low-power card, which I want as it would be running 24/7 and energy prices are pretty bad in Germany, or would it make more sense to buy an Apple product with some M2 chip which apparently is good for these tasks as it shares CPU and GPU memory?
- blendergeek 3y agoIs there a way to run any of these with only 4GB of VRAM?
- washadjeffmad 3y agoAssuming an nvidia GPU and requisite system memory, use llama.cpp compiled with cublas support, then run with the -ngl [n layers] option. You'll need a model quantized after May 12 to work with this. The smallest GPU-only 7B 4-bit model requires 8GB VRAM, so it's either do CPU only or use the GPU offload above.
- blendergeek 3y agoThank you! I'll give it a try.
- akulbe 3y agoI've only ever been a consumer of ChatGPT/Bard. Never set up any LLM stuff locally, but the idea is appealing to me. I have a ThinkStation P620 w/ThreadRipper Pro 3945WX (12c24t) with a GTX 1070 (and a second 1070 I could put in there) and there's 512GB of RAM on the box. Does this need to be bare metal, or can it run in VM? I'm currently running RHEL 9.2 w/KVM (as a VM host) with light usage so far.
- MuffinFlavored 3y agoHow many "B" (billions of parameters) is ChatGPT GPT-4?
- sciolist 3y agoInformation about GPT-4 was not released
- yawnxyz 3y agoCould someone please share a good resource for building a machine from scratch, for doing simple-ish training and running open-source models like Llama? I'd love to run some of these and even train them from scratch, and I'd love to use that as an excuse to drop $5k on a new machine... Would love to run a bunch of models on the machine without dripping $$ to OpenAI, Modal or other providers...
- vonseel 3y agoI am no where near an expert on this subject, and this information is from a few months ago so maybe it's outdated, but people on Reddit[1] are claiming running the llama with 65B parameters would need like 20K+ of GPUs. A 40GB A100 looks like it's almost $8K on Amazon, and I'm sure you could do a lot with just one of those, but that's already beyond your $5K budget. [1] https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_is_it_possible_to_run_metas_llama_65b_model_on/ https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_... I'll let others chime in but you could still probably build something really powerful within your budget that is able to run various AI tasks.
- logicchains 3y agoYou can get around 4-5 tokens per second on the 65B LLaMA with a 32 core 256GB ram Ryzen CPU, not sure how much it costs to build but can rent one from Hetzner for around two hundred bucks a month.
- Joeri 3y agoThere are some threads with hardware recommendations in the LocalLLaMa subreddit. Here’s a recent one: https://www.reddit.com/r/LocalLLaMA/comments/13f5gwn/home_llm_hardware_suggestions/ https://www.reddit.com/r/LocalLLaMA/comments/13f5gwn/home_ll...
- rahimnathwani 3y agoPSA: If you're using oobabooga/text-generation-webui then you need to: 1. Re-install llama-cpp-python with support for CUBLAS: CMAKE_ARGS="-DLLAMA_CUBLAS=on" FORCE_CMAKE=1 pip install llama-cpp-python --no-cache-dir --force-reinstall 2. Launch the web UI with the --n-gpu-layers flag, e.g. python server.py --model gpt4-x-vicuna-13B.ggml.q5_1.bin --n-gpu-layers 24
- BlackLotus89 3y agoThis only uses llama correct? So the output should be the same as if you were only using llama.cpp. Am I the only one who doesn't get nearly the same quality of output using a quantized model compared to GPU? Some models I tried get astounding results when running on a GPU, but create only "garbage" when running on a CPU. Even when not quantized down to 4bit llama.cpp just doesn't compare for me. Am I alone with this?
- Ambix 3y agoNo need to convert models, 4bit LLaMA versions for GGML v2 available here: https://huggingface.co/gotzmann/LLaMA-GGML-v2/tree/main https://huggingface.co/gotzmann/LLaMA-GGML-v2/tree/main
- naillo 3y agoThis is cool but are people actually getting stuff done with these models? I'm enthusiastic about their potential too but after playing with it for a day I'm at a loss for what to use it for anymore at this point
- Zetobal 3y agoWe run some llamas to analyze user content.
- jhbadger 3y agoDepends on what "getting stuff done" means. I find 13B models (running on my M1 Mac) useful for playing AI Dungeon-like games -- where you describe the environment and your character and you can play an RPG.
- fredguth 3y agoGitHub Copilot is (or once was) a 13b model, according to Nat Friedman in the scale.ai interview. (https://youtu.be/lnufceCxwG0 https://youtu.be/lnufceCxwG0)
- xrd 3y agoCan you share more about what you are doing or what you've seen others doing?
- jhbadger 3y agoJust noticed this response, but generally something like "act as a game master of a rpg. The setting is in a medieval style world with magic and my character is a rogue who is good at fast talking and pickpocketing" Then the LLM generally creates an adventure that I can interact with,
- Joeri 3y agoMy experience is that even the "toy" 7B models can do pretty amazing things, but they're very sensitive to prompting in the right way to force the model down the right pathways. My assumption is that the larger models have more redundancy and therefore can be prompted less precisely while still going down the right pathway. For example, the WizardLM model (5_1 quantized, as 4_x quantization makes the 7B models perform terribly): Prompt: I went to the market and bought 10 apples. I gave 2 to my neighbor and 2 to my sister. Then I bought 5 more and ate 1. How many apples do I have? Response: Let me check that for you! You had 9 apples left after giving away 4, so you still have 9 - 4 = 5 apples left. Prompt: I went to the market and bought 10 apples. I gave 2 to my neighbor and 2 to my sister. Then I bought 5 more and ate 1. How many apples do I have? Write out any calculations in full. List the steps used to come to the final result. Response: Sure, let's work through this problem together! Here are the steps we can use: Step 1: Start with the initial number of apples you bought (10). Step 2: Subtract the amount given to your neighbor and sister from the total (2 + 2 = 4). This gives us a new total of 6 apples. Step 3: Add the additional 5 apples you bought later on to the remaining 6 apples, giving us a final total of 11 apples. Step 4: Finally, subtract the apple you ate from the total (1), leaving us with 10 apples in total. So the answer is 10 apples.
- s_dev 3y ago[deleted]
- capableweb 3y agoFirst link: https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp Which in turn has the following as the first link: https://arxiv.org/abs/2302.13971 https://arxiv.org/abs/2302.13971 Is it really quicker to ask here than just browse content for a bit, skimming some text or even using Google for one minute?
- djbusby 3y agoYou gave an awesome answer in 2 minutes! Might be faster than reading!
- capableweb 3y agoIf you cannot click two links in a browser under two minutes, I'm either sorry for you, or scared of you :)
- s_dev 3y ago>Is it really quicker to ask here than just browse content for a bit, skimming some text or even using Google for one minute? I don't know if it's quicker but I trust human assessment a lot more than any machine generated explanations. You're right I could have asked ChatGPT or even Googled but a small bit of context goes a long way and I'm clearly out of the loop here -- it's possible others arrive on HN might appreciate such an explanation or we're better off having lots of people making duplicated efforts to understand what they're looking at.
- capableweb 3y agoWell, I'm saying if you just followed the links on the submitted page, you'd reach the same conclusion but faster.
- rain1 3y ago
- bitL 3y agoHow about reloading parts of the model as the inference progresses instead of splitting it into GPU/CPU parts? Reloading would be memory-limited to the largest intermediate tensor cut.
- moffkalast 3y agoThe Tensor Reloaded, starring Keanu Reeves
- regularfry 3y agoThat would turn what's currently an L3 cache miss or a GPU data copy into a disk I/O stall. Not that it might not be possible to pipeline things to make that less of a problem, but it doesn't immediately strike me as a fantastic trade-off.
- bitL 3y agoOne can keep all tensors in the RAM, just push whatever needed to GPU VRAM, basically limited by PCIe speed. Or some intelligent strategy with read-ahead from SSD if one's RAM is limited. There are even GPUs with their own SSDs.
- hhh 3y agoInstructions are a bit rough. The Micromamba thing doesn’t work, doesn’t say how to install it… you have to clone llama.cpp too
- tikkun 3y agoSee also: https://www.reddit.com/r/LocalLLaMA/comments/13fnyah/you_guys_are_missing_out_on_gpt4x_vicuna/ https://www.reddit.com/r/LocalLLaMA/comments/13fnyah/you_guy... https://chat.lmsys.org/?arena https://chat.lmsys.org/?arena (Click 'leaderboard')
- marcopicentini 3y agoWhat do you use to host these models (like Vicuna, Dolly etc) on your own server and expose them using HTTP REST API? Is there an Heroku-like for LLM models? I am looking for an open source models to do text summarization. Open AI is too expensive for my use case because I need to pass lots of tokens.
- inhumantsar 3y agoWeights and Biases is good for building/training models and Lambda Labs is a cloud provider for AI workloads. Lambda will only get you up to running the model though. You would still need to overlay some job management on top of that. I've heard Run.AI is good on that front but I haven't tried.
- month13 3y agohttps://bellard.org/ts_server/ https://bellard.org/ts_server/ may be what you are after. You can run open-source models, but the software itself is closed-source and free for non-commercial use.
- rain1 3y agoI haven't tried that but https://github.com/abetlen/llama-cpp-python https://github.com/abetlen/llama-cpp-python and https://github.com/r2d4/openlm https://github.com/r2d4/openlm exists
- speedgoose 3y agoThese days I use FastChat: https://github.com/lm-sys/FastChat https://github.com/lm-sys/FastChat It’s not based on llama.cpp but huggingface transformers but can also run on CPU. It works well, can be distributed and very conveniently provide the same REST API than OpenAI GPT.
- itake 3y agoDo you know how well it performs compared to llama.cpp?
- 3y ago
- peatmoss 3y agoFrom skimming, it looks like this approach requires CUDA and thus is Nvidia only. Anyone have a recommended guide for AMD / Intel GPUs? I gather the 4 bit quantization is the special sauce for CUDA, but I’d guess there’d be something comparable for not-CUDA?
- rain1 3y ago4-bit quantization is to reduce the amount of VRAM required to run the model. You can run it 100% on CPU if you don't have CUDA. I'm not aware of any AMD equivalent yet.
- amelius 3y agoLooks like there are several projects that implement the CUDA interface for various other compute systems, e.g.: https://github.com/ROCm-Developer-Tools/HIPIFY/blob/master/README.md https://github.com/ROCm-Developer-Tools/HIPIFY/blob/master/R... https://github.com/hughperkins/coriander https://github.com/hughperkins/coriander I have zero experience with these, though.
- westurner 3y ago"Democratizing AI with PyTorch Foundation and ROCm™ support for PyTorch" (2023) https://pytorch.org/blog/democratizing-ai-with-pytorch/ https://pytorch.org/blog/democratizing-ai-with-pytorch/ : > AMD, along with key PyTorch codebase developers (including those at Meta AI), delivered a set of updates to the ROCm™ open software ecosystem that brings stable support for AMD Instinct™ accelerators as well as many Radeon™ GPUs. This now gives PyTorch developers the ability to build their next great AI solutions leveraging AMD GPU accelerators & ROCm. The support from PyTorch community in identifying gaps, prioritizing key updates, providing feedback for performance optimizing and supporting our journey from “Beta” to “Stable” was immensely helpful and we deeply appreciate the strong collaboration between the two teams at AMD and PyTorch. The move for ROCm support from “Beta” to “Stable” came in the PyTorch 1.12 release (June 2022) > [...] PyTorch ecosystem libraries like TorchText (Text classification), TorchRec (libraries for recommender systems - RecSys), TorchVision (Computer Vision), TorchAudio (audio and signal processing) are fully supported since ROCm 5.1 and upstreamed with PyTorch 1.12. > Key libraries provided with the ROCm software stack including MIOpen (Convolution models), RCCL (ROCm Collective Communications) and rocBLAS (BLAS for transformers) were further optimized to offer new potential efficiencies and higher performance. https://news.ycombinator.com/item?id=34399633 https://news.ycombinator.com/item?id=34399633 : >> AMD ROcm supports Pytorch, TensorFlow, MlOpen, rocBLAS on NVIDIA and AMD GPUs: https://rocmdocs.amd.com/en/latest/Deep_learning/Deep-learning.html https://rocmdocs.amd.com/en/latest/Deep_learning/Deep-learni...
- mozillas 3y agoI ran the 7B Vicuna (ggml-vic7b-q4_0.bin) on a 2017 MacBook Air (8GB RAM) with llama.cpp. Worked OK for me with the default context size. 2048, like you see in most examples was too slow for my taste.
- koheripbal 3y agoGiven the current price (mostly free) off public llms I'm not sure what the use case of running out at home are yet. OpenAIs paid GPT4 has few restrictions and is still cheap. ... Not to mention GPT4 with browsing feature is vastly superior to any home of the models you can run at home.
- toxik 3y agoThe point for me personally is the same as why I find it so powerful to self host SMTP, IMAP, HTTP. It’s in my hands, I know where it all begins and ends. I answer to no one. For LLMs this means I am allowed their full potential. I can generate smut, filth, illegal content of any kind for any reason. It’s for me to decide. It’s empowering, it’s the hacker mindset.
- sagarm 3y agoI think it's mostly useful if you want to do your own fine tuning, or the data you are working with can't be sent to a third party for contractual, legal, or paranoid reasons.
- sroussey 3y agoI’m working on an app to index your life, and having it local is a huge plus for the people I have using it.
- theaussiestew 3y agoSounds interesting, got a link?
- avereveard 3y agoor like download oobabooga/text-generation-webui, any prequantized variant, and be done.