5 ms·
TheBloke doesn’t joke around [1]. I’m guessing we’ll have the quantized ones by the end of the day. I’m super excited to use the 34B Python 4 bit quantized one
by syntaxing 3y ago
TheBloke doesn’t joke around [1]. I’m guessing we’ll have the quantized ones by the end of the day. I’m super excited to use the 34B Python 4 bit quantized one that should just fit on a 3090.
[1] https://huggingface.co/TheBloke/CodeLlama-13B-Python-fp16 https://huggingface.co/TheBloke/CodeLlama-13B-Python-fp16
- suyash 3y agocan it be quantised further so it can run locally on a normal laptop of a developer?
- syntaxing 3y ago“Normal laptop” is kind of hard to gauge but if you have a M series MacBook with 16GB+ RAM, you will be able to run 7B comfortably and 13B but stretching your RAM (cause of the unified RAM) at 4 bit quantization. These go all the way down to 2 bit but I personally I find the model noticeably deteriorate anything below 4 bit. You can see how much (V)RAM you need here [1]. [1] https://github.com/ggerganov/llama.cpp#quantization https://github.com/ggerganov/llama.cpp#quantization
- totallywrong 3y agoJust started playing with this, there's a tool called ollama that runs Llama2 13B on my 16GB M1 Pro really smoothly with zero config.
- stuckinhell 3y agoWhat kind of cpu/gpu power do you need for quantization or these new gguf formats ?
- syntaxing 3y agoI haven’t quantized these myself since TheBloke has been the main provider for all the quantized models. But when I did a 8 bit quantization to see how it compares to the transformers library load_in_8bit 4 months ago(?), it didn’t use my GPU but loaded each shard into the RAM during the conversion. I had an old 4C/8T CPU and the conversion took like 30 mins for a 13B.
- SubiculumCode 3y agoi run llama2 13B models with 4-6 k-quantized oin a 3060 with 12Gb VRam
- selfhoster11 3y agoI can quantize models up to 70B just fine with around 40-50 GB of system RAM, using the GGMLv3 format. GGUF seems not optimised yet, since quantizing with a newer version of llama.cpp supporting the format fails on the same hardware. I expect that to be fixed shortly. For inference, I understand that the hardware requirements will be identical as before.
- mchiang 3y agoOllama supports it already: `ollama run codellama:7b-instruct` https://ollama.ai/blog/run-code-llama-locally https://ollama.ai/blog/run-code-llama-locally More models uploaded as we speak: https://ollama.ai/library/codellama https://ollama.ai/library/codellama
- deleted 3y ago[deleted]
- syntaxing 3y agoWhoa, it’s absolutely astounding how fast the community is reacting to these model release!
- jerrysievert 3y agowhile it supports it, so far I've only managed to get infinite streams of near nonsense from the ollama models (codellama:7b-q4_0 and codellama:latest) my questions were asking how to construct an indexam for postgres in c, how to write an r-tree in javascript, and how to write a binary tree in javascript.
- carbocation 3y agoSimilarly, I had it emit hundreds of blank lines before cancelling it.
- kordlessagain 3y agoWhat fortune, I so happen to need hundreds of blank lines.
- justinsaccount 3y agoMaybe it's outputting https://en.wikipedia.org/wiki/Whitespace_(programming_language) https://en.wikipedia.org/wiki/Whitespace_(programming_langua... :-)
- syntaxing 3y agoSame, just tried it and it would give me infinite amount of blank lines
- UncleOxidant 3y agoIf I don't want to run this locally is it runnable somewhere on huggingface?
- emporas 3y agoReplicate has already hosted Llama2 13B, the chat version. My guess is, in a short span of days or weeks they will host the code version too. They charge a dollar for 2000 generations if i am not mistaken. https://replicate.com/a16z-infra/llama-2-13b-chat https://replicate.com/a16z-infra/llama-2-13b-chat
- deleted 3y ago[deleted]