13 ms·
Quantized Llama models with increased speed and a reduced memory footprint
- newfocogi 2y agoTLDR: Quantized versions of Llama 3.2 1B and 3B models with "competitive accuracy" to the original versions (meaning some degraded performance; plots included in the release notes).
- arnaudsm 2y agoHow do they compare to their original quants on ollama like q4_K_S?
- tcdent 2y agoThese undergo additional fine tuning (QLoRA) using some or all of the original dataset, so they're able to get the weights to align to the nf4 dtype better, which increases the accuracy.
- philipkglass 2y agoThese quantized models show much less degradation compared to a "vanilla post-training-quantization" but there are a bunch of PTQ schemes that people have already applied to Llama models [1]. I didn't see any details about the vanilla PTQ they used as a baseline. Has it been written about elsewhere? [1] https://ollama.com/library/llama3.2/tags https://ollama.com/library/llama3.2/tags
- nisten 2y agoIt's pretty interesting that the new SpinQuant method did not manage to be better than good old nf4bit QLORA training (Tim Dettmers really cooked with that one). Really appreciate that Meta published both results+model quants and didn't just make some bs claim about a new sota quant like most other bigger companies would've done.
- Aeolun 2y agoIt’s a little bizarre that I feel like I’m actually starting to respect this little bit of Meta…
- FuckButtons 2y agoI think meta and facebook before it have always valued a very high standard of engineering, and have also been generally pretty good about open sourcing a lot of that work in a way that allows a lot of people to work with their tools. This doesn’t seem all that out of character.
- ipaddr 2y agoIt's a huge company with a lot of different voices. One may create react and open source it while another would add a clause that if you sue facebook over anything your react license disappears. When they are good they are really good.
- miven 2y agoI mean, it's no free lunch, you still need to expend significantly more compute for the QLoRA training compared to any usual PTQ method, be it SpinQuant or any other more conventional quantization approaches.
- formalsystem 2y agoThe naming is unfortunate but in this blog QLoRA is referring to Quantization-Aware Training with LoRA adaptor
- ipsum2 2y ago
- EliBullockPapa 2y agoAnyone know a nice iOS app to run these locally?
- Arcuru 2y agoI access them by running the models in Ollama (on my own hardware), and then using my app Chaz[1] to access it through my normal Matrix client. [1] - https://github.com/arcuru/chaz https://github.com/arcuru/chaz
- simonw 2y agoMLC Chat is a great iPhone app for running models (it's on Android too) and currently ships with Llama 3.2 3B Instruct - not the version Meta released today, its a quantized version of their previous release. I wouldn't be surprised to see it add the new ones shortly, it's quite actively maintained. https://apps.apple.com/us/app/mlc-chat/id6448482937 https://apps.apple.com/us/app/mlc-chat/id6448482937
- Havoc 2y agoSeems much more stable than the last time I tried it too
- behnamoh 2y agoI've been using PocketGPT.
- drilbo 2y agohttps://github.com/a-ghorbani/pocketpal-ai https://github.com/a-ghorbani/pocketpal-ai This was just recently open sourced and is pretty nice. Only issue I've had is very minor UI stuff (on Android, sounds like it runs better on iOS from skimming comments)
- evbogue 2y agoI'm on Android, however my somewhat elaborate solution was to install Ollama on my home laptop computer and then ssh in when I want to query a model. I figured that'd be better for my phone battery. Since my home computer is behind NAT I run yggdrasil on everything so I can access my AI on the go.
- theanonymousone 2y agoMay I ask if anyone has successfully used 1B and 3B models in production and if yes, in what use cases? I seem to be failing even in seemingly simpler tasks such as word translation or zero-shot classification. For example, they seem to not care about instructions to only write a response and no explanation, thus making it impossible to use them in a pipeline :/
- wswope 2y agoI’ve only toyed with them a bit, and had a similar experience - but did find I got better output by forcing them to adhere to a fixed grammar: https://github.com/ggerganov/llama.cpp/tree/master/grammars https://github.com/ggerganov/llama.cpp/tree/master/grammars For context, I was playing with a script to bulk download podcasts, transcribe with whisper, pass the transcription to llama.cpp to ID ads, then slice the ads out with ffmpeg. I started with the generic json_array example grammar, then iteratively tweaked it.
- accrual 2y agoNot in production, but I've used a 3B model to test a local LLM application I'm working on. I needed a full end-to-end request/response and it's a lot faster asking a 3B model than an 8B model. I could setup a test harness and replay the responses... but this was a lot simpler.
- jdthedisciple 2y agoIf for testing then why not just mock the whole thing for ultimate performance ... ?
- nkozyra 2y agoProbably faster to use off the shelf model with llama.cpp than to mock it
- com2kid 2y ago3B models are perfectly capable, I've had great luck with Phi 3.5. > For example, they seem to not care about instructions to only write a response and no explanation You need to use tools to force the model to adhere to a schema. Or you can learn to parse out the part of the response you want, both work. You'll also need to make good use of robust examples in your initial prompt, and give lots of examples of how you want the output to look. (Yes this quickly burns up the limited context length!) Finally, embrace the fact that these models are tuned for chat, so the more conversational you make the back and forth the less you are stretching the models abilities. I wrote a very small blog post at https://meanderingthoughts.hashnode.dev/unlock-the-full-potential-of-llms-ditch-json-for-dsls https://meanderingthoughts.hashnode.dev/unlock-the-full-pote... explaining some of this.
- ngamboa 2y ago[dead]
- mmaunder 2y ago[flagged]
- pryelluw 2y agoI don’t get the comment. For one I’m excited for developments in the field. Not afraid it will “replace me” as technology has replaced me multiple times over. I’m looking towards working with these models more and more.
- mmaunder 2y agoNo, I meant that a lot of us are working very fast on a pre-launch product, implementing some cutting edge ideas using e.g. the incredible speedup in a small fast inference model like quantized 3B in combination with other tools, and I think there's quite a bit of paranoia out there that someone else will beat you to market. And so not a lot of sharing going on in the comments. At least not as much as previously, and not as much technical discussion vs other non-AI threads on HN.
- mattgreenrocks 2y agoThis thread attracts a smaller audience than, say, a new version of ChatGPT.
- pryelluw 2y agoOk, thank you for pointing that out. I’m focused on making models play nice with each other rather than building a feature that relies on it. That’s where I see the more relevant work being. Why such news are exciting!
- accrual 2y agoTwo days ago there was a pretty big discussion on this topic: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku https://news.ycombinator.com/item?id=41914989 1421 points, 717 comments
- flawn 2y ago
- behnamoh 2y agoDoes anyone know why the most common method to speed up inference time is quantization? I keep hearing about all sorts of new methods but nearly none of them is implemented in practice (except for flash attention).
- o11c 2y agoBecause the way LLMs work is more-or-less "for every token, read the entire matrix from memory and do math on it". Math is fast, so if you manage to use only half the bits to store each item in the matrix, you only have to do half as much work. Of course, sometimes those least-significant-bits were relied-upon in the original training.
- slimsag 2y agoHas anyone worked on making tokens 'clusters of words with specific semantic meaning'? e.g. instead of tokens ['i', 'am', 'beautiful'] having tokens ['I am', 'beautiful'] on the premise that 'I am' is a common set of bytes for a semantic token that identifies a 'property of self'? Or taking that further and having much larger tokens based on statistical analysis of common phrases of ~5 words or such?
- dragonwriter 2y agoMuch larger tokens require a much larger token vocabulary.
- visarga 2y agoyes, look up Byte Pair Encoding https://huggingface.co/learn/nlp-course/chapter6/5 https://huggingface.co/learn/nlp-course/chapter6/5
- pizza 2y agoI think you might be thinking of applying a kind of low-rank decomposition to the vocabulary embeddings. A quick search on Google Scholar suggests that this might be useful in the context of multilingual tokenization.
- deleted 2y ago[deleted]
- justanotheratom 2y agoAny pointers no how to finetune this on my dataset and package and run it in my swift ios app?
- tveita 2y agoSo SpinQuant learns a rotation for activations and weights that, to my understanding, "smear" the outlier weights out so you don't get extreme values in any one weight. Random anecdote warning - In the old days, before vector search became AI and everyone and their dog offered a vector database, I had a task that required nearest neighbour search in a decent amount of high-dimensional vectors. I tried quantizing them to bit vectors in an index and scanning through it to get an initial set of candidates. Performance was actually quite decent - reading through RAM linearly is fast! But the selectivity wasn't great. Somewhere along the way I found this paper[1] that iteratively finds a rotation to apply before quantization to reduce the quantization error. Very similar goal to SpinQuant, but focused on bit quantization only. As it turns out the 'random rotation' baseline they benchmark against worked great for my use case, so I never tried implementing the fancier algorithm. But it's a pretty rare day at work that "apply a random rotation matrix to a 128-dimensional vector" is the solution to my problem. [1] https://ieeexplore.ieee.org/abstract/document/6296665 https://ieeexplore.ieee.org/abstract/document/6296665 / https://slazebni.cs.illinois.edu/publications/ITQ.pdf https://slazebni.cs.illinois.edu/publications/ITQ.pdf
- arijo 2y agoI find the geometrical intuition of rotating a vector in high dimensional space to minimize its largest values (vector basis projections) beautiful. I'm no expert and I'm sure this has been tried by many people already - but would it be possible to reduce the computational effort instead by using SVD decomposition, spreading the singular values and then reapplying the original singular values and recomposing the matrix using the quantized versions of the SVD matrices?
- govg 2y agoTangentially related to the idea of "apply a random rotation matrix" is one where you apply a random matrix to a set of points to preserve distances between them but transform them into a lower dimensional space. This is captured by the JL Lemma [1]. [1] - https://en.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_lemma https://en.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_...
- derefr 2y ago
- ed 2y agoOh cool! I’ve been playing with quantized llama 3B for the last week. (4-bit spinquant). The code for spinquant has been public for a bit. It’s pretty adept at most natural language tasks (“summarize this”) and performance on iPhone is usable. It’s even decent at tool once you get the chat template right. But it struggles with json and html syntax (correctly escaping characters), and isn’t great at planning, which makes it a bad fit for most agenetic uses. My plan was to let llama communicate with more advanced AI’s, using natural language to offload tool use to them, but very quickly llama goes rogue and starts doing things you didn’t ask it to, like trying to delete data. Still - the progress Meta has made here is incredible and it seems we’ll have capable on-device agents in the next generation or two.
- tucnak 2y ago>But it struggles with json You should customise your sampler to mandate JSON grammar after ```json tokens.
- ed 2y agoGrammar samplers are clever! But in the case of a missing escape character you’ll end up with a corrupted string. Take for example: "A dog says \"Woof!\"" With a grammar, you’ll end up with "A dog says " when the model forgets to escape. Which is valid JSON, but not what the model intended. So it’s usually better to catch the exception and ask the model to try again. Unless you’ve come across a sampler with backtracking? That would be cool
- formalsystem 2y agoHi I'm Mark I work on torchao which was used for the quantization aware training and ARM kernels in this blog. If you have any questions about quantization or performance more generally feel free to let me know!
- philipkglass 2y agoWhat was the "vanilla post-training quantization" used for comparison? There are 22 GGUF quantization variants smaller than 16 bits per weight and I can't tell which one is being compared with: https://huggingface.co/docs/hub/en/gguf#quantization-types https://huggingface.co/docs/hub/en/gguf#quantization-types It might even mean a non-GGUF quantization scheme; I'm just an intermediate user of local models, not an expert user or developer.
- formalsystem 2y agoSo this should be referring to w8a8 (weights and activations in 8 bit) So this is gonna be 8 bit weights, 8 bit activations, group size of 256, symmetric quantization. Not sure how to map this to the GGUF variants because they don't mention how they don't do activation quantization
- imjonse 2y agoWere there comparisons made to AWS, Smoothquant, GPTQ or other non-vanilla PTQ methods? Thanks.
- formalsystem 2y agoNot that I know of for this study, at least for the specific scope torchao we want to make it easier for researchers to create new quantization algorithms in python and have those algorithms run fast and you can see a lot of those algorithms here https://github.com/pytorch/ao/tree/main/torchao/prototype https://github.com/pytorch/ao/tree/main/torchao/prototype So for example for AWQ and GPTQ we can accelerate them by using a fast int4 kernel called tinygemm
- nikolayasdf123 2y agowhat's your opinion on LlamaStack? for me it is nothing short of bad experience. it is way over-engineered with poor quality and just plain does not work, and maintainers are questionable. I would rather call HuggingFace python code for inference or anything else. is ExecuTorch any better?
- SoLoMo123 2y agoHi, I'm Mergen and I work on ExecuTorch. ExecuTorch is a runtime for mobile and embedded devices to run PyTorch models directly. Currently it runs pretty fast on CPU, but expanding our use-case for mobile accelerators and GPUs. We're still in our early stages (just turned beta status). But try it out and let us know. Regarding Llama Stack, it is built by my colleagues. What were some concrete issues have you experienced? If you have error/bug reports, I'll happy to pass along.
- nikolayasdf123 2y agowill give executorch a try. with llamastack, well making it work with CUDA for starters would be great. it is also bloated. something that supposed to take direct 100 lines of code and a couple files, takes dozens of files, multiple frameworks, generators.. which in the end do not work at all, and nobody knows why. very obscure framework. can't believe this code is coming from Meta.
- Evidlo 2y agoWhy don't they actually say what the size of the model is in GB? That and average inference times on common hardware is what I'm curious about.
- Ardren 2y agoThe last table shows memory usage and performance on an Android phone. > Decode latency improved by 2.5x and prefill latency improved by 4.2x on average, while model size decreased by 56% and memory usage reduced by 41% on average. The benchmarks can be reproducible today via ExecuTorch Llama instructions. The table above shows results using an Android OnePlus 12 device—however, we’ve also verified similar relative performance on Samsung S24+ for 1B and 3B and Samsung S22 for 1B.
- Tepix 2y agoFrom TFA: > At Connect 2024 last month, we open sourced Llama 3.2 1B and 3B No you did not. There is no source (in this case: training data) included. Stop changing the meaning of "open source", Meta!
- cmsj 2y agoIt really bugs me that every time I see posts about new models, there is never any indication of how much VRAM one needs to actually run them.
- qeternity 2y agoThat's because it's easily calculable and also somewhat impossible to say in any meaningful sense. Most weights are released as fp16/bf16 so 2 bytes per weight. So just double the number of parameters = the number of gigabytes of VRAM. Llama 3.1 8B ~= 16GB weights in fp16. At 4bit quantization, it would be half the number of parameters so Llama 3.1 8B ~= 4GB weights. But this is just weights. The real issue is context and output length: how much data are you feeding in? This is where VRAM can explode, and it's entirely use-case dependent. So for a 128k context model, the range of VRAM usage is huge. The reality is, if you're not able to quickly estimate the above, you're probably not running local models anyway.
- bick_nyers 2y agoPerhaps I'm being charitable but I read OP's comment in the light of what you described with context length. Batching, context length, and attention implementation vary these numbers wildly. I can fit a 6bit quant Mistral Small (22b) on a 3090 with ~10-12k context, but Qwen2VL (7b, well 8.3b if you include vision encoder) also maxes out my 3090 VRAM with an 8bit quant and ~16k context. I do think it would be good to include some info. on "what we expect to be common deployment scenarios, and here's some sample VRAM values". Tangentially, whenever these models get released with fine-tuning scripts (FFT and Lora) I've yet to find a model that provides accurate information on the actual amount of VRAM required to train the model. Often times it's always 8x80GB for FFT, even for a 7B model, but you can tweak the batch sizes and DeepSpeed config. to drop that down to 4x80GB, then with some tricks (8bit Adam, Activation Checkpointing), drop it down to 2x80GB.
- formalsystem 2y agoYou can estimate context length impact by doing back of the envelope calculations on KV cache size: 2 * layers * attention heads * head_dim * byte_per_element * batch_size * sequence_length Some pretty charts here https://github.com/pytorch/ao/issues/539 https://github.com/pytorch/ao/issues/539
- yuvalr1 2y agoLooking at how to deploy 1B and 3B Llama models on Android for inference. Some posts online recommend using Termux (an amazing app) to have an emulated shell and then install as if it's Linux, using ollama for example. However, this forces you into a manual installation process, and also most of the people don't know what Termux is, and would be afraid to install it from F-Droid. Maybe someone can recommend a way to deploy Llama to Android without Termux, maybe even something that can be potentially fully implemented inside an app? I'm currently looking into compiling llama.cpp for Android and bundling it inside an app. Is that a viable path? Would love to hear from someone who tried something similar.
- tugdual 2y agoI actually did something similar using llama.cpp a while back, would be curious to see the speedup with this model. https://github.com/TugdualKerjan/bunny/tree/main https://github.com/TugdualKerjan/bunny/tree/main
- antonvs 2y agoThis might be of use: https://github.com/a-ghorbani/pocketpal-ai https://github.com/a-ghorbani/pocketpal-ai
- deleted 2y ago[deleted]
- niutech 2y agoYou can use MLC LLM: https://llm.mlc.ai/ https://llm.mlc.ai/
- itsTyrion 2y agoWait, so I can get incorrect information and text summaries with things added or cut off even faster and on mobile now? that's amazing.