34 ms·
LLaMA now goes faster on CPUs
- llm_trw 3y ago[flagged]
- KnightHawk3 3y agodid a search and I have no idea what you are talking about
- llm_trw 3y agohttps://github.com/ggerganov/llama.cpp/issues/91 https://github.com/ggerganov/llama.cpp/issues/91 Justine was kicked out of llama.cpp for introducing changes before the rest of the maintainers approved them. Much drama has been had since then and I've lost interest in both projects. It's just been exhausting wanting to build good software without ego in this space. Everyone is trying to get rich quick before the inevitable AI ice age starts and all these skills are again useless.
- Rexxar 3y agoThe PR associated with the issue was merged after the approval of the owner of the repo (https://github.com/ggerganov/llama.cpp/pull/613 https://github.com/ggerganov/llama.cpp/pull/613). And there is other more recent contributions. Where do you see any drama ?
- llm_trw 3y agohttps://news.ycombinator.com/item?id=35411909#35412589 https://news.ycombinator.com/item?id=35411909#35412589
- cryptonector 3y agoFinally a useful link in demonstrating all the drama.
- vidarh 3y agoThe last commit from Justine I can see to the llama.cpp repo is a week ago, so whatever drama there was appear to have been at a minimum partially resolved.
- cryptonector 3y ago> https://github.com/ggerganov/llama.cpp/issues/91 https://github.com/ggerganov/llama.cpp/issues/91 > Justine was kicked out of llama.cpp for introducing changes before the rest of the maintainers approved them. I followed your link and found nothing regarding this. What am I missing? Seems to me that you're just casting aspersions w/o backup.
- sneak 3y agoEvery release a perf improvement?
- llm_trw 3y agoCount how many times I and me were said in the article. I have a nice grease monkey script that blocks both domains on my old machine. I guess I forgot to import it on this one.
- deleted 3y ago[deleted]
- mlyle 3y agoSomeone doing performance work can state they did it. It's a public service we all benefit from. If the rivalry further adds urgency to improve: great.
- jrflowers 3y agoI feel your pain here. I hate it when I am forced to post online about a website I don’t like because I’ve both forgotten to stop myself from being able to look at it and forgotten not to click on it.
- deleted 3y ago[deleted]
- bottlepalm 3y agoI think it's a good idea for everyone to download and be able to run a LLM locally, even if you have the minimum of requirements. As a pseudo-backup of a large chunk of human knowledge.
- TaylorAlexander 3y agoI contend that most human knowledge is not written down or if it is written down it’s not publicly available on the internet and so does not exist in these datasets. There’s so much subtle knowledge like the way a mother learns to calm her child or the way a carpenter learns to work different kinds of wood which may be written down in part, but may also be learned through lived experience or transferred from human to human such that little of it gets written down and posted online.
- mickdarling 3y agoWait till all the videos ever created are tokenized and ingested into a training dataset. Carpentry techniques are certainly there. The subtleties of parenting maybe harder to derive from that, but maybe lots of little snippets of people’s lives will add up to a general understanding of parenting. There have certainly been bigger surprises in the field.
- oblio 3y agoWhat about smells or tastes? Or feelings? I can't help but feel we're at the "aliens watch people eat from space and recreate chemically identical food that has no taste" phase of AI development.
- skeledrew 3y agoIf the food is chemically identical then the taste would be the same though, since taste (and smell) is about chemistry. I do get what you're saying though.
- 3y ago
- deleted 3y ago[deleted]
- 1-6 3y agoQuestion is, how much of an improvement has it gotten to over a GPU or ASIC?
- dartos 3y agoNothing in software will ever beat an equivalent ASIC.
- postalrat 3y agoSure there is. Software is easy to change.
- dartos 3y agoBy “beat” I meant in performance. Obviously you can’t change an asic
- fragmede 3y agoan asic is fixed function, so it'll never be able to boot my pc and then be the CPU, even though an asic beats the pants off anything else computing Sha hashes for Bitcoin mining.
- dartos 3y agoBy “beat” I meant performance. Obviously an ASIC is not a general purpose machine like a cpu.
- fulafel 3y agoMost ASICs are cost or power optimizations.
- dartos 3y agoExactly. They’re much faster for their specific tasks and thus are more power efficient and potentially cost efficient
- discordance 3y ago"As for disk speed, dd if=/dev/zero of=/tmp/output bs=128k count=50k; rm -f /tmp/output reports 1.6 GB/s which is 3.6x slower than my Mac Studio, and 3x slower than my Intel (which has the same M.2 stick). I'm told that Intel and Apple are just better at this, but I wish I understood why. " Can anyone here answer why this is?
- pstrateman 3y agoApple made fsync a noop. You have to make a different call to get sync on macos. So tons is stuff is faster because it's not actually writing to disk.
- bishfish 3y agoPlus he isn’t using oflag=direct, so since output file is small it isn’t even making it to disk. I think it would only be sent to page cache. I’m afraid he is testing CPU and memory (bus) speeds here. oflag=direct will write direct and bypass page cache.
- pama 3y agoSuper nice story on the matmul optimization that gave 810 gflops for 512x512. Thanks for the write up and the contributions to llama.cpp and the community more broadly.
- kiratp 3y agoIt fascinating to me that coming up on a year since Sapphire Rapids has been available in the public cloud, developers are still targeting AVX512 when they should be targeting VNNI and AMX. https://github.com/ggerganov/llama.cpp/issues/2555 https://github.com/ggerganov/llama.cpp/issues/2555
- yjftsjthsd-h 3y agoThis project in particular seems to care about the long tail of hardware; note that the very first machine in this post is a box from 2020 with spinning rust disk. Granted, adding support for newer extensions is likely also good, but cost/benefit is in play.
- taneq 3y agoIs four years really 'long tail' these days? Our VM host box is from 2010 (and I had to rebuild llama.cpp locally without AVX to get it working :P )
- yjftsjthsd-h 3y agoFor cutting-edge LLM work, probably? I mean, I run mine on older hardware than that, but I'm a total hobbyist...
- d416 3y agoIt should be noted that while the HP Prodesk was released in 2020, the CPU’s Skylake architecture was designed in 2014. Architecture is a significant factor in this style of engineering gymnastics to squeeze the most out of silicon.
- refulgentis 3y agoFor LLMs...yeah. I imagine you're measuring in tokens/minute with that setup. So its possible, but...do you use it much? :)
- luyu_wu 3y agoI don't believe that is the target for a local LLM... Pretty sure we're talking about client-side computing, of which the newest supports only AVX-512 (and even that sketchily on Intel's side).
- aniijbod 3y agoA way of thinking about what's inside any of the top LLMs right now: even if they never learn another single fact, even if they get ridiculously out of date as a result, even if they are even more riddled with errors and prone to biases than we know them to be, even if they are as prone to hallucinations as we know they they are and they never develop the capacity to cure themselves of this, they are more knowledgeable and capable of more reasoned response, despite their capacity for error, to more questions than any single human being that has ever lived.
- JKCalhoun 3y agoPicturing "LLM Jeopardy". You know, a game show.
- samus 3y agoWe shouldn't choose LLMs for how many facts they support, but their capability to process human language. There is some overlap between these two though, but an LLM that just doesn't know something can always be augmented with RAG capabilities.
- talldayo 3y agoIf you ignore my capacity for error, I bet I'd put up a good score too. Hell, maybe Markov chains are smarter than LLMs by this definition.
- ajtulloch 3y ago- https://www.cs.utexas.edu/users/flame/laff/pfhp/index.html https://www.cs.utexas.edu/users/flame/laff/pfhp/index.html (e.g. here https://www.cs.utexas.edu/users/flame/laff/pfhp/week2-blocking-for-registers.html https://www.cs.utexas.edu/users/flame/laff/pfhp/week2-blocki...) - https://gist.github.com/nadavrot/5b35d44e8ba3dd718e595e40184d03f0 https://gist.github.com/nadavrot/5b35d44e8ba3dd718e595e40184... might be of interest
- kpw94 3y agoGreat links, especially last one referencing the Goto paper: https://www.cs.utexas.edu/users/pingali/CS378/2008sp/papers/gotoPaper.pdf https://www.cs.utexas.edu/users/pingali/CS378/2008sp/papers/... >> I believe the trick with CPU math kernels is exploiting instruction level parallelism with fewer memory references It's the collection of tricks to minimize all sort of cache misses (L1, L2, TLB, page miss etc), improve register reuse, leverage SIMD instructions, transpose one of the matrices if it provides better spatial locality, etc.
- larodi 3y agoThe trick is indeed to somehow imagine how the CPU works with the Lx caches and keep as much info in them as possible. So its not only about exploiting fancy instructions, but also thinking in engineering terms. Most of the software written in higher level langs cannot effectively use L1/L2 and thus results in this constant slowing down otherwise similarly (from asymptotic analysis perspective) complexity algos.
- wokwokwok 3y ago> You don't need a large computer to run a large language model While running tiny llama does indeed count as running a language model, I’m skeptical that the capabilities of doing so match what most people would consider a baseline requirement to be useful. Running 10 param model is also “technically” running an LM, and I can do it by hand with a piece of paper. That doesn’t mean “you don’t need a computer to run an LM”… I’m not sure where LM becomes LLM, but… I personally think it’s more about capability than parameter count. I don’t realllly believe you can do a lot of useful LLM work on a pi
- mlyle 3y agoTinyllama isn't going to be doing what ChatGPT does, but it still beats the pants off what we had for completion or sentiment analysis 5 years ago. And now a Pi can run it decently fast.
- jerrygenser 3y agoYou can fine-tune a 60mm parameter (e.g. distilBERT) discriminative (not generative) language model and it's one or two order of magnitude more efficient for classification tasks like sentiment analysis, and probably similar if not more accurate.
- mlyle 3y agoYup, I'm not saying TinyLLAMA is minimal, efficient, etc (indeed, that is just saying that you can take models even smaller). And a whole lot of what we just throw LLMs at is not the right tool for the job, but it's expedient and surprisingly works.
- anentropic 3y agoit seems that BERT can be run on the llama.cpp platform https://github.com/ggerganov/llama.cpp/pull/5423 https://github.com/ggerganov/llama.cpp/pull/5423 so presumably those models could benefit from the speed ups described in OP article when running on CPU
- none_to_remain 3y agoFrom the example: "--temp 0 turns off the random number generator (we don't want improvisation for a spam filter)" I've been thinking for a while about how many applications of LLMs need this adjustment and aren't getting it
- mvkel 3y agoIs that what it does, though? I thought setting temperature to 0 would (extremely simple example) equate to a spam filter seeing: - this is a spam email But if the sender adapts and says - th1s is a spam email It wouldn't be flagged as spam.
- none_to_remain 3y agoMy understanding is that temperature applies to the output side and allows for some randomness in the next predicted token. Here Justine has constrained the machine to start with either "yes" or "no" and to predict only one token. This makes the issue stark: leaving a non-zero temperature here would just add a chance of flipping a boolean.
- refulgentis 3y agoIt's more nuanced than that, in practice: this is true for the shims you see from API providers (ex. OpenAI, Anthropic, Mistral). With llama.cpp, it's actually not a great idea to have temperature purely at 0: in practice, especially with smaller models, this leads to pure repeating or nonsense. I can't remember where I picked this up, but, a few years back, without _some_ randomness, the next likely token was always the last token.
- samus 3y agoThe output of an autoregressive model is a probability for each token to appear next after the input sequence. Computing these is strictly deterministic from the prior context and the model's weights. Based on that probability distribution, a variety of text generation strategies are possible. The simplest (greedy decoding) is picking the token with the highest probability. To allow creativity, a random number generator is used to choose among the possible outputs, biased by the probabilities of course. Temperature scales the output probabilities. As temperature increases, the probabilities approach 1/dictionary size, and the output becomes completely random. For very small temperature values, text generation approaches greedy sampling. If all you want is a spam filter, better replace the output layer of an LLM with one with just two outputs, and finetune that on a public collection of spam mails and some "ham" from your inbox.
- bee_rider 3y agoIs it easy to find where the matvecs are, in LLaMA (if you are someone who is curious and wants to poke around at the “engine” without understanding the “transmission,” so to speak)? I was hoping to mess around with this for Stable Diffusion, but it seemed like they were buried under quite a few layers of indirection. Which is entirely reasonable, the goal is to ship software, not satisfy people who’d just want to poke things and see what happens, haha.
- fragmede 3y agodid you see tiny grad can run llama and stable diffusion? it's an intentionally extremely simple framework vs pytorch or even micrograd, which helped me dig into the underlying math. though https://spreadsheets-are-all-you-need.ai/ https://spreadsheets-are-all-you-need.ai/ is a good one for learning LLMs.
- Ono-Sendai 3y agoMultithreading support in llama.cpp is probably still pretty busted, assuming it uses the same underlying NN inference code as whisper.cpp: https://github.com/ggerganov/whisper.cpp/issues/200#issuecomment-1484025515 https://github.com/ggerganov/whisper.cpp/issues/200#issuecom...
- imtringued 3y agoFrom what I have heard they use manual spin locks. Generally, spin locks are not a good idea unless you want to dedicate the entire machine to a single application. If the process a spinlock waits on gets suspended, you're burning CPU time for nothing. The OS thinks a spinlock making zero progress is actually a high priority process, so it is starving the suspended process from making progress.
- Ono-Sendai 3y agoYeah the code looks like a spinlock. It behaves terribly under contention, resulting in performance falling off a cliff as the number of threads increases. Adding more threads actually slows down the total performance. I would fix it if I could be bothered. Instead I will just use the Cuda whisper backend which is pretty nice and fast.
- jongjong 3y agoThat's interesting because I built a simple ANN library and I was playing around with GPU acceleration and came to a similar conclusion as this article. To be fair, my ANN library was faster (up to 2x) with GPU acceleration in some scenarios were ANN was shallow (as opposed to deep with many hidden layers). I thought the marginal gain may have been because, the way it's set up in my library, it has to load all the values into the GPU from RAM for each pass of forward and back propagation in each layer during training. I believe there is a way to allocate memory on the GPU chip itself but it's a lot more challenging to do, especially in a modular, fully portable way (which was one of the goals of my library). But anyway, even the 2x best-case figure seemed disappointing. In my mind, I expected to see at least 10x speed improvement... And I was surprised that the CPU version was actually slightly faster in the scenario I was testing at the time which was a relatively deep network. It makes sense since the different layers cannot be parallelized as the input of one layer depends on the output of the previous layer... So the more layers you have, the more serial bottlenecks you have, the less you can benefit from GPU acceleration... And unfortunately, deep networks also happen to be those which tend to perform best for a lot of use cases.
- kristianp 3y agoNice to see such speedups for CPUs. Are these changes available as a branch or pull request in llama.cpp itself? I'd like to make use of them in that form if possible (as I'm used to using that).
- dagaci 3y agoYes, this is really a phenomenal effort! And what open source is about: Bringing improvements to so many use cases. So that Intel and AMD chip uses can start to perform while taking advantage of their high-performance capabilities, making even old parts competitive. There are two PRs raised to merge to llama.cpp: https://github.com/ggerganov/llama.cpp/pull/6414 https://github.com/ggerganov/llama.cpp/pull/6414 https://github.com/ggerganov/llama.cpp/pull/6412 https://github.com/ggerganov/llama.cpp/pull/6412 Hopefully these can be accepted, without drama! as there are many downstream dependencies on llama.cpp can will also benefit. Though of course everyone should also look directly at releases from llamafile https://github.com/mozilla-Ocho/llamafile https://github.com/mozilla-Ocho/llamafile.
- wtallis 3y agoI know this post is focused specifically on CPU performance, but the section on the performance on the Mac Studio seems to be deliberately avoiding directly mentioning that machine's GPU, let alone benchmark against it. I think it would have been interesting to see a straightforward comparison of what compute performance and memory bandwidth (as measured by the prompt processing and token generation speeds, respectively) are achievable with reasonable optimization effort on the CPU vs GPU when they're attached to the same memory subsystem.
- politelemon 3y agoThis is great work. I've always thought it would be great if running LLM could be commoditized for regular average Joe hardware. I had thought that llamafile was like dockerfile for llama.cpp but looks like that's a mistake? Will definitely be giving this a try.
- seangrogg 3y agoMmm, I wonder how well this would work on a mobile device. Maybe I'll try grabbing my ubuntu touch here in a sec...
- seangrogg 3y ago(For any who were curious: it does not for memory reasons)
- speps 3y agoRegarding this bit at the end: > I learned how to write math kernels by renting Vast VMs and watching Gautham Venkatasubramanian and mrdomino develop CUDA kernels in a tmux session. They've been focusing on solving a much more important challenge for llamafile, which is helping it not have a mandatory dependency on the cuBLAS If I'm reading this right, they're trying to rewrite cuBLAS within CUDA itself. I'm guessing the next step would be removing CUDA dependency and go with directly using Vulkan or Metal compute shaders. Am I correct?
- WithinReason 3y agoYes, but none of these have performance portability across GPU vendors, so it's probably seen as pointless. You would need an AMD Vulkan shader, an nvidia one, and intel one, etc. It's not like C code on CPUs.
- surge 3y agoMaybe its a dumb question, but isn't something like OpenCL meant to solve this problem?
- jvanderbot 3y agoFrom my understanding, using triangle / shaders to do HPC has given way to a specific, more general purpose GPU programming paradigm which is CUDA. Of course this knowledge is superficial and probably outdated, but if I'm not too far off base, it's probably more work to translate a general CUDA-like layer or CUDA libs to OpenCL.
- slackito 3y agoThe fact that you're comparing CUDA to using triangles and shaders makes me think you might be confusing OpenCL with OpenGL. OpenCL is meant for general computation (the C is for "computing") rather than graphics, like CUDA.
- VHRanger 3y ago
- pknerd 3y agoSo, I can now run it on my 2015 Macbook with 8GB RAM?
- isusmelj 3y agoIs there somewhere an overview of the progress we made on the software side for training and inference of LLMs? It feels like we squeezed 10-100x more out of the hardware since llama appeared. This crazy progress will probably saturate though as we reach theoretical limits, no?
- mijoharas 3y agoHas Justine written anywhere about her disassembly setup? > I configured Emacs so I can push a button, and the disassembly for the C++ code I'm working on will pop up on the screen in a few milliseconds. I assume it's something project specific rather than being able to get the disassembly for an arbitrary section of code or something? It seems very handy, so I'd love to see the implementation (I couldn't find anything googling)
- pelletier 3y agoThis is probably what they are referring to https://github.com/jart/disaster https://github.com/jart/disaster
- moffkalast 3y ago> the Raspberry Pi Odd how there were no Mistral 7 benchmarks for the Pi 5 in that table (I doubt anyone is seriously considering using TinyLlama for anything at all), so I went to re-test it out myself on the Pi 5 8G. llamafile 0.7: 52 predicted, 150 cached, 430ms per token, 2.32 tokens per second llama.cpp + OpenBLAS: 36 predicted, 124 cached, 381ms per token, 2.62 tokens per second It does seem to inch closer to the speed you get with blas acceleration which is quite impressive, but in practical terms the Pi 5 is so heavily limited by its memory throughput bottleneck that it saturates the required compute with 3 threads already. So while fancy kernels will make it more efficient it won't really save you from that fundamental bandwidth limit. The Pi foundation messed up going with a 32 bit memory bus, simple as.
- 6r17 3y agotoday being today ; I must ask ; anyone has actually tried this ?
- tomp 3y agoTL;DR: unroll the outer two loops of matrix multiplication
- amelius 3y agoShouldn't this have been done in a library instead of a specific project? Then others could also profit from it.
- AbuAssar 3y agoregarding AMD zen4 with avx512: "Here we see that, despite only being twice the price, the 7995WX x86 ISA offers 7x more raw compute power than the M2 Ultra ARM ISA, and nearly the same token generation speed, which is likely thanks to its 384mb L3 cache. When I bought this chip, I had to expand support in llama.cpp for bfloat16 and AVX512 before I could fully test its capabilities. My work means you can now run LLaMA 2.8x faster on Zen4 than you could before."
- reckless 3y agoDoes this also count platform costs or just chip cost? I'd imagine the threadripper motherboard and ram costs aren't insignificant
- KennyBlanken 3y agoA complete desktop computer with the M2 Ultra w/64GB of RAM and 1TB of SSD is $4k. The 7995WX processor alone is $10k, the motherboard is one grand, the RAM is another $300. So you're up to $11300, and you still don't have a PSU, case, SSD, GPU....or heatsink that can handle the 300W TDP of the threadripper processor; you're probably looking at a very large AIO radiator to keep it cool enough to get its quoted performance. So you're probably up past $12k, 3x the price of the Studio...more like $14k if you want to have a GPU of similar capability to the M2 Ultra. Just the usual "aPPle cOMpuTeRs aRE EXpeNsIVE!" nonsense.
- incrudible 3y agoSo from a CPU perspective you get 7x the CPU throughput for 3x to 4x the price, plus upgradable RAM that is massively cheaper. The M2 uses the GPU for LLMs though, and there it sits in a weird spot where 64GB of (slower) RAM plus midrange GPU performance is not something that exists in the PC space. The closest thing would probably be a (faster) 48GB Quadro RTX which is in the $5000 ballpark. For other use cases where VRAM is not such a limiting factor, the comparably priced PC will blow the Mac out of the water, especially when it comes to GPU performance. The only reason we do not have cheap 96GB GDDR GPUs is that it would cannibalize NVIDIA/AMDs high margin segment. If this was something that affected Apple, they would act the same.
- aimonster2 3y agoPosted too early.
- sublimefire 3y agore:funding my friend suggested to nominate Justine for the open source contributions in an internal Microsoft programme (the winner takes $10k). They did not even want to add her to the potential list of nominees because her software is not used in MSFT. It speaks volumes about the corporate culture and shows what they really think about OSS support.
- deleted 3y ago[deleted]
- miki123211 3y agoIf I'm reading the post correctly, Llamafile is faster than llama.cpp, despite the author upstreaming some of the changes. What's the reason for this?
- tiffanyh 3y agoPixar uses CPUs … I wonder if we’ll end up in a situation like rendered movies. Where the big studios like Pixar uses CPUs (not GPUs) to render their movies due to the cost/perf (and access to larger amounts of RAM). https://news.ycombinator.com/item?id=25616372 https://news.ycombinator.com/item?id=25616372
- kreco 3y ago> Where the big studios like Pixar uses CPUs (not GPUs) to render their movies due to the cost/perf (and access to larger amounts of RAM). I wonder if (or when) this will change once integrated GPUs become "mainstream", the CPU/GPU share the same RAM AFAIK.
- rockwotj 3y agoI expect GPU hardware to specialize like Google’s TPU. The TPU feels like ARM in these AI workloads where when you start to run these at scale, you’ll care about the cost perf tradeoff for most usecases. > CPU/GPU share the same RAM AFAIK. This depends on the GPU I believe Apple has integrated memory, but most GPUs from my limited experience writing kernels have their own memory. CUDA pretty heavily has a device memory vs host memory abstraction.
- talldayo 3y agoOn top of that, Nvidia has provided a unified addressing abstraction over PCI for a looooong time via CUDA: https://developer.nvidia.com/blog/unified-memory-in-cuda-6/ https://developer.nvidia.com/blog/unified-memory-in-cuda-6/ Customers like Pixar could probably push this even further, with a more recent Nvidia rack and Mellanox networking. Networking a couple Mac Studios over Thunderbolt doesn't have a hope of competing, at that scale.
- CaptainOfCoit 3y agoI'm not sure how true that is anymore, from the outside it seems they're at least moving to a CPU/GPU hybrid (which makes a lot of sense), at least judging by new features landing in RenderMan that continues to add more support for GPUs (like XPU).
- 4bpp 3y agoIt would be good to see some independent verification of this claim. HN has previously [1] fallen for a claim by the same author to have reduced llama.cpp memory usage for a dense model way below the size of the model, which should have failed a basic smell test and indeed was debunked shortly after. Justine Tunney appears to enjoy extreme superstar status here, and it's hard to overstate the degree of social pressure that needed to be overcome at the time for the skeptic position to reach fixation (to begin with, what other LLM developments even hit upvote numbers like the +1300ish there or the +712 here at the time of writing?). [1] https://news.ycombinator.com/item?id=35393284 https://news.ycombinator.com/item?id=35393284
- freedomben 3y ago> Justine Tunney appears to enjoy extreme superstar status here This is true, and for sure pretty much all humans can benefit from increased skepticism (though not cynicism), but that superstar status is achieved from numerous impressive works. Cosmopolitan C and Actually Portable Executable were some of the things in the past that alone were worthy of significant respect, and for many people (like myself) these were our first introduction. Speaking only for myself, I have a high opinion of Justine on technical merits. I'm sure she makes mistakes like all humans. I can tell she gets excited by discoveries and the chase, and that probably does sometimes cause premature celebration (this is something I struggle with so it's recognizable to me haha), but being wrong sometimes doesn't erase when you're right, and she has been spectacularly right a lot more times than most people I know. There have been some personality clashes between Justine and others at times, and unfortunately it's situations where only part (sometimes a small part) of it was public, meaning we can only take people's word for what happened. Given my ignorance, I choose to withhold judgment here, but even if I didn't (and assumed she was guilty) it doesn't change the technical merits and it certainly wouldn't dissuade me from seeing what she's working on now. So when I see stuff from Justine come out like this, it gets my attention. Would it get my attention if the same thing were posted by somebody whose name I don't recognize? Likely not, but I think that is (unfortunately) part of being a human. We aren't capable (yet!) of evaluating everything on technical merit alone because the shear volume of material far exceeds our time. Therefore we use other (less reliable to be true) signalling mechanisms as a way to quickly decide what is worthy of our time investment and what may not be. Reputation/name recognition is a much imperfect, but better than random chance, indicator.
- s_Hogg 3y agoI'd pay good money to watch jart in conversation with Carmack
- Solvency 3y agoCarmack is great but completely irrelevant here. He missed the entire AI/LLM/ML boat to help Zuckerberg hawk virtual reality fantasies for years.
- vinkelhake 3y agoCompletely irrelevant is probably overstating it. He's been working on AI for the last 4+ years.
- cactusplant7374 3y agoHe's striving for AGI though, right? So he's not really working on anything because he certainly hasn't discovered AGI.
- Solvency 3y agoHe literally squandered the last 10 years of his life working on absolutely nothing for Zuckerberg. And only after the rest of the world innovated on AI (transformers, etc) did he clearly feel embarrassed and had to proclaim he's going to focus on AGI in a "one-up" way.
- talldayo 3y ago> He literally squandered the last 10 years of his life working on absolutely nothing Speak for yourself, the Oculus Quest is the coolest piece of sub-$500 tech in my home.
- fkyoureadthedoc 3y agoHe got paid a lot to do something he was presumably passionate about and enjoyed. It also might surprise you to find out that there's quite a lot of people that just work as a means to an end, and find value and enjoyment primarily from other parts of their life.
- m3kw9 3y agoSo Nvidia in trouble now because intel can be used instead for faster/cheaper? inference?
- tubs 3y agoThe ram is not on the cpu on a mac. It's in the same can but it's still regular ddr dimms.
- marshallward 3y agoThere is an implication here that the Fortran implementation of `SGEMM` is somehow inadequate. But any modern Fortran compiler will quite easily apply the AVX and FMA optimizations presented here without any additional changes. Both GNU and Intel make these substitutions with the correct flags. The unrolling optimization is also just another flag away (`-funroll-all-loops`). The Intel Compiler will even do this without prompting. In fact, it appears to only do a modest 2x unroll on my machine, suggesting that the extreme unroll in this article would have been overkill. Parallelization certainly a lot to ask of Fortran 77 source, but there there is little stopping you from adding OpenMP statements to the `SGEMM` function. In fact, modern Fortran even offers its own parallelization constructs if you're willing to go there. Which is to say: Let's not belittle this old Fortran 77 function. Yes it is old, and does not even resemble modern Fortran. But the whole point of Fortran is to free the developer from these platform-specific details, and hand the job off to the compiler. If you don't like that approach, then you're welcome to go to C or C++. But this little block of Fortran code is already capable of doing just about everything in this article.
- steppi 3y agoThe Fortran implementation is just a reference implementation. The goal of reference BLAS [0] is to provide relatively simple and easy to understand implementations which demonstrate the interface and are intended to give correct results to test against. Perhaps an exceptional Fortran compiler which doesn't yet exist could generate code which rivals hand (or automatically) tuned optimized BLAS libraries like OpenBLAS [1], MKL [2], ATLAS [3], and those based on BLIS [4], but in practice this is not observed. Justine observed that the threading model for LLaMA makes it impractical to integrate one of these optimized BLAS libraries, so she wrote her own hand-tuned implementations following the same principles they use. [0] https://en.wikipedia.org/wiki/Basic_Linear_Algebra_Subprograms https://en.wikipedia.org/wiki/Basic_Linear_Algebra_Subprogra... [1] https://github.com/OpenMathLib/OpenBLAS https://github.com/OpenMathLib/OpenBLAS [2] https://www.intel.com/content/www/us/en/developer/tools/oneapi/onemkl.html#gs.6q6b7q https://www.intel.com/content/www/us/en/developer/tools/onea... [3] https://en.wikipedia.org/wiki/Automatically_Tuned_Linear_Algebra_Software https://en.wikipedia.org/wiki/Automatically_Tuned_Linear_Alg... [4]https://en.wikipedia.org/wiki/BLIS_(software) https://en.wikipedia.org/wiki/BLIS_(software)
- hrkfmud50k 3y ago> It's clearly optimal since my CPU is listed as only being capable of going 780 gigaflops 780 GFLOP is the iGPU spec. Is this a valid comparison? https://nanoreview.net/en/cpu/intel-core-i9-14900k https://nanoreview.net/en/cpu/intel-core-i9-14900k
- arendtio 3y agoDoes someone else see llamafile using Wine on Linux? Edit: After the download I did a simple chmod +x llava-v1.5-7b-q4.llamafile; ./llava-v1.5-7b-q4.llamafile
- jart 3y agoThere's a simple fix for that. sudo wget -O /usr/bin/ape https://cosmo.zip/pub/cosmos/bin/ape-$(uname -m).elf sudo chmod +x /usr/bin/ape sudo sh -c "echo ':APE:M::MZqFpD::/usr/bin/ape:' >/proc/sys/fs/binfmt_misc/register" sudo sh -c "echo ':APE-jart:M::jartsr::/usr/bin/ape:' >/proc/sys/fs/binfmt_misc/register" https://github.com/mozilla-ocho/llamafile/?tab=readme-ov-file#gotchas https://github.com/mozilla-ocho/llamafile/?tab=readme-ov-fil...
- yieldcrv 3y agonote, this is "goes faster on CPUs than before", not faster than GPUs.
- TimPC 3y agoStrange title. My first read of the title thought the author was arguing the model is now faster on CPU than GPU. Would be much nicer if they titled this something closer to "Performance Improvement for LLaMa on CPU".
- utopcell 3y agoSame here.
- aaronscott 3y ago> I like to define my subroutines using a modern language like C++, which goes 47 gigaflops. This means C++ is three orders of a magnitude faster than Python. That's twenty years of progress per Moore's law. This is great. I love the idea of measuring performance differences in “years of Moore’s law.” Twenty years puts the delta in an easy to understand framework.
- JohnKemeny 3y agoI doubt that you get Python to run faster than C++ at 2004 hardware.
- mrtranscendence 3y agoPython on 2024 hardware vs C++ on 2004 hardware ... I don't think it's obvious that C++ always wins here, though it would depend on the use case, how much of the Python is underpinned by native libraries, and the specific hardware in question.
- JohnKemeny 3y agoIf we allow native libraries, it's not clear that C++ would win, even on modern hardware.
- michaelt 3y agoI think we all know that, when someone writes "C++ is three orders of a magnitude faster than Python" they're not including native libraries.
- mrtranscendence 3y agoYou can't not include native libraries, at least if you want your benchmark to be realistic. Almost every Python library where performance matters is written (at least partially) in a compiled language.
- 3y ago
- ein0p 3y agoAs someone who has tried to beat MKL-DNN, and was unsuccessful at doing so even for constrained matrix sizes, I’m curious how they pulled off such a massive improvement. But as someone who routinely estimates picojoules per flop at $DAY_JOB - there’s simply no way this is energy efficient. That is not even physically possible with a CPU.
- janwas 3y agoI think the previous code was using dot products, f32 instead of bf16.
- rbnsl 3y agoDefinitely wild we’re in the timeline you can run a 1.1 bn param model on a raspberry pi, but its still tough to justify because the 1.1 is kinda useless compared to the beefier models. Sick for home builds/hobbyists though I might wanna get one of the new Pis just to try this out
- JohnnyHerz 3y agoAwesomeness. thank you for sharing!
- deleted 3y ago[deleted]
- column 3y agoUnfortunately BitDefender (corporate) blocks llamafile as a ransomware "atc.heur.crypt" and it seems there is no workaround. :(
- apitman 3y agoAre the executables not signed? That would be surprising to me for something coming from Mozilla. EDIT: I just realized that the cross-platform single-binary thing might actually cause issues with code signing. I'm curious about this.
- saagarjha 3y ago> One important thing to know if you're considering buying a Mac Studio is that, like the Windows Executive, XNU does a really good job keeping your desktop stable, and that means protecting your system from you. It takes me 45 seconds on Mac Studio to compile the Cosmo monorepo, due to all these safety features; but if I fork bombed it, I'd be surprised if Netflix skipped a single frame. Clearly nobody actually tried this, because on XNU if you fork bomb the system it reliably goes down every single time. There are no "safety features" here but extra overhead when spawning processes.
- Dobiasd 3y agoAre there any benchmarks on the performance of these new matrix multiplication kernels compared to the Eigen library (ideally for float32)?
- Dobiasd 3y agoWhile I did not succeed in making the matmul code from https://github.com/Mozilla-Ocho/llamafile/blob/main/llamafile/sgemm.cpp https://github.com/Mozilla-Ocho/llamafile/blob/main/llamafil... work in isolation, I compared eigen, openblas, and mkl: https://gist.github.com/Dobiasd/e664c681c4a7933ef5d2df7caa87cb94 https://gist.github.com/Dobiasd/e664c681c4a7933ef5d2df7caa87... In this (very primitive!) benchmark, MKL was a bit better than eigen (~10%) on my machine (i5-6600). Since the article https://justine.lol/matmul/ https://justine.lol/matmul/ compared the new kernels with MLK, we can (by transitivity) compare the new kernels with Eigen this way, at least very roughly for this one use-case.
- jart 3y agoHere's a complete working example for POSIX systems on how to reproduce my llamafile tinyBLAS vs. MKL benchmarks: https://gist.github.com/jart/640231a627dfbd02fb03e23e8b01e592 https://gist.github.com/jart/640231a627dfbd02fb03e23e8b01e59... This new generalized kernel does even better than what's described in the blog post. It works well on oddly shaped matrices. It needs however a good malloc function, which I've included in the gist. Since having the good memory allocator is what makes the simple implementation possible.
- DrNosferatu 3y agoAny performance benchmark against intel's 'IPEX-LLM'[0] or others? [0] - https://github.com/intel-analytics/ipex-llm https://github.com/intel-analytics/ipex-llm