16 ms·
AI PCs Aren't Good at AI: The CPU Beats the NPU
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- fancyfredbot 2y agoThe write up on the GitHub repo is much more informative than the blog. When running int8 matmul using onnx performance is ~0.6TF. https://github.com/usefulsensors/qc_npu_benchmark https://github.com/usefulsensors/qc_npu_benchmark
- dang 2y agoThanks—we changed the URL to that from https://petewarden.com/2024/10/16/ai-pcs-arent-very-good-at-ai/ https://petewarden.com/2024/10/16/ai-pcs-arent-very-good-at-.... Readers may way want to look at both, of course!
- dhruvdh 2y agoOh, maybe also change the title? I flagged it because of the title/url not matching.
- dmitrygr 2y agoIn general MAC unit utilization tends to be low for transformers, but 1.3% seems pretty bad. I wonder if they fucked up the memory interface for the NPU. All the MACs in the world are useless if you cannot feed them.
- moffkalast 2y agoI recall looking over the Ryzen AI architecture and the NPU is just plugged into PCIe and thus gets completely crap memory bandwidth. I would expect it might be similar here.
- PaulHoule 2y agoI spent a lot of time with a business partner and an expert looking at the design space for accelerators and it was made very clear to me that the memory interface puts a hard limit on what you can do and that it is difficult to make the most of. Particularly if a half-baked product is being rushed out because of FOMO you’d practically expect them to ship something that gives a few percent of the performance because the memory interface doesn’t really work, it happens to the best of them: https://en.wikipedia.org/wiki/Cell_(processor) https://en.wikipedia.org/wiki/Cell_(processor)
- wtallis 2y agoIt's unlikely to be literally connected over PCIe when it's on the same chip. It just looks like it's connected over PCIe because that's how you make peripherals discoverable to the OS. The integrated GPU also appears to be connected over PCIe, but obviously has access to far more memory bandwidth.
- Hizonner 2y agoIt's a tablet. It probably has like one DDR channel. It's not so much that they "fucked it up" as that they knowingly built a grossly unbalanced system so they could report a pointless number.
- dmitrygr 2y agoWell, no. If the CPU can hit better numbers on the same model then the bandwidth from the DDR IS there. Probably the NPU does not attach to the proper cache level, or just has a very thin pipe to it
- Hizonner 2y agoThe CPU is only about twice as good as the NPU, though (four times as good on one test). The NPU is being advertised as capable of 45 trillion operations per second, and he's getting 1.3 percent of that. So, OK, yeah, I concede that the NPU may have even worse access to memory than the CPU, but the bottom line is that neither one of them has anything close to what it needs to to actually delivering anything like the marketing headline performance number on any realistic workload. I bet a lot of people have bought those things after seeing "45 TOPS", thinking that they'd be able to usefully run transformers the size of main memory, and that's not happening on CPU or NPU.
- dmitrygr 2y agoYup, sad all round. We are in agreement.
- pram 2y agoI laughed when I saw that the Qualcomm “AI PC” is described as this in the ComfyUI docs: "Avoid", "Nothing works", "Worthless for any AI use"
- deleted 2y ago[deleted]
- LikelyABurner 2y agoDidn't believe that anyone would be bridge-burning-happy enough to put this in their official docs, but you're not kidding: https://github.com/comfyanonymous/ComfyUI/wiki/Which-GPU-should-I-buy-for-ComfyUI https://github.com/comfyanonymous/ComfyUI/wiki/Which-GPU-sho... In retrospect, the fact that Intel and AMD's stock prices both closed slightly up when Microsoft announced the Snapdragon X on Windows 11 was a dead giveaway that the major players knew behind the scenes that it was being released seriously under baked.
- jsheard 2y agoThese NPUs are tying up a substantial amount of silicon area so it would be a real shame if they end up not being used for much. I can't find a die analysis of the Snapdragon X which isolates the NPU specifically but AMDs equivalent with the same ~50 TOPS performance target can be seen here, and takes up about as much area as three high performance CPU cores: https://www.techpowerup.com/325035/amd-strix-point-silicon-pictured-and-annotated#g325035-2 https://www.techpowerup.com/325035/amd-strix-point-silicon-p...
- Kon-Peki 2y agoModern chips have to dedicate a certain percentage of the die to dark silicon [1] (or else they melt/throttle to uselessness), and these kinds of components count towards that amount. So the point of these components is to be used, but not to be used too much. Instead of an NPU, they could have used those transistors and die space for any number of things. But they wouldn't have put additional high performance CPU cores there - that would increase the power density too much and cause thermal issues that can only be solved with permanent throttling. [1] https://en.wikipedia.org/wiki/Dark_silicon https://en.wikipedia.org/wiki/Dark_silicon
- IshKebab 2y agoIf they aren't being used it would be better to dedicate the space to more SRAM.
- a2l3aQ 2y agoThe point is parts of the CPU have to be off or throttled down when other components are under load to maintain TDP, adding cache that would almost certainly be being used defeats the point of that.
- jsheard 2y agoDoesn't SRAM have much lower power density than logic with the same area though? Hence why AMD can get away with physically stacking cache on top of more cache in their X3D parts, without the bottom layer melting.
- tromp 2y ago> the 45 trillion operations per second that’s listed in the specs Such a spec should be ideally be accompanied by code demonstrating or approximating the claimed performance. I can't imagine a sports car advertising a 0-100km/h spec of 2.0 seconds where a user is unable to get below 5 seconds.
- deleted 2y ago[deleted]
- dmitrygr 2y agomost likely multiplying the same 128x128 matrix from cache to cache. That gets you perfect MAC utilization with no need to hit memory. Gets you a big number that is not directly a lie - that perf IS attainable, on a useless synthetic benchmark
- kmeisthax 2y agoSounds great for RNNs! /s
- tedunangst 2y agoI have some bad news for you regarding how car acceleration is measured.
- isusmelj 2y agoI think the results show that just in general the compute is not used well. That the CPU took 8.4ms and GPU took 3.2ms shows a very small gap. I'd expect more like 10x - 20x difference here. I'd assume that the onnxruntime might be the issue. I think some hardware vendors just release the compute units without shipping proper support yet. Let's see how fast that will change. Also, people often mistake the reason for an NPU is "speed". That's not correct. The whole point of the NPU is rather to focus on low power consumption. To focus on speed you'd need to get rid of the memory bottleneck. Then you end up designing your own ASIC with it's own memory. The NPUs we see in most devices are part of the SoC around the CPU to offload AI computations. It would be interesting to run this benchmark in a infinite loop for the three devices (CPU, NPU, GPU) and measure power consumption. I'd expect the NPU to be lowest and also best in terms of "ops/watt"
- AlexandrB 2y ago> Also, people often mistake the reason for an NPU is "speed". That's not correct. The whole point of the NPU is rather to focus on low power consumption. I have a sneaking suspicion that the real real reason for an NPU is marketing. "Oh look, NVDA is worth $3.3T - let's make sure we stick some AI stuff in our products too."
- kmeisthax 2y agoYou forget "Because Apple is doing it", too.
- jamesy0ung 2y agoWhat exactly does Windows do with a NPU? I don't own an 'AI PC' but it seems like the NPUs are slow and can't run much. I know Apple's Neural Engine is used to power Face ID and the facial recognition stuff in Photos, among other things.
- DrillShopper 2y agoIt supports Microsoft's Recall (now required) spyware
- Janicc 2y agoPlease remind me again how Recall sends data to Microsoft. I must've missed that part. Or are you against the print screen button too? I heard that takes images too. Very scary.
- throwaway314155 2y ago[dead]
- cmeacham98 2y agoWhile calling it spyware like GP is over-exaggeration to a ridiculous level, comparing Recall to Print Screen is also inaccurate: Print Screen takes images on demand, Recall does so effectively at random. This means Recall could inadvertently screenshot and store information you didn't intend to keep a record of (To give an extreme example: Imagine an abuser uses Recall to discover their spouse browsing online domestic violence resources).
- Terr_ 2y ago> Please remind me again how Recall sends data to Microsoft. I must've missed that part. Sure, just post the source code and I'll point out where it does so, I somehow misplaced my copy. /s The core problem here is trust, and over the last several years Microsoft has burned a hell of a lot of theirs with power-users of Windows. Even their most strident public promises of Recall being "opt-in" and "on-device only" will--paradoxically--only be kept as long as enough people remain suspicious. Glance away and MS go back to their old games, pushing a mandatory "security update" which reset or entirely-removes your privacy settings and adding new "telemetry" streams which you cannot inspect.
- eightysixfour 2y agoI thought the purpose of these things was not to be fast, but to be able to run small models with very little power usage? I have a newer AMD laptop with an NPU, and my power usage doesn't change using the video effects that supposedly run on it, but goes up when using the nvidia studio effects. It seems like the NPUs are for very optimized models that do small tasks, like eye contact, background blur, autocorrect models, transcription, and OCR. In particular, on Windows, I assumed they were running the full screen OCR (and maybe embeddings for search) for the rewind feature.
- conradev 2y agoThat is my understanding as well: low power and low latency. You can see this in action when evaluating a CoreML model on a macOS machine. The ANE takes half as long as the GPU which takes half as long as the CPU (actual factors being model dependent)
- nickpsecurity 2y agoTo take half as long, doesn’t it have to perform twice as fast? Or am I misreading your comment?
- eightysixfour 2y agoNo, you can have latency that is independent of compute performance. The CPU/GPU may have other tasks and the work has to wait for the existing threads to finish, or for them to clock up, or have slower memory paths, etc. If you and I have the same calculator but I'm working on a set of problems and you're not, and we're both asked to do some math, it may take me longer to return it, even though the instantaneous performance of the math is the same.
- refulgentis 2y agoIn isolation, makes sense. Wouldn't it be odd for OP to present examples that are the opposite of their claim, just to get us thinking about "well the CPU is busy?" Curious for their input.
- wmf 2y agoThis headline is seriously misleading because the author did not test AMD or Intel NPUs. If Qualcomm is slow don't say all AI PCs are not good.
- protastus 2y agoDeploying a model on an NPU requires significant profile based optimization. Picking up a model that works fine on the CPU but hasn't been optimized for an NPU usually leads to disappointing results.
- catgary 2y agoYeah whenever I’ve spoken to people who work on stuff like IREE or OpenXLA they gave me the impression that understanding how to use those compilers/runtimes is an entire job.
- CAP_NET_ADMIN 2y agoBeauty of CPUs - they'll chew through whatever bs code you throw at them at a reasonable speed.
- marginalia_nu 2y agoI don't think this is correct. The difference between well optimized code and unoptimized code on the CPU is frequently at least an order of magnitude performance. Reason it doesn't seem that way is that the CPU is so fast we often bottleneck on I/O first. However, for compute-workloads like inference, it really does matter.
- consteval 2y agoWhile this is true, the most effective optimizations you don't do yourself. The compiler or runtime does it. They get the low-hanging fruit. You can further optimize yourself, but unless your design is fundamentally bad, you're gonna be micro-optimizing. gcc -O0 and -O2 has a HUGE performance gain. We don't really have anything to auto-magically do this for models, yet. Compilers are intimately familiar with x86.
- marginalia_nu 2y agoWhile the compiler is decent at producing code that is good in terms of saturating the instruction pipeline, there are many things the compiler simply can't help you with. Having cache friendly memory access patterns is perhaps the biggest one. Though automatic vectorization is also still not quite there, so in cases where there's a severe bottleneck, doing that manually may still considerably improve performance, if the workload is vectorizable.
- lostmsu 2y agoThe author's benchmark sucks if he could only get 2 tops from a laptop 4080. The thing should be doing somewhere around 80 tops. Given that you should take his NPU results with a truckload of salt.
- hkgjjgjfjfjfjf 2y agoSutherland's wheel of reincarnation turns.
- downrightmike 2y agoThey should have just made a pci card and not tried to push whole new machines on us. We are all good with the machines we already have. If you want to sell a new feature, then it needs to be an add-on
- Mistletoe 2y ago>The second conclusion is that the measured performance of 573 billion operations per second is only 1.3% of the 45 trillion ops/s that the marketing material promises. It just gets so hard to take this industry seriously.
- m00x 2y agoNPUs are efficient, not especially fast. The CPU is much bigger than the NPU and has better cache access. Of course it'll perform better.
- acdha 2y agoIt’s more complicated than that (you’re assuming that the bigger CPU is optimized for the same workload) but it’s also irrelevant to the topic at hand: they’re seeing this NPU within a factor of 2-4 of the CPU, but if it performed half as well as Qualcomm claims it would be an order of magnitude faster. The story here isn’t another round of the specialized versus general debate but that they fell so far short of their marketing claims.
- llm_nerd 2y agoNPUs are actually incredibly fast for standard inference operations. This benchmark is horribly flawed in many ways, and was so evidently useless that I'm surprised that they still decided to "publish" this. When your test gets 1% of the published performance, it's a good indication that things aren't being done correctly.
- Havoc 2y ago>We see 1.3% of Qualcomm's NPU 45 Teraops/s claim To me that suggests that the test is wrong. I could see intel massaging results, but that far off seems incredibly improbable
- rationalfaith 2y ago[dead]
- p1necone 2y agoI might be overly cynical but I just assumed that the entire purpose of "AI PCs" was marketing - of course they don't actually achieve much. Any real hardware that's supposedly for the "AI" features will actually be just special purpose hardware for literally anything the sales department can lump under that category.
- teilo 2y agoActual article title: Benchmarking Qualcomm's NPU on the Microsoft Surface Tablet Because this isn't about NPUs. It's about a specific NPU, on a specific benchmark, with a specific set of libraries and frameworks. So basically, this proves nothing.
- iml7 2y agoBut you can’t get more clicks. You have to attack enough people to get clicks.I feel like this place is becoming more and more filled with posts and titles like this.
- gerdesj 2y agoInternet points are a bit crap but HN generally discusses things properly and off topic and downright weird stuff generally gets downvoted to doom.
- gnabgib 2y agoThe title is from the original article (https://petewarden.com/2024/10/16/ai-pcs-arent-very-good-at-ai/ https://petewarden.com/2024/10/16/ai-pcs-arent-very-good-at-...), the URL was changed by dang: https://news.ycombinator.com/item?id=41863591 https://news.ycombinator.com/item?id=41863591
- NoPicklez 2y agoFairly misleading title, boiling down AI PCs to just the Microsoft Surface running Qualcomm
- cjbgkagh 2y ago> We've tried to avoid that by making both the input matrices more square, so that tiling and reuse should be possible. While it might be possible it would not surprise me if a number of possible optimizations had not made it into Onnx. It appears that Qualcomm does not give direct access to the NPU and users are expected to use frameworks to convert models over to it, and in my experience conversion tools generally suck and leave a lot of optimizations on the table. It could be less of NPUs suck and more of the conversions tools suck. I'll wait until I get direct access - I don't trust conversion tools. My view of NPUs is that they're great for tiny ML models and very fast function approximations which is my intended use case. While LLMs are the new hotness there are huge number of specialized tasks that small models are really useful for.
- jaygreco 2y agoI came here to say this. I haven’t worked with the Elite X but the past gen stuff I’ve used (865 mostly) the accelerators - compute DSP and much smaller NPU - required _very_ specific setup, compilation with a bespoke toolchain, and communication via RPC to name a few. I would hope the NPU on Elite X is easier to get to considering the whole copilot+ thing, but I bring this up mainly to make the point that I doubt it’s just as easy as “run general purpose model, expect it to magically teleport onto the NPU”.
- Hizonner 2y ago> While LLMs are the new hotness there are huge number of specialized tasks that small models are really useful for. Can you give some examples? Preferably examples that will run continuously enough for even a small model to stay in cache, and are valuable enough to a significant number of users to justify that cache footprint? I am not saying there aren't any, but I also honestly don't know what they are and would like to.
- consteval 2y agoiPhones use a lot of these. There's a bunch of little features that run on the NPU. Suggestions, predictive text, smart image search, automatic image classification, text selection in images, image processing. These don't run continuously, but I think they are valuable to a lot of users. The predictive text is quite good, and it's very nice to be able to search for vague terms like "license plate" and get images in my camera roll. Plus, selecting text and copying it from images is great. For desktop usecases, I'm not sure.
- stanleykm 2y agothe ARM SME could be an interesting alternative to NPUs in the future. Unlike the NPUs which have at best some fixed function API it will be possible to program the SMEs more directly
- piskov 2y agoSnapdragon touts 45 TOPS but it’s only int8. For example Apple's m3 neural engine is mere 18 TOPS but it’s FP16. So windows has bigger number but it’s not apple to apple comparison. Did author test int8 performance?
- freehorse 2y agoI always thought that the main point of NPUs is energy efficiency (and being able to run ML models without taking over all computer resources, making it practical to integrate ML applications in the OS itself in ways that it does not disturb the user or the workflow) rather than being exceptionally faster. At least this has been my experience with running stable diffusion on macs. Similar to using other specialised hardware like media encoders; they are not necessarily faster than a CPU if you throw a dozen+ cpu cores on the task, but it will draw a minuscule part of the power.
- tokyolights 2y agoNot sure why this isn't discussed more here. I think exactly the same, the NPU occupies more silicon area because it has custom circuits specifically to reduce the number of cycles (and thus energy) needed to perform those calculations. Doesn't necessarily mean that a CPU wouldn't be able to bulldoze through it faster (with much more energy).
- guelermus 2y agoOne should pay attention also to power efficiency, a direct comparison could be misleading here.
- _davide_ 2y agoThe RTX 4080 should be capable of ~40 TFLOPS, yet they only report 2,160 billion operations per second. Shouldn't this be enough to reconsider the benchmark? They probably made some serious error in measuring FLOPS. Regarding the fact that CPU beats NPU is possible but they should benchmark many matrix multiplications without any application synchronization in order to have a decent comparison.
- Grimblewald 2y agoThat isnt the half of it. A quick skim of the documentation shows that the cpu inference wasnt done in a comparable way either.
- ein0p 2y agoMemory bound workload is memory bound. Doesn’t matter how many TOPS you have if you’re sitting idle waiting on DRAM during generation. You will, however notice a difference in prefill for long prompts.
- woadwarrior01 2y agoIMO, benchmarking accelerator hardware with onnxruntime is like benchmarking a CPU with a Python script. > We've seen similar performance results to those shown here using the Qualcomm QNN SDK directly. Why not include those results?
- fschutze 2y agoIs there a possibility to use the Qualcomm SNPE SDK? I thought this SDK isn't bad. Also, for those who have access to the Qualcomm NPU: Is the Hexagon SDK working properly? Do apps still need to be signed (which i never got to work) when using Hexagon?
- pzo 2y agoHaven't played much with Qualcomm NPU but Apple Neural Engine available in iOS and MacOS for many Computer Vision models was significantly faster than when running on CPU or GPU (e.g. mediapipe models, yolo, depth-anything) - to the point that inference was much faster on Macbook M2 Max using its NPU that is the same as on older iPhones rather than executing on all 38 GPU cores. This all depends on model architecture, conversions and tuning. Apple provides good tooling in XCode for benchmarking models up to execution time of single operators and where such operator got executed (CPU, GPU, NPU) in case couldn't been executed on NPU and have to fallback to CPU/GPU. Sometimes model have be tweaked to slightly different operator if it's not available in NPU. On top of that ML frameworks/runtimes such as ONNX/Pytorch/TensorflowLite sometimes don't implement all operators in CoreML or MPS.
- irusensei 2y agoAre NPUs the VLIW of our times in terms of hype?
- lambda-research 2y agoThe benchmark is matrix multiplcation with the shapes `(6, 1500, 256) X (6, 256, 1500)`, which just aren't that big in the AI world. I think the gap would be larger with much larger matrices. E.g. Llama 3.1 8B which is one of the smaller models has matrix multiplications like `(batch, 14336, 4096) x (batch, 4096, 14336)`. I just don't think this benchmark is realistic enough.
- bsmartt 2y agowhat are all these folks hoping to accomplish? By crying and starting shit about windows recall, all you did was signal to their shareholders and the financial analysts that windows recall actually substance and not just a marketing facade. Otherwise, why would all those nerds be so angry? So microsoft takes some of the criticisms on twitter and gets them in before shipping. Free appsec, nice. Now, microsoft doesnt care about your benchmarks, dude. Grandma isnt gonna notice these workloads finish faster on a different compiled program utilizing different chips. Her last PC was EOL'd 10 years ago, it certainly cant keep up with this new ai laptop.
- bsmartt 2y agoYou dont seriously think MSFT expects this shit to benefit consumers do you? Their datacenters are overheating and the billing meter is still ticking while they burn, they need to figure out how to get consumers to start paying for this shit before they go broke and wall st sells them off for parts.. Either way, these are some of the first personal computers to have NPUs. They will improve. CPUs are 20 years optimized, this is literally the first try for some of these companies
- bsmartt 2y agoif anything this is a very promising benchmark for the new tech,. we're getting close. so what this means if NPUs are anywhere close to CPUs in the benchmarks is that NPUs are going to blow past CPUs very soon, because CPUs dont have much more weight to shed whereas NPUs are just getting started.
- NebulaTrek 2y agoWe ran qprof (a Qualcomm NPU profiler) on this benchmark. The profiling results indicate that the workload was distributed to the vector cores instead of the tensor core, which provide the vast majority of the compute power in the NPU (my back of napkin math suggests that HMX is 30x stronger than HVX). The workload is relatively small, which results in underutilization of the hardware capacity due to the overhead associated with input/output quantization-dequantization and NCHW-NHCW mapping. Padding the weights and inputs to be a multiple of 64 would also help the performance. Edit: Link to the profiling graph https://imgur.com/a/2OKR93e https://imgur.com/a/2OKR93e Estimated HVX compute capability 421.43*1024/8 = 1.46TOPS in int8, in which 4 is number of vector cores 2 is number operation per cycle 1.43GHz is HVX frequency 1024bit is vector register width 8bit is precision
- NebulaTrek 2y agoThe formula was formatted wrong, it should be 4 * 2 * 1.43 * 1024 / 8
- cloudhan 2y agoOK, I am one of the developers in onnxruntime team. Perviously working on ROCm EP now has been transfered to QNN EP. The following is purely devrant and the opinions are mine. So ROCm already sucks whereas QNN sucks even harder! The conclusion here is NVIDIA knows how to make software that just works. AMD makes software that might work. Qualcomm, however, knows zero piece of shit of how to make a useful software. The dev experience is just another level of disaster with Qualcomm. Their tools and APIs return absolutely zero useful infomation about what error you are getting, just an error code that you can grep from their include headers from SDK. To debug an error code, you need strace to get the internal error string on the device. Their profiler merely gives you a trace that cannot be associated back to original computing logic with very high stddev on the runtime. Their docs website is not indexed by the MF search engine, not to say LLMs, so if you have any question, good luck then! So if you don't have a reason to use QNN, just don't use it (and any other NPU you name it). Back to the benchmark script. There is a lot of flaws as I can see. 1. the session is not warmed up and the iteraion is too small. 2. the onnx graph is too small, I suspect the onnxruntime overhead cannot be ignored in this case. Try stack more gemm in the graph instead of increasing the iteration naively. 3. the "htp_performance_mode": "sustained_high_performance" might give a lower perf compare to "burst" mode. A more reliable way to benchmark might just dump the context binary[1] and dump context inputs[2] and run this with qnn-net-run to get rid of the onnxruntime overhead. [1]: https://github.com/cloudhan/misc-nn-test-driver/blob/main/qnn/dump_qnn_ctx.py https://github.com/cloudhan/misc-nn-test-driver/blob/main/qn... [2]: https://github.com/cloudhan/misc-nn-test-driver/blob/main/qnn/dump_qnn_inputs.py https://github.com/cloudhan/misc-nn-test-driver/blob/main/qn...
- cloudhan 2y agoNPU folks offen time say > it's not enough time to get new silicon designs specifically for <blahblah> Where blahblah stands for a model architecture that has caused a paradigm shift. When you need a new silicon for a new model, you are already losing.