9 ms·
M2's Neural Engine had 15TOPS, M3's 18TOPS (+20%) vs. M4's 38TOPS (+111%). In transistor counts, M2 had 20BTr, M3 25BTr (+25%) and M4 has 28BTr (+12%). M2 use
by dhx 2y ago
M2's Neural Engine had 15TOPS, M3's 18TOPS (+20%) vs. M4's 38TOPS (+111%).
In transistor counts, M2 had 20BTr, M3 25BTr (+25%) and M4 has 28BTr (+12%).
M2 used TSMC N5P (138MTr/mm2), M3 used TSMC N3 (197MTr/mm2, +43%) and M4 uses TSMC N3E (215MTr/mm2, +9%).[1][2]
[1] https://en.wikipedia.org/wiki/5_nm_process#%225_nm%22_process_nodes https://en.wikipedia.org/wiki/5_nm_process#%225_nm%22_proces...
[2] https://en.wikipedia.org/wiki/3_nm_process#%223_nm%22_process_nodes https://en.wikipedia.org/wiki/3_nm_process#%223_nm%22_proces...
- ttul 2y agoAn NVIDIA RTX 4090 generates 73 TFLOPS. This iPad gives you nearly half that. The memory bandwidth of 120 GBps is roughly 1/10th of the NVIDIA hardware, but who’s counting!
- kkielhofner 2y agoTOPS != TFLOPS RTX 4090 Tensor 1,321 TOPS according to spec sheet so roughly 35x. RTX 4090 is 191 Tensor TFLOPS vs M2 5.6 TFLOPS (M3 is tough to find spec). RTX 4090 is also 1.5 years old.
- imtringued 2y agoYeah where are the bfloat16 numbers for the neural engine? For AMD you can at least divide by four to get the real number. 16 TOPS -> 4 tflops within a mobile power envelope is pretty good for assisting CPU only inference on device. Not so good if you want to run an inference server but that wasn't the goal in the first place. What irritates me the most though is people comparing a mobile accelerator with an extreme high end desktop GPU. Some models only run on a dual GPU stack of those. Smaller GPUs are not worth the money. NPUs are primarily eating the lunch of low end GPUs.
- lemcoe9 2y agoThe 4090 costs ~$1800 and doesn't have dual OLED screens, doesn't have a battery, doesn't weigh less than a pound, and doesn't actually do anything unless it is plugged into a larger motherboard, either.
- talldayo 2y agoFrom Geekbench: https://browser.geekbench.com/opencl-benchmarks https://browser.geekbench.com/opencl-benchmarks Apple M3: 29685 RTX 4090: 320220 When you line it up like that it's kinda surprising the 4090 is just $1800. They could sell it for $5,000 a pop and it would still be better value than the highest end Apple Silicon.
- nicce 2y agoA bit off-topic since not applicable for iPad: Adding also M3 MAX: 86072 I wonder the results if the test would be done on Asahi Linux some day. Apple implementation is fairly unoptimized AFAK.
- haswell 2y agoComparing these directly like this is problematic. The 4090 is highly specialized and not usable for general purpose computing. Whether or not it's a better value than Apple Silicon will highly depend on what you intend to do with it. Especially if your goal is to have a device you can put in your backpack.
- talldayo 2y agoI'm not the one making the comparison, I'm just providing the compute numbers to the people who did. Decide for yourself what that means, the only conclusion I made on was compute-per-dollar.
- aurareturn 2y agoI think it would be simpler to compare cost/transistor.
- pulse7 2y agoThis is true, but... RTX 4090 has only 24GB RAM and M3 can run with 192GB RAM... A game changer for largest/best models...
- 2y ago
- brigade 2y agoIt would also blow through the iPad’s battery in 4 minutes flat
- jocaal 2y ago> The memory bandwidth of 120 GBps is roughly 1/10th of the NVIDIA hardware, but who’s counting Memory bandwidth is literally the main bottleneck when it comes to the types of applications gpus are used for, so everyone is counting
- anvuong 2y agoThis comment needs to be downvoted more. TFLOPS is not TOPS, this comparison is meaningless, the 4090 is about 40x TOPS of the M4.
- ttul 2y agoMany thanks for the encouraging comments.
- bearjaws 2y agoWe will have M4 laptops running 400B parameter models next year. Wild times.
- visarga 2y agoAnd they will fit in the 8GB RAM with 0.02 bit quant
- gpm 2y agoYou can get a macbook pro with 128 GB of memory (for nearly $5000). Which still implies... a 2 bit quant?
- freeqaz 2y agoThere are some crazy 1/1.5 bit quants now. If you're curious I'll try to dig up the papers I was reading. 1.5bit can be done to existing models. The 1 bit (and less than 1 bit iirc) requires training a model from scratch. Still, the idea that we can have giant models running in tiny amounts of RAM is not completely far fetched at this point.
- gpm 2y agoYeah, I'm broadly aware and have seen a few of the papers, though I definitely don't try and track the state of the art here closely. My impression and experience trying low bit quants (which could easily be outdated by now) is that you are/were better off with a smaller model and a less aggressive quantization (provided you have access to said smaller model with otherwise equally good training). If that's changed I'd be interested to hear about it, but definitely don't want to make work for you digging up papers.
- moneywoes 2y agoeli5 quant?
- gpm 2y ago
- adrian_b 2y ago> The Most Powerful Neural Engine Ever While it is true that the claimed performance for M4 is better than for the current Intel Meteor Lake and AMD Hawk Point, it is also significantly lower (e.g. around half) than the AI performance claimed for the laptop CPU+GPU+NPU models that both Intel and AMD will introduce in the second half of this year (Arrow Lake and Strix Point).
- whynotminot 2y ago> will introduce Incredible that in the future there will be better chips than what Apple is releasing now.
- hmottestad 2y agoDon’t worry. It’s Intel we’re talking about. They may say that it’s coming out in 6 months, but that’s never stopped them from releasing it in 3 years instead.
- adrian_b 2y agoAMD is the one that has given more precise values (77 TOPS) for their launch, their partners are testing the engineering samples and some laptop product listings seem to have been already leaked, so the launch is expected soon (presentation in June, commercial availability no more than a few months later).
- spxneo 2y agoI literally don't give a fck about Intel anymore they are irrelevant The taiwanese silicon industrial complex deserves our dollars. Their workers are insanely hard working and it shows in its product.
- benced 2y agoThere's no Taiwanese silicon industrial complex, there's TSMC. The rest of Taiwanese fabs are irrelevant. Intel is the clear #3 (and looks likely-ish to overtake Samsung? We'll see).
- paulpan 2y agoThe fact that TSMC publishes their own metrics and target goals for each node makes it straightforward to compare the transistor density, power efficiency, etc. The most interesting aspect of the M4 is simply it's debuting on the iPad lineup, whereas historically it's always been on the iPhone (for A-series) and Macbook (for M-series). Makes sense given low expected yielded for the newest node for one of Apple's lower volume products. For the curious, the original TSMC N3 node had a lot of issues plus was very costly so makes sense to move away from it: https://www.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-even https://www.semianalysis.com/p/tsmcs-3nm-conundrum-does-it-e...
- spenczar5 2y agoiPads are actually much higher volume than Macs. Apple sells about 2x to 3x as many tablets as laptops. Of course, phones dwarf both.
- andy_xor_andrew 2y agoThe iPad Pros, though? I'm very curious how much iPad Pros sell. Out of all the products in Apple's lineup, the iPad Pro confuses me the most. You can tell what a PM inside Apple thinks the iPad Pro is for, based on the presentation: super powerful M4 chip! Use Final Cut Pro, or Garageband, or other desktop apps on the go! Etc etc. But in reality, who actually buys them, instead of an iPad Air? Maybe some people with too much money who want the latest gadgets? Ever since they debuted, the general consensus from tech reviewers on the iPad Pro has been "It's an amazing device, but no reason to buy it if you can buy a MacBook or an iPad Air" Apple really wants this "Pro" concept to exist for iPad Pro, like someone who uses it as their daily work surface. And maybe some people exist like that (artists? architects?) but most of the time when I see an iPad in a "pro" environment (like a pilot using it for nav, or a nurse using it for notes) they're using an old 2018 "regular" iPad.
- transpute 2y agoiPadOS 16.3.1 can run virtual machines on M1/M2 silicon, https://old.reddit.com/r/jailbreak/comments/18m0o1h/tutorial_tiny11arm64_vm_on_ipad_m1/?rdt=38058 https://old.reddit.com/r/jailbreak/comments/18m0o1h/tutorial... Hypervisor support was removed from the iOS 16.4 kernel, hopefully it will return in iPadOS 18 for at least some approved devices. If not, Microsoft/HP/Dell/Lenovo Arm laptops with M3-competitive performance are launching soon, with mainline Linux support.
- barbariangrunge 2y agoMy m2 pro is already more powerful than I can use. The screen is too small to do big work like using a daw or doing video editing, the Magic Keyboard is uncomfortable so I stopped writing on it. All that processing power, I don’t know what it will be used for on a tablet without even a good file system. Lousy ergonomics