10 ms·
Apple’s M1 processor and the full 128-bit integer product
- thesz 6y agoMultiplier in M1 can be pipelined or relicated (or both), so issuing two instructions can be as fast as issuing one. Instruction recognition logic (DAG analysis, BTW) is harder to implement than to implement pipelined multiplier. Former is a research project, while latter was done at the dawn of computing.
- acje 6y agoSo this is why the integer multiply accumulate instruction mullah, only delivers the most significant bits? Ironic if you aren't religious about these things.
- thefourthchime 6y agoI love my M1, but does anyone else have horrific performance when resuming from wake? It’s like it swaps everything to disk and takes a full minute to come back to life.
- thefourthchime 6y agoI tried reinstalling OSX, now it seems fine! ¯\_(ツ)_/¯
- robbiep 6y agoVery happy with my Air 16gb on resume - much faster than my 2018 air
- deleted 6y ago[deleted]
- mwint 6y agoHaven’t seen this on my M1 Mac mini; my wake times are faster than my monitors can wake. No performance issues immediately after sleep. Maybe desktop platforms sleep differently than laptops?
- singhkays 6y agoMost likely. Same here with my Mini
- SCUSKU 6y agoI haven't noticed poor wake times, but my laptop does kernel panic and reboot a fair amount. Maybe 4 times in the past week. My hunch is that it's Spotify's fault but I haven't dug into the logs.
- Reason077 6y agoI've had mine panic and reboot twice, both times happened shortly after disconnecting other Macs that were connected via a Thunderbolt cable (target disk mode).
- mhh__ 6y agoHow does spotify panic the kernel itself?
- SigmundA 6y agoA user-land program cannot be at fault for a kernel panic, that is the kernels fault, always.
- ridiculous_fish 6y agoxnu will panic if it doesn't receive periodic check-ins from userspace. For example if WindowServer hangs, then the kernel may deliberately panic so that the system reboots. See man watchdogd for (a tiny bit) more.
- saagarjha 6y agoTrue, but a third-party program shouldn’t really be able to do this, at least not easily.
- siquick 6y agoI have this problem but only if I've been plugged into a monitor, and then unplugged and gone onto battery. Rebooting after unplugging stops it but its annoying.
- Reason077 6y agoDefinitely no issue with waking my M1 8GB MacBook Air. Takes a fraction of a second, every time. In fact, this is specifically something that Apple were bragging about when they launched the M1 Macs!
- floatingatoll 6y agoIt’s inappropriate of you to post an offtopic end user technical support question on this post about CPU microarchitecture performance.
- xiphias2 6y agoNo, it's perfect, instant wake for me. I have never seen anything like this. I have buggy apps (like Facebook Messenger) locking up, but I guess that's normal, I just uninstall them.
- thefourthchime 6y agoWhich M1 do you have?
- xiphias2 6y ago16 GB RAM pro 1TB SSD. I guess the only drawback compared to MB Air is that it's a bit heavier.
- zitterbewegung 6y agoDo you have it plugged into a monitor upon wake? How many programs do you have when it is resuming? I noticed this really infrequently.
- bombcar 6y agoI have a non-M1 with Big Sir and it takes more than full minute from sleep to usable. I suspect it’s because I have five monitors and 20 million pixels (actually more as that’s the post-retina resolution).
- monsieurbanana 6y agoHow does the number of pixels (or even monitors, for that matter), affect that much sleeping time? Rendering a FPS game at 1080p is 2 million pixels per frame. At 60fps, that's rendering 120 million pixels per second. What am I missing?
- bzzzt 6y agoDetecting the monitors, negotiating the correct resolutions, setting scaling factors and window positions after coming out of sleep will take some time. Maybe macOS does monitor setup sequentially? (no idea, just got 1 big external screen that also takes a few seconds to light up - handshake speed seems to vary between monitor brands)
- bombcar 6y agoThis is almost certainly it - the screens show all sorts of weird graphical artifacts (some clearly a Retina display in "native" resolution) as it starts and loads. I assume it's having to fire up all the GPU memory, etc.
- rckoepke 6y agoInstant wake for me. However any time I come across a password field in a website the computer freezes for a painfully long 10 seconds or so while it presumably decrypts my password vault. Sometimes this will happen multiple times per page load if I deselect and reselect the password field.
- Waterluvian 6y agoDid you somehow accumulate a bajillion passwords?
- CameronNemo 6y agoCrypto optimizations are no joke. My Pinebook Pro takes several seconds longer than my T430 to decrypt my keepassxc database.
- coldtea 6y agoWhy would it need to decrypt a whole database though, and not just the required password? To avoid leaking the key name?
- CameronNemo 6y agoYes I think it is an all or nothing deal.
- coldtea 6y agoPerhaps it could be a two-level thing then, where you first decrypt e.g. the list of keys and then read the one you want and get a "start/end" offset to read/decrypt just the value for that key from another file. So keys and values are still decrypted, but you don't have to decrypt the whole bundle in one step to get the value.
- akvadrako 6y agoMaybe the time is hashing your password. This is designed to take as long as it can while still being reasonable. 1 second on a fast machine isn't unheard of.
- kureikain 6y agoIt way faster than my old macbook pro. But one thing is the external monitor won't be open upon resuming. I have to plug/unplug the cable to re-active the external monitor. It seems HDMI handshake was failed somehow
- rootusrootus 6y agoYes. This is actually a known issue, provided you have an external monitor attached; lots of people complaining about it. The Mac actually wakes up instantly if you lift the screen, but it usually takes 5-10 seconds before it will wake up the external monitor. Worse, for some of us when it does finally wake up the monitor, sometimes it wakes it up with all the wrong colors, and rebooting is the only reliable fix. (and before anyone asks, yes, I tried a different HDMI cable)
- benhurmarcel 6y agoDoes this happen with the Mac Mini?
- matwood 6y agoI think they are still working through some external monitor driver issues. The colors on my LG 4k were initially way off until I did an advanced calibration. Occasionally waking up from sleep it will revert, but opening the Display preference and swapping between the calibrations fixes it. My connection is through usbc. I don't have any performance issues waking up though.
- rootusrootus 6y agoI think there are a couple issues going on with the colors. For some people it is calibration. But what I experience is almost like a color inversion (but it's not a full inversion, it looks like maybe one or two channels got inverted). Makes it difficult to even find the mouse pointer so I can get to the menu and reboot the machine. Then it comes up fine.
- matwood 6y agoInteresting. I did submit a bug to Apple about my issue. Mainly because I never did a calibration with my 2017 mbp and the monitor looked great. With the M1 MBA I had to do the advanced calibration just to make the same monitor usable. Otherwise the colors were completely washed out. It almost seemed like the setting to auto-dim the laptop monitor was also being oddly applied to the external.
- TheRealSteel 6y agoThis is anecdotal and I don't have anything to prove it, but I really feel like my old spinning-hard-drive 2010 MacBook Pro woke faster from sleep running Snow Leopard, than the Retina models ever did (or the old ones did after a few software updates). Of course for general tasks it was slower, but I really remember that thing waking up instantly when I raised the lid, every time.
- 2muchcoffeeman 6y agoMy nostalgia agrees with you, but I think it’s probably wrong. The wake from sleep was what finally convinced me to get my first Mac. Even now they have the best hibernation and wake from sleep.
- wil421 6y agoNope, my M1 Air is connected to an external monitor most of the day and I don’t have issues.
- mindajar 6y agoYeah. To me it looks like macOS goes so deep into sleep it disconnects the external display. On wake, the system rediscovers the external and resizes the desktop across both displays. With a bunch of apps/windows open, half your apps simultaneously resizing all their windows can peg all CPU cores for a number of seconds. (It's still way faster than the same set of apps on an Intel Mac laptop, where it could sometimes take on the order of 30 seconds to get to a usable desktop after a long sleep. On Intel Macs it seemed more obvious that the GPU was the bottleneck)
- dcow 6y agoI am using the LG Ultrafine 5K (so it’s a TB monitor) and it takes maybe 1.5s to 5s longer than the built in display (which wakes instantly) to turn on. I do occasionally have an issue where the brightness on the built in display is borked and won’t adjust back to the correct level for anywhere between 30s to a few minutes. And then I don’t know if it’s my monitor or the M1, but sometimes there will be a messed up run of consecutive pixel columns about 1/10th of the screen wide starting about 30% from the left of the display. The entire screen in that region is shifted a few pixels upwards. Sometimes it’s hard to notice it but once you do it can’t be unseen. Replugging the monitor into the M1 resolves the issue.
- LAMike 6y agoAnyone want to take a guess at how long it will be until Apple has their own fab in the US making M1 chips?
- tedyoung 6y agoNot in the next 10 years. Why would they? Fabs are extremely capital-intensive and take years to get up and running, when (like Taiwan Semi) knows how to do it. Intel has shown how hard it can be to do this right. Let TSM work on production (and hopefully get more/larger fabs in the USA up and running) and getting better at packing in the transistors, and let Apple improve the design (and software).
- londons_explore 6y agoBut Apple has a lot of capital, and could win massive political brownie points for doing so, especially if they promised that some percentage of fab capacity would be sold to other American firms.
- deleted 6y ago[deleted]
- robjan 6y agoTSMC is based in on the soil of one of America's allies.
- mschuster91 6y agoAn ally who China is dreaming of re-assimilating (or taken over, depending on which side of the view you are) since its inception. China is engaging in ever more aggressive saber rattling and the total lack of any measurable reaction to their takeover of Hong Kong only has emboldened them. Who can guarantee Taiwan won't end up the same fate?
- robjan 6y ago
- p1mrx 6y agoRISC-V does this too: https://five-embeddev.com/riscv-isa-manual/latest/m.html https://five-embeddev.com/riscv-isa-manual/latest/m.html "If both the high and low bits of the same product are required, then the recommended code sequence is [...]. Microarchitectures can then fuse these into a single multiply operation instead of performing two separate multiplies."
- lincpa 6y ago# Apple M1 chip is a warehouse/workshop model Copyright © 2018 Lin Pengcheng. All rights reserved. My computer hardware architecture ("warehouse/workshop model") was published on Feb. 06, 2019. One or two years later, the Apple M1 chip adopted the warehouse/workshop model and was released on Nov. 11, 2020. - Warehouse: unified memory - Workshop: CPU, GPU and other cores - Products (raw materials): information, data there's also a new unified memory architecture that lets the CPU, GPU, and other cores exchange information between one another, and with unified memory, the CPU and GPU can access memory simultaneously rather than copying data between one area and another. Accessing the same pool of memory without the need for copying speeds up information exchange for faster overall performance. reference: Developer Delves Into Reasons Why Apple's M1 Chip is So Fast - From the introduction - M1 has not done global optimization of various core scheduling. - M1 only optimizes the access to memory data. - Apple needs to further improve the programming language and compiler to support and promote my programming methodology. - My architecture supports a wider range of workshop types than Apple M1, with greater efficiency, scalability and flexibility. - Conclusion M1 still needs a lot of optimization work, now its optimization level is still very simple, after all, it is only the first generation of works, released in stages. ## Forecast(2021-01-19): Intel, AMD, ARM, supercomputer, etc. will adopt the warehouse/workshop model With more and more CPU and GPU cores, and the number and types of peripherals, the communication, coordination, and management of cores (or peripherals) have become more and more important, They become a key factor in performance. Finally, the "warehouse/workshop model" will surely replace the "von Neumann architecture" and become the first architecture in the computer field.
- tzs 6y agoI wonder if order matters? That is, would mul followed by mulh be the same speed as mulh followed by mul? How about if there is an instruction between them that does not do arithmetic? (What I'm wondering here is if the processor recognizes the specific two instruction sequence, or if it something more general like mul internally producing the full 128 bits, returning the lower 64, and caching the upper 64 bits somewhere so this if there is a mulh before something overwrites that cache it can use it).
- sgtnoodle 6y agoIt seems like something that would be arbitrary depending on how the optimization was implemented. There wouldn't be an inherent need for that amount of generalization. Apple can tightly control their compiler to follow the rules, and there seemingly wouldn't be any compelling reason not to stick those two instructions back to back in a consistent order, since the second instruction is effectively free. It would be fun to experiment with, for someone that has the hardware. My guess is that swapping the order will make it slower, but adding an independent instruction or two between them probably won't have a measureable effect. It would be fun to try and consistently interrupt the CPU between the two instructions as well somehow, to see if that short-circuits the optimization.
- mmaunder 6y agoA common misconception about RISC processors.
- chrisseaton 6y agoIt’s not a RISC thing - CISC implementations do exactly the same kind of fusion for similar pairs of operations.
- userbinator 6y agoIt has a bit less gain on a RISC due to the code density (or lack thereof), since it requires more fetch bandwidth. Apple works around this by using a very wide front-end: https://news.ycombinator.com/item?id=25257932 https://news.ycombinator.com/item?id=25257932
- _chris_ 6y agox86-64 code density is more than 4 bytes / instruction.
- saagarjha 6y agoThat looks like inverse density.
- chrisseaton 6y agoIsn’t that the point?
- titzer 6y agoDid anyone actually look at the machine code generated here? 0.30ns per value? That is basically 1 cycle. Of course, there is no way that a processor can compute so many dependent instructions in one cycle, simply because they generate so many dependent micro-ops, and every micro-op is at least one cycle to go through an execution unit. So this must mean that either the compiler is unrolling the (benchmarking) loop, or the processor is speculating many loop iterations into the future, so that the latencies can be overlapped and it works out to 1 cycle on average. 1 cycle on average for any kind of loop is just flat out suspicious. This requires a lot more digging to understand. Simply put, I don't accept the hastily arrived-at conclusion, and wish Daniel would put more effort into investigation in the future. This experiment is a poor example of how to investigate performance on small kernels. You should be looking at the assembly code output by the compiler at this point instead of spitballing.
- baryphonic 6y agoTotally agree. I was thinking he'd get there, and then the post abruptly ended.
- brigade 6y agoThe dependency chain is state += 0x60bee2bee120fc15ull or (state += UINT64_C(0x9E3779B97F4A7C15)); the rest of the calculations are independent per iteration. Anyway, the more important fact is that 64x64b -> 128b mul might be one instruction on x86, but it's broken into 2 µops. Because modern CPUs generally don't design around µops being able to write two registers in the same set.
- titzer 6y agoIt's a shame we can't see the rest of the code. What is happening to the result value? Is it being compared to something? Put into an array, or what? All of that code probably totally outweighs what you pointed out here. Or, at least it should. I have a bad feeling it might be being dead-code eliminated, since compilers are super aggressive about that nowadays, but I hope he's somehow controlled for that.
- Daho0n 6y agoThe amount of bugs in the M1 and MacOS posted on HN in a week could keep developers working for months at Apple.
- zelon88 6y agoYou mean to tell me that a $2000 Macbook is almost as performant as a $1000 PC? Tell me more!
- neogodless 6y agoBased on U.S. prices, it's more like $999 vs $609 for similar specs (but no doubt a nicer machine and much better screen/touchpad with the Air.) https://www.apple.com/shop/buy-mac/macbook-air https://www.apple.com/shop/buy-mac/macbook-air https://www.amazon.com/Lenovo-IdeaPad-Laptop-Newest-Display/dp/B08BDJQKH1 https://www.amazon.com/Lenovo-IdeaPad-Laptop-Newest-Display/...
- mpweiher 6y agoBoth Minis and Airs start at under $1000, and they're all the same speed.
- immigrantsheep 6y agoAir starts at $1500 if you're not in the USA
- Toutouxc 6y ago$1345 in my country, $1265 with edu discount.
- 6y ago
- ben_bai 6y agoThat's great if you App is compute bound. "May all your Processes be compute bound." Back in the real world most of the time your Process will be io bound. I think that's the real innovation of the M1 chip.
- isitdopamine 6y agoExplain please. What does the M1 do to IO loads?
- 1_player 6y agoNothing. Compute speed isn't that important if you're waiting on IO is GP's point.
- K0balt 6y agoOn die memory and storage. No bottlenecks, very little latency.
- yuhong 6y agoI believe that ARMv8 NEON crypto extensions has a special instruction for 64-bit multiply to 128-bit product, which is useful for Monero mining for example.
- namibj 6y agoFor the interested, LLVM-MCA says this Iterations: 10000 Instructions: 100000 Total Cycles: 25011 Total uOps: 100000 Dispatch Width: 4 uOps Per Cycle: 4.00 IPC: 4.00 Block RThroughput: 2.5 No resource or data dependency bottlenecks discovered. , which to me seems like 2.5 cycles per iteration (on Zen3). Tigerlake is a bit worse, at about 3 cycles per iteration, due to running more uOPs per iteration, by the looks of it. For the following loop core (extracted from `clang -O3 -march=znver3`, using trunk (5a8d5a2859d9bb056083b343588a2d87622e76a2)): .LBB5_2: # =>This Inner Loop Header: Depth=1 mov rdx, r11 add r11, r8 mulx rdx, rax, r9 xor rdx, rax mulx rdx, rax, r10 xor rdx, rax mov qword ptr [rdi + 8*rcx], rdx add rcx, 2 cmp rcx, rsi jb .LBB5_2