34 ms·
Dissecting the Apple M1 GPU, part I
- osamagirl69 6y agoExciting! I can't believe the incredible rate at which progress is being made on this front. And by a sophomore(?!) undergraduate student no less!
- jesse_cureton 6y agoAlyssa was one of the primary developers behind the Panfrost open source drivers for a subset of Mali GPUs - she's a brilliant engineer and great to work with.
- pantalaimon 6y agoShe also wrote panfrost while still in high school.
- azinman2 6y agoIs there an interview / write up with her somewhere? That's incredible for most people let alone a high schooler.
- leetreveil 6y agohttps://blogs.gnome.org/engagement/2019/10/03/alyssa-rosenzweig-panfrost/ https://blogs.gnome.org/engagement/2019/10/03/alyssa-rosenzw...
- l1k 6y agoXDC 2018 - Lyude Paul & Alyssa Rosenzweig - Introducing Panfrost https://www.youtube.com/watch?v=qtt2Y7XZS3k https://www.youtube.com/watch?v=qtt2Y7XZS3k
- ralfd 6y agoI guess both presenters are trans? Interesting how many notable trans* programmers/engineers the field attracts.
- ognarb 6y agoYou would be surprised by the amount of high-schoolers doing incredible work in open source software.
- tomnipotent 6y agoLike George Hotz, who at 17 removed the SIM lock on the iPhone and a few years later got sued by Sony for breaking security on the PlayStation 3.
- Medox 6y agoObligatory George Hotz (geohot) PS3 rap from back then: https://www.youtube.com/watch?v=9iUvuaChDEg https://www.youtube.com/watch?v=9iUvuaChDEg First time he appeared in the PS3 jailbreak scene I was like "wait, is this a Romanian?! Some boy-genius expat?!" because George is a common first name here and (fun fact) Hotz means "thief". Actually it is written hoț but Romanian teens sometimes write tz instead of ț. He was meant to be a key thief, I guess. I still wonder if they got to examine the video in the courtroom...
- marcan_42 6y agoAs someone who has seen both of them work (and who got sued on the same lawsuit)... I wish people would stop putting geohot up on a pedestal. His entire career has been about self-promotion (something which he is very good at), but he got his start taking credit for other people's work (that SIM lock stuff was largely just implementing stuff other people told him about). He even took credit from my/our (fail0verflow's) work when he did the thing that got him sued (published some PS3 private keys), which was just him implementing the ECDSA attack we had detailed in a presentation a week prior, without a single mention until I emailed him to ask him to do so. Alyssa's Panfrost work is leaps and bounds ahead of anything geohot has ever personally done. Pretty much everyone of our profile gets started at around that age (I was porting Linux to a weirdo ISP router at 14 and working on Wii hardware and homebrew, the first "well known" thing I did, starting at 16-17), but somehow geohot has marketed himself to be some kind of genius when he's one of the more mediocre hackers I've known. He's not bad these days, but he has always represented his skills as being way better than they actually are, and taken solo credit for work that involved other people, his entire career.
- vaxman 6y agoProgress? They're forcibly hacking something that Apple doesn't want them to have. If this was the iPhone, that would be one thing (hint: Apple has a monopoly on the productive consumers and via AppStore, developers), but the Mac is and has always been private and the suggestion it is a monopoly is laughable. There is no justification for this, other than "I can haz Linux on my Air/Mini hilk hilk."
- Veedrac 6y ago> Yet Metal optimization resources imply 16-bit arithmetic should be significantly faster, in addition to a reduction of register usage leading to higher thread count (occupancy). I believe this is a difference between the A14 GPU and the M1 GPU; the former's 32 bit throughput is half its 16 bit throughput, whereas on the latter they are equal.
- MayeulC 6y ago> This suggests the hardware is superscalar, with more 16-bit ALUs than 32-bit ALUs To me, it sounds like it might mean 32-bit ALUs can be used as two 16-bit ones; that's how I would approach it, unless I'm missing something? The vectorization can also happen at the superscalar level, if borrowing the instruction queue concept from out-of-order designs: buffer operations for a while until you've filled a vector unit's worth, align input data in the pipeline, execute. A smart compiler could rearrange opcodes to avoid dependency issues, and insert "flushes" or filling operations at the right time.
- bullen 6y agoI recently added half floats to my 3D MMO engine, and I was very dissappointed when I discovered that very few GPUs support it in hardware! Desktop Nvidia simply converts them to 32-bit floats up front so the performance is kept but memory is wasted, but raspberry does the 16 -> 32 bit calculations every time resulting in horrible performance. I still have to test the engine on Jetson Nano with half floats, but I'm pretty sure I will be dissapointed again, and since raspberry doesn't support it I need to backtrack the code anyway! After some further research I heard snapdragon has 16-bit support in the hardware and hopefully this is where we are heading! 32-bit is completely overkill for model data, it wastes memory and cycles! 16-bit for model data: vertex, normal, texture and index and 8-bit for bone-index and weights! Back to the 80-90s! This is the last memory & performance increase we can grab now at "5"nm without adding too much complexity! You can try the engine without half float here: http://talk.binarytask.com/task?id=5959519327505901449 http://talk.binarytask.com/task?id=5959519327505901449
- kllrnohj 6y ago> Desktop Nvidia simply converts them to 32-bit floats up front so the performance is kept but memory is wasted Pascal has native FP16 operations and can execute 2 FP16's at once ( https://docs.nvidia.com/cuda/pascal-tuning-guide/index.html#fp16 https://docs.nvidia.com/cuda/pascal-tuning-guide/index.html#... ) BUT, and this is where things get fucked up, Nvidia then neutered that in the GeForce lineup because market segmentation. In fact, it's slower than FP32 operations: "GTX 1080’s FP16 instruction rate is 1/128th its FP32 instruction rate" https://www.anandtech.com/show/10325/the-nvidia-geforce-gtx-1080-and-1070-founders-edition-review/5 https://www.anandtech.com/show/10325/the-nvidia-geforce-gtx-...
- Anka33 6y ago"He a brilliant engineer and great to work with." Fixed that for you.
- ksec 6y ago>Some speculate it might descend from PowerVR GPUs, as used in older iPhones, while others believe the GPU to be completely custom. But rumours and speculations are no fun when we can peek under the hood ourselves! As per IMG CEO, Apple has never not been an IMG customer. ( Referring to period between 2015 and 2019.) Unfortunately that quote, along with that article has simply vanished. It was said during an Interview on a British newspaper / web site if I remember correctly. "On 2 January 2020, Imagination Technologies announced a new multi-year license agreement with Apple including access to a wider range of Imagination's IP in exchange for license fees. This deal replaced the prior deal signed on 6 February 2014." [1] The Apple official Metal Feature set document [2], All Apple A Series SoC including the latest A14 supports PVRTC, which stands for PowerVR Texture Compression[3]. It could be a Custom GPU, but it still has plenty of PowerVR tech in it. Just like Apple Custom CPU, it is still an ARM. Note: I am still bitter at what Apple has done to IMG / PowerVR. [1] https://www.imaginationtech.com/news/press-release/imagination-and-apple-sign-new-agreement/ https://www.imaginationtech.com/news/press-release/imaginati... [2] https://developer.apple.com/metal/Metal-Feature-Set-Tables.pdf https://developer.apple.com/metal/Metal-Feature-Set-Tables.p... [3] https://en.wikipedia.org/wiki/PVRTC https://en.wikipedia.org/wiki/PVRTC
- remexre 6y ago> Note: I am still bitter at what Apple has done to IMG / PowerVR. I'm unfamiliar with this; are you bitter about a lack of attribution that they produced lots of the IP their GPUs are built on?
- masklinn 6y agoThey're probably talking about Apple suddenly announcing they'd be dropping Imagination out of the blue in 2017.
- raverbashing 6y agoBut in the end it seems they didn't drop? Or are they just licensing a basic IP and building on top? I remember the IMG stuff being full of bugs
- MangoCoffee 6y agohopefully, Apple M1(ARM) and AMD Zen(x86) can push Intel, Intel stuck at 14nm is what happened when there is no viable competitor.
- wil421 6y agoCan someone explain why they would buy a Mac and install linux? I used to dual boot windows for school and when I first switched to windows I had an old laptop's backup in bootcamp. Cross-platform software is much more ubiquitous than 5-10 years ago. For linux I always used another box or just run a VM. Nowadays my laptop can ssh or Remote Desktop into a more powerful machine. I have a custom built widows box, a custom built NAS running FreeNas (FreeBSD), 2 RPi's running Raspbian, and a not always linux box based on old hardware. There is a machine big or small to do things or play with. My VPN allows me to connect from anywhere. What are you guys doing that you have to install linux instead of running a VM or remotely connecting to a linux box? If it's just for the sake of knowledge I can understand it. Apple's touch pad experience in MacOS is the best in the market and it is always very different in Windows and Linux. The XPS, Lenovo and smaller vendors really make killer Linux/Windows laptops that have much more options than Macs.
- londons_explore 6y agoRight now the draw is the M1 CPU... If Linux had software support for it, it would probably be the best platform for Linux (server) software development.
- wmf 6y agoThe result may be the fastest Linux machine available (by some metrics).
- sliken 6y agoSingle core performance, maybe. Not much else though. But the apple's do quite well on perf/$, perf/watt. The M1 mini is pretty competitive at $700 for a fast silent desktop.
- wlesieutre 6y agoYou mentioned performance/watt - in user terms, battery life is a huge advantage of this hardware. The official specs of the 13" MBP are 17 hours of web browsing, 20 hours of video playback. I assume it won't be as long in Linux but it'll still be longer than a lightweight and high performance x86 laptop.
- londons_explore 6y agoWhy start with the shader compiler rather than simply whatever commands are necessary to get the screen turned on and a simple memory mapped framebuffer? It would seem easier to get Linux booting (by just sending the same commands apples software does) before worrying about 3d acceleration and shaders...
- wmf 6y agoIIRC the firmware turns the screen on for you so it's already there when Linux boots.
- StillBored 6y agoYes, this is a standard UEFI feature. The firmware basically hands off a structure indicating horizontal/vert resolution, pixel depth and a buffer. Which is fine until you need to change the mode, or start doing heavy BLT/etc operations. At which point your machine will feel like something from the early 1990s. So yes, you can get a full linux DE running with mesa and CPU GLES emulation, but its not really particularly usable.
- DCKing 6y agoIt's worth noting that the M1 does not use EFI like Intel Macs do, so the kernel would still need to boot on this proprietary new mechanism. Apple previously indicated that the M1 Macs boot using a derivative of the iPhone/iPad boot procedure. There's a reasonable chance it'll still have something like the EFI buffer allocation mechanism, but it'll not be using a mainline standard. While it sucks that Apple ditched EFI, the good news is that it's probably reasonably well understood since Linux has already been booted on an iPhone 7 using Checkra1n [1] (and apparently with framebuffer support too!). [1]: https://tuxphones.com/iphone-7-now-boots-postmarketos-linux/ https://tuxphones.com/iphone-7-now-boots-postmarketos-linux/
- q3k 6y agoI'm not nearly as experienced as the blog post author, but I suspect the M1's GPU does not have any way of blitting to the screen other than driving it 'fully' - ie. using the standard command queue and shader units to even get the simplest 2D output.
- frozenport 6y ago>>There are no convoluted optimization tricks, but doing away with the trickery is creating a streamlined, efficient design that does one thing and does it well. Maybe Apple’s hardware engineers discovered it’s hard to beat simplicity. What? They shifted the complication from software to hardware.
- camdenlock 6y agoI wonder how much time is left before Apple locks down macOS so hard that this type of analysis becomes impossible without a jailbreak. :/
- deleted 6y ago[deleted]
- IfOnlyYouKnew 6y agoBased on some similarities between this idea and a certain strand of thinking in nuclear engineering, I believe MacOS will be locked down at around the same time as fusion power becomes commercially viable, i. e. twenty years from now.
- rswail 6y agoIt's been Real Soon Now for at least 10 years...
- weeboid 6y agoHow many FPS does it achieve on … is there an FPS standard (FPS) game for Mac GPU benchmarking?
- Synaesthesia 6y agoIt's about equivalent to an RX560 or NVIDIA 1030 I believe
- pabs3 6y agoThere is the implication upthread that the M1 GPU is similar to or derived from the IMGTEC PowerVR GPUs, anyone know if that is the case? Also, I noticed that IMGTEC are going to write an open source mesa driver, I wonder how much code the two drivers will be able to share. https://riscv.org/blog/2020/11/picorio-the-raspberry-pi-like-small-board-computer-for-risc-v/ https://riscv.org/blog/2020/11/picorio-the-raspberry-pi-like...
- hajile 6y agoThey pay multi-millions to license Imagination Tech patents and poached most of their best designers. Even with a custom design, those designers are likely to think along the same lines (and Imagination Tech IP is pretty good in it's own right). In any case, I doubt Apple would pay that much money if they didn't need rights to the IP. The only real question is why they didn't acquire the company outright.
- phire 6y agoApple deliberately designed their custom GPU to have the same performance characteristics as the PowerVR gpus that early iphones used. Apple GPUs support the proprietary IMGTEC compressed texture format that nobody else uses and they use the same deferred rendering method. Apple's deception was good that nobody realised they were shipping custom GPUs. Even with hindsight, nobody can really tell when apple shipped their first custom GPU. Apple may have even shipped GPUs which were a mixture of IMG and Apple RTL. But at this point the internals are far enough appart that I doubt much code can be shared.
- kbob 6y agoI wonder whether this work could eventually enable a native Vulkan port for MacOS on this GPU.
- mlindner 6y agoDissecting something without a single graph or illustration or image. For shame....