9 ms·
GPU advancements in M3 and A17 Pro [video]
- gary_0 3y agoI skimmed the video but a lot of it sounded more like advertising than technical information to me. On the other hand, I'm looking forward to watching the Asahi folks crack this stuff open.
- wmf 3y agoThe beginning does sound like marketing but eventually it gets into technical information.
- reroute22 3y agoI'd say the beginning sounds like an introduction to GPU architectures in general, not marketing. Somewhere in the ballpark of 5:30-6:00 or so it describes prior hardware design of the Apple's shader core, and starting 7:00 it goes into hardware design of the new M3/A17Pro shader core. It's actually surprisingly detailed, e.g. Nvidia's whitepapers provide less detail on the actual organization of their SMs.
- GuestHNUser 3y agoIt is an overview of what is new about their hardware and technical advice (but not overly technical) on how to maximize the GPUs performance. I watched the full video and thought it was excellent. I wish other CPU/GPU manufacturers made technical overview videos like this. I've never programmed graphics targeting metal before, but I feel much more inclined after watching this so I guess it was good advertising.
- ytch 3y agoI didn't watch too many developer conference, but IHMO the Apple WWDC sessions are fairly easy and enjoyable for learning. especially they provide full transcript.
- pjmlp 3y agoMaybe it is my former graphics background, but I find GTC videos from NVidia and their developer site quite understandable, as AMD GPUOpen as well. Have you dived into them?
- ytch 3y agoIt's one of ̶W̶W̶D̶C̶ developer sessions. Some of them at WWDC are basic 101 introduction for rookies. There are advanced (for me) sessions like: https://developer.apple.com/videos/play/wwdc2023/10127/ https://developer.apple.com/videos/play/wwdc2023/10127/ https://developer.apple.com/videos/play/wwdc2023/10042/ https://developer.apple.com/videos/play/wwdc2023/10042/ Although it's true that they won't discuss at hardware level.
- tux1968 3y ago> On the other hand, I'm looking forward to watching the Asahi folks crack this stuff open. It will still be years before it is practical for Linux developers to target these features. Eventually, the rate of change in GPU design will slow and Linux will catch up once and for all. But it's hard to not drool over the hardware that proprietary OSs get to use today.
- kanwisher 3y agoit literally has code and diagrams of how to organize SIMD instructions to fill out all the shader cores
- saagarjha 3y agoI think it's important to understand that this video is aimed at a broad audience of developers. Obviously some of them are GPU experts but a lot of people are quite literally app developers who are going to watch this video as an introduction to "hmm everyone keeps telling about this GPU thing, what can I do with it?" So the video has to provide them with some context they can use to explore more.
- xign 3y agoWhat's not technical about it? It tells you what the new advancements in the M3/A17 GPUs are, what you should look out for, and how you can take advantage of them. It provides enough information so you can understand what technical tradeoffs you need to make when you target M3 / A17 GPUs. E.g. Register pressure is a real concern in large shaders, and this helps explain how the behavior would be different under a dynamic allocation scheme. It explains how the ray tracing acceleration works, and how it reorders the different intersection calls and how you should avoid intersection queries. There are some hyperbole interjected about how incredible the performance is but that's only in between the useful data. (I did chuckle at the… enthusiasm of the speaker though) This isn't a technical document for GPU designers. Apple doesn't really need or want you to understand exactly how the implementation works because that's basically trade secret for them. This is aimed at letting app / game developers know how they should optimize for the new GPUs, since previously Apple just made some ambiguous remarks about some of these new technology ("Dynamic Caching") without explaining what they meant. But yes, I do like how the Asahi folks tend to end up documenting a lot of how these hardware works, but they also only have public information like this to start from so these are still useful info to have for them.
- diimdeep 3y agoReally awful narration with overuse of unnecessary pitch glides
- goosinmouse 3y agoIt is pretty funny how amateur a multi trillion dollar can be, its so distracting and sounds like an 80's instructional VHS.
- dylan604 3y agoyou could never hear the audio that cleanly on a VHS even with HiFi tracks.
- diimdeep 3y agoIt is about attitude in voice, not about audio quality. It is very noticeable at 1.5x speed, 1x speed is too slow anyway with infinite pauses for such low density technical details.
- tambourine_man 3y agoIt’s an engineer and they provide a transcription.
- riscy 3y agothe narrator is an actual engineer, not a television personality
- diimdeep 3y agobut he is cosplaying television personality for no reason
- deleted 3y ago[deleted]
- alphanullmeric 3y agoIt’s amazing how bad the competition is. The A17 pro has 2 performance cores and 4 efficiency cores. The Google G3 has 9 cores of 3 different types, the fastest being slower than Apple’s performance cores, the most efficient being less efficient than apple’s efficiency cores. And it’s a phone so you don’t take advantage of the extra parallelism. You just get the worst of both worlds. no wonder these android phones have 50% more battery and 50% less battery life. Is it that hard to just copy the winning formula?
- zmmmmm 3y agohere have some Qualcomm Kool-Aid to go with that sweet Apple juice :-) https://www.youtube.com/watch?v=h_vh7_n_OPs https://www.youtube.com/watch?v=h_vh7_n_OPs
- hutzlibu 3y agoI don't know about your facts, but "Is it that hard to just copy the winning formula?" yes it is, thanks to IP law. And back in the day Steve Jobs already wanted thermonuclear war on Samsung, because he felt their flagship at the time was too close to the IPhone.
- KingLancelot 3y agoSteve Jobs wanted nuclear war with Google (not Samsung) because Eric Schmidt was on Apple’s board of directors while the iPhone was being developed, so Jobs felt Schmidt was basically doing insider trading for google to develop Android.
- hutzlibu 3y agoAh, I did not remember that one about Eric Schmidt, but I thought there was also something with Samsung. Btw. it appears you are shadow banned. You might want to check some of your comments and then contact dang.
- rootusrootus 3y agoTo be fair, there was a moment when Samsung was in full copy mode. All the way down to having their own version of the dock connector and a retail box that closely mimicked Apple. In retrospect, a bit embarrassing for a company we know is capable of much more.
- zmmmmm 3y agoDoes apple document exactly how many actual true cores there are inside their GPUs? It is always confusing they say "40 core GPU" but I assume these are shader cores which each inside them can execute (per the video) "many thousands" of parallel execution paths. So how does one translate to an equivalent in "CUDA cores" type terminology?
- kergonath 3y ago> So how does one translate to an equivalent in "CUDA cores" type terminology? I don’t really think we can, even if we knew exactly what is in a M3 GPU core, which we don’t. Both architectures are very different, and different again from AMD GPUs. We have to count Tflops.
- wmf 3y agoEach Apple core (heh) has 128 FPUs so 40 cores would be akin to 5120 CUDA "cores".
- pixelpoet 3y agoNot quite, because Nvidia counts dual issue as a flat doubling of "core" (which previously you could accurately call a vector lane) count.
- pavlov 3y agoWhat Apple calls a GPU core seems to be roughly the same as what Nvidia calls a “stream multiprocessor”. For example a 1080 GTX GPU has 20 stream multiprocessors (SM), each containing 128 cores, each of which supports 16 threads. Meanwhile Apple describes the M1 GPU as having 8 cores, where “each core is split into 16 Execution Units, which each contain eight Arithmetic Logic Units (ALUs). In total, the M1 GPU contains up to 128 Execution units or 1024 ALUs, which Apple says can execute up to 24,576 threads simultaneously and which have a maximum floating point (FP32) performance of 2.6 TFLOPs.” So one option to get a single number for a rough comparison is to count threads. The 1080 GTX supports 40,960 threads while the M1 supports 24,576 threads. There’s obviously a lot more to a GPU — for starters, varying clock speeds, ALUs can have different capabilities, memory bandwidth, etc. But at least counting threads gives a better idea of the processing bandwidth than talking about cores.
- w10-1 3y agoThe video is just enough of a peek into the GPU's to encourage people to write using Metal API's (and by the way, use the new APIs and FP16).
- jeffybefffy519 3y agoThey should just support directx. Devs will never support two graphics api’s. it costs too much especially to grab the marginal mac os share that has powerful enough gpu’s. Id bed in 4 years apple moves to directx.
- nemothekid 3y agoDirectX is exclusive to the Windows platform. At this point, it's probably deeply tied into Windows. I don't see how you can make that bet.
- reroute22 3y agoI've no idea exactly how MS licenses uses of DX, but just for context Imagination Technologies just released a custom GPU design that implements DirectX Feature Level 11_0 (which corresponds to earlier versions of DX 12 [1]): https://www.imaginationtech.com/news/imagination-launches-brand-new-line-of-high-performance-gpu-ip-with-directx/ https://www.imaginationtech.com/news/imagination-launches-br... Imagination Technologies is a near 40 year old British silicon IP company that has been doing GPUs for quite some time, just not ones supporting DX up until now, and it has nothing to do with MS (in terms of ownership / rights / etc). [1] https://learn.microsoft.com/en-us/windows/win32/direct3d11/overviews-direct3d-11-devices-downlevel-intro#direct3d-12-feature-support-feature-levels-12_2-through-11_0 https://learn.microsoft.com/en-us/windows/win32/direct3d11/o...
- monocasa 3y agoThe way that works isn't that IMG ships a full directx implementation, but that they ship some kernel and user mode components that plug in to Microsoft's directx implementation such that when taken together, directx is accelerated by IMG's hardware. Similarly, Microsoft would need to release the non GPU specific bits for macos to fit the same model.
- Reason077 3y agoI’m always impressed with the speech synthesis that Apple uses to make the voiceovers in these videos. Some of them almost sound like real people!
- stingraycharles 3y agoThe guy introduced himself by name, I was really confused for a while if it was just a human trained to sound like an AI, or an AI trained to sound like a human.
- deleted 3y ago[deleted]
- Reason077 3y agoPerhaps they ask the AI speech/content generator to give itself a name as part of its training prompt.
- franzb 3y agoFrom the variety of intonations based on context, I doubt it’s speech synthesis.
- SushiHippie 3y agoThey even created a Twitter profile for this AI persona https://nitter.net/jhaberstro https://nitter.net/jhaberstro
- leloctai 3y agoDoes the complex block in the diagram refer to complex numbers? That doesn't sound typical, does it? What type of work load that typically run on the GPU that would require complex numbers?
- make3 3y agoquaternions? for rotations
- reroute22 3y agoJudging by the output/GUI of their GPU profiler, "complex" there is more like "complex instructions", think f32 (floating point) ops that aren't additions and multiplications (and FMAs), but trigonometry, square roots, that sort of thing.
- mattsan 3y agoFFT plus some game stuff requires complex numbers to do partial rendering (e.g. do some now and then do more next frame - I've lost the link to the talk but IIRC EA did a talk on how they made a shader that emulates lights in the background that are out of focus (not Guassian but the actual cool effect as if it was a real camera))
- sccxy 3y agoDoes M3 still outputs garbage which make external displays flicker? https://forums.macrumors.com/threads/m1-m2-flickering-ghosting-with-external-display-merged.2271670/ https://forums.macrumors.com/threads/m1-m2-flickering-ghosti... https://www.benq.com/en-us/knowledge-center/knowledge/how-to-fix-mac-m1-m2-external-monitor-flicker.html https://www.benq.com/en-us/knowledge-center/knowledge/how-to... https://www.howtogeek.com/805459/mac-flickering-external-screen/ https://www.howtogeek.com/805459/mac-flickering-external-scr...
- monocasa 3y agoInterestingly, that's a different component than the GPU on these chips, which is the typical architecture in SoCs. In fact, even in discrete GPUs, the display scanout engine is generally a nearly completely independent block relative to the rest of the GPU.
- lwkl 3y agoThis issue can be fixed by switching the display to RGB [1]. So I think it’s a software bug but it‘s really annoying since the fix sometimes resets and the bug only occurs when there is a lot of black on the screen. [1] https://gist.github.com/GetVladimir/c89a26df1806001543bef4c8d90cc2f8 https://gist.github.com/GetVladimir/c89a26df1806001543bef4c8...
- sccxy 3y agoI have two monitors connected to M1 mac. One works perfectly fine and is automatically RGB. Other flickers and when changing to RGB mode it is lime green. Wonder why it is not problem with Intel Macs and if M3 fixes those bugs? Maybe it is Apples feature to sell more of their own monitors. They make sure other high end brands do not work with macOS. It does not make sense that this kind of bug is 3 years active.
- Someone 3y ago> Wonder why it is not problem with Intel Macs and if M3 fixes those bugs? Let’s guess: maybe they have different drivers? Would be far from surprising, given that they’re running on different processors.
- reroute22 3y agoOkay, I went through the other video they reference ("Discover new Metal profiling tools for M3 and A17 Pro" [1]), and there is actually a whole bunch of extra very relevant (IMO) information on the subject, starting about 13:30 or so. [1] https://developer.apple.com/videos/play/tech-talks/111374?time=818 https://developer.apple.com/videos/play/tech-talks/111374?ti...
- pjmlp 3y agoThere are additionally related videos, "Discover new Metal profiling tools for M3 and A17 Pro" https://developer.apple.com/videos/play/tech-talks/111374/ https://developer.apple.com/videos/play/tech-talks/111374/ "Learn performance best practices for Metal shaders" https://developer.apple.com/videos/play/tech-talks/111373/ https://developer.apple.com/videos/play/tech-talks/111373/ "Bring your high-end game to iPhone 15 Pro" https://developer.apple.com/videos/play/tech-talks/111372/ https://developer.apple.com/videos/play/tech-talks/111372/
- runeks 3y ago> I'm excited to tell you about the new Apple family 9 GPU architecture in A17 Pro and the M3 family of chips, which are at the heart of iPhone 15 Pro and the new Max. "The new Max"? He clearly meant "the new Macs". Kinda weird that Apple can't properly transcribe its own content.
- Tijdreiziger 3y agoI’ve also noticed this on lots of YouTube videos, where the creator clearly meant one thing, but the subtitles substitute a more common, similarly-sounding word with a different meaning. I suspect they have the videos transcribed externally, and don’t check the transcription (or only do so in a cursory manner).
- TheCapeGreek 3y agoOr automated transcription. For YT vids, especially shorts, it's because churning out shorts/reels/tiktoks of clips from longer form videos (and/or with the split screen gameplay of some mobile game/minecraft platforming run) is now a common tactic for trying to gain tons of views on your account for monetisation later.
- Tijdreiziger 3y agoI’ve also frequently seen it on long-form videos. I think the transcriptions must be at least partially reviewed by humans, because YouTube already has automatic transcription for videos without subtitles.
- eviks 3y agoWhy is it surprising, it's not like content ownership gives you any advantage in the typical transcription algorithms
- runeks 3y agoIt’s surprising since I would expect Apple to check whatever they get back from a transcription service.
- TradingPlaces 3y agoMedia is so focused on CPUs, they are missing the fact that Apple focused on the GPU and Neural Engine for this round of chips
- wincy 3y agoIt was wild seeing Linus Tech Tips demoing resident evil village on the iPhone 15 Pro.
- rsynnott 3y agoThis is unfortunately inevitable; CPUs are just so much easier to benchmark in a broadly useful way. And the extreme leakiness of geekbench is helpful (I suspect Apple sees this as a feature; most recent Apple chip iterations have leaked on geekbench)
- javaunsafe2019 3y ago[flagged]
- josu 3y agoDoes the narration sound like AI to anyone else?
- frogblast 3y agoIf you're interested in more background about one user-visible problem being directly attacked by this new GPU architecture, that could be "shader compilation stutter" (although there are many others). These are two excellent posts that go deep on this: The Shader Permutation Problem - Part 1: How Did We Get Here? The Shader Permutation Problem - Part 2: How Do We Fix It? In particular, the second post has the line: We probably should not expect any magic workarounds for static register allocation: if a callable shader requires many registers, we can likely expect for the occupancy of the entire batch to suffer. It’s possible that GPUs could diverge from this model in the future, but that could come with all kinds of potential pitfalls (it’s not like they’re going to start spilling to a stack when executing thousands of pixel shader waves). ... And some kind of 'magic workaround for static register allocation' is pretty much what has been done. https://therealmjp.github.io/posts/shader-permutations-part1/ https://therealmjp.github.io/posts/shader-permutations-part1... https://therealmjp.github.io/posts/shader-permutations-part2/ https://therealmjp.github.io/posts/shader-permutations-part2...
- RantyDave 3y agoSo ... the registers are dynamically allocated from a chunk of cache? Does this mean there, effectively, are no registers? Does this cache have one clock latency?
- reroute22 3y agoI doubt anyone will be able to answer questions this fine grained, not now (if the implementation is architecturally exposed - leaks into the ISA - and Asahi Linux group figures some of it out), or possibly not ever (if it's architecturally transparent and thus entirely micro-architectural). > Does this mean there, effectively, are no registers? I can only point out just for context that if by any chance you're asking whether the registers are implemented as actual hardware design "registers" - individually routed and and individually accessible small strings of flip-flops or D-latches - then the history of the question is actually "it never was registers in the first place" - architectural (ISA) registers in GPUs are implemented by a chunk of addressable ported SRAM, with an address bus, data bus, and limited number of accesses at the same time and limited b/w [1]. [1] see the diagram at https://www.renesas.com/us/en/products/memory-logic/multi-port-memory/asynchronous-dual-port-rams/7007-32k-x-8-dual-port-ram#overview https://www.renesas.com/us/en/products/memory-logic/multi-po...
- RantyDave 3y agoOh! Well, that explains that then. Wild!
- reroute22 3y agoThere is a fairly informative survey on the subject: https://www.osti.gov/servlets/purl/1332070 https://www.osti.gov/servlets/purl/1332070 (A Survey of Techniques for Architecting and Managing GPU Register File)
- reroute22 3y agoAn easier to read research article that's narrower in subject and seemingly more relevant to the OP: https://research.nvidia.com/sites/default/files/pubs/2012-12_Unifying-Primary-Cache/Gebhart_MICRO_2012.pdf https://research.nvidia.com/sites/default/files/pubs/2012-12... ("Unifying Primary Cache, Scratch, and Register File Memories in a Throughput Processor", 2012)