19 ms·
Temptation of the Apple: Dolphin on macOS M1
- toyg 5y agoFtr, this is Dolphin the games emulator, not Dolphin the KDE file manager.
- amelius 5y agoAnd also not Dolphin, the MySQL logo.
- agustif 5y agoAlso not Delphi, for the dyslexics like myself out there
- drrotmos 5y agoAnd also not Ecco the Dolphin (the game).
- hotpickles 5y agoAnd also not Dolph Lundgren (the actor and chemical engineer).
- ihuman 5y agoAlthough you might be able play the the Wii virtual console version using the Dolphin emulator
- hnlmorg 5y agoNor Flipper -- the cult TV show that featured a crime fighting dolphin (I kid you not).
- deleted 5y ago[deleted]
- djhonovak 5y agoAlso not Dolph Lundgrin.
- xvector 5y agoNor Flipper, the tamogatchi-like hacking device [1]. [1]: https://flipperzero.one/ https://flipperzero.one/
- Andrex 5y agoNor Flipper, the GameCube's GPU [1]. 1. https://en.wikipedia.org/wiki/GameCube_technical_specifications https://en.wikipedia.org/wiki/GameCube_technical_specificati... But I guess it's somewhat related. :P
- deleted 5y ago[deleted]
- Macha 5y agoAlthough I would love to replace Finder with Dolphin.
- pantulis 5y agoI truly came here thinking about Dolphin the mobile browser.
- elondaits 5y agoNot Dolphin Smalltalk :-(
- MonkeyIsNull 5y agoYup, sad -- that's EXACTLY what I thought it was as well.
- jhgb 5y agoWould have been weird for it to support Mac on ARMv8 before Windows on AMD64.
- dvfjsdhgfv 5y agoAlso not Dolphin, a web browser for Android.
- tdonovic 5y agoIncredible the perf they get out of it. Bit confused with the graph towards the end, is perf better under Rosetta than natively?!
- Macha 5y agoThe rosetta vs native vs 9900k vs 8559h graph? The only game rosetta is beating native on is rogue squadron 2. Since Dolphin is a JIT, this seems to be a case of where Rosetta's JIT is smarter than Dolphin's in terms of which ARM instructions are chosen when converting from the Intel instructions than Dolphin when converting from the emulated PPC instructions. Unless you're comparison is the 8559h and not the "native" bar. I mean, the 8559h is a mid range older Intel CPU and it's hard to understate how much Intel stagnated since Sandy Bridge (and especially since Skylake).
- leoetlino 5y agoRosetta is faster than native in that case because the AArch64 JIT has to fall back to the interpreter for memchecks (unlike the x86-64 JIT).
- Aaargh20318 5y ago> Since Dolphin is a JIT, this seems to be a case of where Rosetta's JIT is smarter than Dolphin's in terms of which ARM instructions are chosen when converting from the Intel instructions than Dolphin According to the article, the AArch64 JIT isn’t as complete as the x86 one so some less common instructions are emulated, not JITed. I imagine a game that uses a lot of these is slower with the native ARM version.
- rvz 5y ago> There's undeniable excitement for the next generation of AArch64 hardware to see how much further that this can go. This is what I am looking forward to. In the case for Apple Silicon, the next generation will be even better and is not far off from announcing the newer processors that will supersede M1 in WWDC. The M1 only shows what's possible for Apple Silicon and the newer generation of ARM-based Macs will impress further. So will skip this one for now and wait what WWDC has to offer for the next generation.
- rationalData 5y agoIn 5 years there will be 1nm transistors. If you skip this generation, you might as well wait another year to get high end stuff and you won't have to deal with Apple's less than friendly practices.
- dimitrios1 5y agoI traded in my inferior Intel based 2020 13 inch macbook pro for the M1 MBP and it only costed me 100 dollars, so I went ahead and pulled the trigger. The battery life was abysmal on the intel based and I was desperate. I love the form factor and the size, I just never could get into any real programming tasks without being strapped to a power source.
- faitswulff 5y agoYour comment made me realize I’m waiting for the next generation Apple Silicon SoC as a second data point. Even if I don’t buy it, it will tell us something about the expected performance trajectory.
- dougmwne 5y agoWow, that frames per watt graph is an eye opener for sure. What an incredible advancement in mobile computing.
- ninjinxo 5y agoWhilst very impressive, it's a bit exaggerated, they should have been locked to the same framerate for comparison: * 9900k is boosting to 5ghz which is sacrificing efficiency. * 9900k PC is delivering a much higher framerate, so it'd also have much higher GPU utilisation. * Afaik RTX3090 will have high power draw even at low utilisation (large card, lots of memory). From anandtech: >Should users be interested, in our testing at 4C/4T and 3.0 GHz, the Core i9-9900K only hit 23W power. Doubling the cores and adding another 50%+ to the frequency causes an almost 7x increase in power consumption. https://www.anandtech.com/show/13400/intel-9th-gen-core-i9-9900k-i7-9700k-i5-9600k-review/21 https://www.anandtech.com/show/13400/intel-9th-gen-core-i9-9... Look at the 3090s power consumption during media playback: https://www.techpowerup.com/review/zotac-geforce-rtx-3090-trinity/29.html https://www.techpowerup.com/review/zotac-geforce-rtx-3090-tr...
- Synaesthesia 5y agoWell that actually shows how impressive the M1 is because it hits faster CPU than the 9900k at 5ghz using only 10W-20W total. And GPU it's much faster than the Intel integrated.
- peoplefromibiza 5y agoAs I understand it, OP is annoyed by the graph because they are comparing different things at different scale. I will use another John Deere metaphor: a Prius can cover a much longer distance on the same amount of fuel, but if I need a John Deere it's because the Prius can't do the same job and I am willing to sacrifice fuel efficiency for raw power. In other words: how much more does the i9 consumes to produce the same FPS of an M1? We don't know, but we know power consumption increase on these CPUs is non linear, meaning that the 60-65% of the frame rate could potentially lead to 5-6 times less energy used.
- marsven_422 5y agoThis reads like a native ad.
- Shadonototro 5y agoefficiency test speaks for itself the m1 is the renaissance of laptops
- afavour 5y agoA renaissance would be a revival, surely? Laptops have been dominant for a long time now.
- stu2b50 5y agoLaptops have dominated personal computing (by the mildly unintuitive definition of "PC"), but you could certainly argue that smartphones and tablets have eaten their lunch in the overall computing space.
- G3rn0ti 5y agoThe real challenge is running F-Zero GX. I’d love to see some benchmarks for this game — the hardest game to emulate.
- umanwizard 5y agoWhy’s it hard to emulate that game in particular?
- G3rn0ti 5y agoThe heated air effect on the „Sand Ocean“ course seems to make the emulator sweat. My 2013 Intel Core i7 with HD Graphics can’t render that without massive slow down.
- bri3d 5y agoBy what metric? AFAIK the Factor 5 games, especially Rogue Squadron III, are considered the most challenging, both due to their obscure tricks (iirc Rogue Squadron uploads an outdated audio microcode to get a "loop counter" feature back which no other games use, for example) and most complete use of the MMU mechanisms (I believe they even use the ARAM as swap transparently to the game engine, using some goofy allocator trick) - which is why Rogue Squadron III was chosen for this benchmark.
- G3rn0ti 5y ago> Rogue Squadron III Ok, never played that game. I’ll better try not to play that on my ancient Intel powered laptop ... Different games run differently well on Dolphin if you got older hardware. While Mario Kart Double Dash runs perfectly fine in full screen, „F-Zero GX“ suffers massive slowdown in some levels on my 7 years old CPU/GPU combination. Interestingly, both games employ the „heated air“ effect on similarly looking levels — but still I got 40 FPS vs. 60 FPS in that case. I wouldn’t mind but the sound needs to be in sync with the graphics subsystem on the Game Cube — audio is broken with even slightly slower frame rates, unfortunately.
- jeroenhd 5y agoIt's clear AS is a great advancement in general computing, but every piece like this reads as "the hardware is amazing, and it's totally worth it to work around these arbitrary software restrictions". This performance would've been available on iPads years ago if it wasn't for Apple's blanket ban on JIT and the likes. Apple is one of those companies whose hardware I'd love to have if it wasn't for their software and general corporate decisions. Until I can run a proper version of Firefox on iPad, I'll have to stick with the objectively inferior hardware for the coming years.
- meepmorp 5y agoThis is a really unpopular take on Apple for HN, so I can see why you're getting downvoted.
- threeseed 5y agoYou are either being disingenuous or completely oblivious because the “I want a fully open iOS platform” is the most popular take on Apple.
- meepmorp 5y agoMy god, you're right - it has to exactly one of the two options you posit!
- gjsman-1000 5y agoAnd also it is quite possible that we are in an echo chamber, because outside of hacker news and a few circles like it, there might be less support for this than we wish.
- PragmaticPulp 5y agoThe Write XOR Execute restriction discussed in the article is a security feature, and it’s greatly beneficial from a security standpoint. > Until I can run a proper version of Firefox on iPad FireFox for iOS works just fine. The Gecko vs WebKit difference doesn’t really matter in practice. If you want general purpose computing, just get a Mac. You can run Firefox and any other program you’d like. It would be great if there was an opt-in developer mode on iOS that bypassed certain restrictions, but I also understand why Apple chose to go with security and simplicity as 99.9+% of their customers have no need nor desire to go beyond the security and platform restrictions. I have both an iPad Pro with keyboard case and a MacBook. Even if I could run whatever I wanted on the iPad, I’d still be reaching for the MacBook because it’s just a better physical platform for doing anything other than simple touchscreen and stylus work.
- floatboth 5y ago> [mapping memory WX] hasn't been forbidden on any of the prior platforms that Dolphin supports Well, rarely completely forbidden, but e.g. I think OpenBSD has been W^X by default for quite some time (though IIRC with a WX allowed flag per… FS mount?). Now on FreeBSD it's not default but it's there, and if you turn it on, you have to mark WX-mapping binaries by running `elfctl -e +wxneeded`. Firefox actually became W^X compliant all the way back in 2015: https://jandemooij.nl/blog/wx-jit-code-enabled-in-firefox/ https://jandemooij.nl/blog/wx-jit-code-enabled-in-firefox/
- jolux 5y agoI didn't realize SpiderMonkey was W^X compliant. Does that mean Apple's arguments about third party browser security on iOS are less well-founded than I had believed? My impression was that performant JITs were incompatible with W^X.
- voxic11 5y agoNo the issue on ios isn't that W^X is enforced, its that you can't mark a page that was writable as executable (whereas W^X just implies that a page can't be both writable and executable at the same time). Firefox has been W^X compliant by default since 2016 as its considered more secure in general.
- jolux 5y agoAh I see. The iOS restriction makes sense, even though it's more aggressive.
- jchw 5y agoIt’s always nice to see Dolphin news. I dunno why they’re so surprised over the JITs syncing in some games, though. I suppose a lot could go wrong, but only a few games seemed to have especially strong reliance on floating point behaviors to begin with, and I sort of expect the behavior of JITs to be influenced by the interpreter a bit due to the way things are laid out in dolphin. I tried porting a much simpler JIT to M1 and ran into the problem that Rosetta 2 was simply better at translating an AMD64 JIT than my attempt at a JIT. It could’ve been related to W^X performance, but I actually suspect the real answer is that Rosetta’s optimization passes were doing things the JIT did not do natively. I don’t know how to debug that, though, because from the debugger’s PoV, emulated processes look just like native Intel processes.
- xrd 5y agoWhere is the Linux ARM equivalent laptop? When I read about Pine laptops, it never seems like they tout the amazing performance like the M1.
- fulafel 5y agohttps://arstechnica.com/gadgets/2021/04/apple-m1-hardware-support-merged-into-linux-5-13/ https://arstechnica.com/gadgets/2021/04/apple-m1-hardware-su...
- xrd 5y agoThat's interesting, but I don't want to buy apple hardware. When I buy an Apple product, I'm paying for the integration the software and all the other stuff in addition to the hardware. That's a steep tax. I just want an arm chip performance and free software on top of it
- djrogers 5y agoThe M1 doesn't smoke Intel chips just because it's ARM - the latest chips from Broadcom and Samsung don't even come close. The M1 is good because it's good.
- xrd 5y agoThat's what I'm a little confused about. It isn't just because it is RISC, it's Apple magic? It seems weird that you can emulate other instruction sets with RISC underneath and get the performance they do. I assumed if you could recompile to the native instruction set you would get a really optimized app, but it seems like the interesting work always operates at a different layer. Fascinating stuff.
- NobodyNada 5y agoI am absolutely not an expert on microarchitecture, but I’ve had the same questions and tried my best to figure out answers. Here’s my understanding of the situation: > It isn't just because it is RISC, it's Apple magic? It’s both. We’ve known for decades that RISC was the “right” design, but x86 was so far ahead of everyone else that switching architectures was completely infeasible (even Intel themselves tried and failed with Itanium). It would have taken years to design a new CPU core that could match existing x86 designs, and breaking backwards compatibility is a non-starter in the Windows world. So we ended up with a 20-year-long status quo where ARM dominated the embedded world (due to its simplicity and efficiency) and x86 dominated the desktop world due to its market position. However, with Apple, all the stars lined up perfectly for them to be able to pull off this transition in a way that no other company was able to accomplish. - Apple sells both PCs and smartphones, and the smartphone market gave them a reason to justify spending 10 years and billions of dollars on a high-performance ARM core. The A series slowly evolved from a regular smartphone processor, into a high-end smartphone processor, and then into a desktop-class processor in a smartphone. - Apple (co-)founded ARM, giving them a huge amount of control over the architecture. IIRC they had a ton of influence on the design of AArch64 and beat ARM’s own chips to market by a year. - Intel’s troubles lately have given Apple a reason to look for an alternative source of processors. - Apple’s vertical integration of hardware and software means they can transition the entire stack at once, and they don’t have to coordinate with OEMs. - Apple does not have to worry about backwards compatibility very much compared to a Windows-based manufacturer. Apple has a history of successfully pulling off several architecture transitions, and all the software infrastructure was still in place to support another one. Mac users also tend to be less reliant on legacy or enterprise software. > It seems weird that you can emulate other instruction sets with RISC underneath and get the performance they do. As far as I understand it, the only major distinction between RISC and CISC is in the instruction decoder. CISC processors do not typically have any more advanced “hardware acceleration” or special-purpose instructions; the distinction between CISC and RISC is whether you support advanced addressing modes and prefix bytes that let you cram multiple hardware operations into a single software instruction. For instance, on x86 you can write an instruction like ‘ADD [rax + 0x1234 + 8*rbx], rcx’. In one instruction you’ve performed a multi-step address calculation with two registers, read from memory, added a third register, and written the result back to memory. Whereas on a RISC, you would have to express the individual steps as 4 or 5 separate instructions. Crucially, you don’t have to do any more actual hardware operations to execute the 4 or 5 RISC as compared to the one CISC instruction. All modern processors convert the incoming instruction stream into a RISCy microcode anyway, so the only performance difference between the two is how much work the processor has to spend decoding instructions. x86 requires a very complex decoder that is difficult to parallelize, whereas ARM uses a much more modern instruction set (AArch64 was designed in 2012) that is designed to maximize decoder throughput. So this helps us understand why Apple can emulate x86 code so efficiently: the JIT/AOT translator is essentially just running the expensive x86 decode stage ahead of time and converting it to a RISC instruction stream that is easier for a processor to digest. You’re right, though, that native code can always be more tightly optimized since the compiler knows much more about the program than the JIT does and can produce code bettor tailored to the quirks and features of the target processor.
- georgestephanis 5y agoBest excerpt of the post: > We really didn't expect this to work or we probably would have tried it sooner.
- NaN1352 5y agoAs an aside I’m thinking voxel based games, and generally games that render via CPU should do really well with a native M1 port, right? (with scaling, because the 4.5k resolution gotta hurt :))
- Rhedox 5y agoWhat do you mean by voxel based games? And why would software rendering be fast on the M1?
- gjsman-1000 5y agoFor anyone shocked that ARM chips could get this far, and for the Android users in the audience, remember that Apple cofounded ARM.
- mengineer 5y agoI think android users are more aware that GPUs are where you'd spend money. No one was competing for fastest single thread because no one needs it. Well maybe marketers need it.
- girvo 5y ago> because no one needs it. Engineers making pronouncements like this bring me no end of amusement. Surely we've learned by now that we're pretty terrible at these sorts of guesses?
- spideymans 5y ago> No one was competing for fastest single thread because no one needs it. A big part of the reasons why web applications are so darn fast on the M1 is it’s single-threaded performance. Remember that JavaScript itself is single-threaded
- astrange 5y ago> No one was competing for fastest single thread because no one needs it. https://en.wikipedia.org/wiki/Amdahl%27s_law https://en.wikipedia.org/wiki/Amdahl%27s_law
- etaioinshrdlu 5y agoI wonder if there are any gains to be had on the M1 because it uses shared memory between the CPU and GPU - much like the actual Gamecube architecture. From reading this blog, Gamecube games often made heavy usage of the memory-sharing capability of the hardware - which made emulation on PCs a performance challenge.
- astrange 5y agoThis certainly isn’t the first shared memory GPU that Dolphin runs on, although it might be the fastest and it’s definitely the fastest TBDR one. A problem that I don’t think applies to consoles is that GPUs don’t use the same texture format CPUs do - they swizzle them in proprietary ways and it needs conversion even if there are no memory transfers.
- Rhedox 5y agoYou most likely still need to do all the work to manage texture caches because first of all, you need to create texture objects in the host graphics api based on what's essentially just memory. On top of that, GC/Wii texture formats might be different from what the host can support. From my understanding it's not that useful for reading back either as the main bottleneck there is the fact that you need to sync gpu and cpu rather than transfer speed.
- etaioinshrdlu 5y agoNow you've made me wonder why Dolphin can't store textures and all GPU data in the same form in memory as the Gamecube -- and do any translation on-the-fly in shaders.
- deleted 5y ago[deleted]