5 ms·
I think the conclusions are a bit far fetched: "A generic Snapdragon ARM SoC, say, would deliver notably less performance in this specific scenario that is crit
by rockdoe 4y ago
I think the conclusions are a bit far fetched: "A generic Snapdragon ARM SoC, say, would deliver notably less performance in this specific scenario that is critically important to Mac users."
The usage of these flags is NOT common.
And as observed above, if they're not used, then you can "have an “unused flags” optimisation that avoids the computation a lot of the time".
So I don't see how it follows that this is a "critically important" scenario. Reality is that adding the logic to compute these flags is almost free, so it's a cheap (in terms of hardware) optimization to do to squeeze some marginal extra performance. But to say it's a deal breaker? If it is, I don't think you can conclude that based on the evidence presented.
- klelatti 4y agoBroadly agree but how easy is it for Rosetta to prove to itself that these flags aren't used somewhere in the code path?
- delusional 4y agoI'd think it would be fairly easy considering it transpiles the entire binary in a single shot. I don't see a lot of cases where you'd branch to read either of these flags. I suppose it would be much harder for JITs and other dynamically generated code.
- klelatti 4y agoFair comment. If these flags are hardly used then I guess it is a bit surprising that this is necessary.
- toast0 4y agoThe flags are hardly read, but computed often. Computing them while doing operations is nearly free in hardware, so it makes sense to add them in hardware if you can. It's not nearly free in software, but it's important to do it where not doing it might be obervable in normal flow (tricky things with interrupts are out of luck, even if you always do the software calculation, you could interrupt in the middle of it). In many cases, it's easy to determine the flags aren't observable and you can skip software computation.
- delusional 4y agoI have no real information on the topic, but to me it looks like a "why not" optimization. It was probably pretty cheap to put into the hardware (seeing as it was already an ARM extension), the engineers probably figured there was some risk that it would be really important for some workload, and once they had included it they might as well use it. In other words, it's probably not necessary but was included early on out of an abundance of caution. They knew that if it turned out to be used in some binary somewhere, it would be a major performance killer. I'd be interested in seeing a benchmark with the hardware flag turned off and the translation/optimization setup used for linux enabled. I bet the difference would be negligible.
- TeMPOraL 4y ago> I suppose it would be much harder for JITs and other dynamically generated code. Which is... just about anything these days, I suppose? Half of the desktop apps are Electron, half of new CLI apps are in Node.JS, half of old CLI apps, including near-ubiquitous ones like git, are a random assortment of half a dozen scripting languages... I haven't actually counted it properly, but ad-hoc random sampling gives me an impression that a third of typical Linux distro userspace is in Python, and most of it not even compiled AOT.
- piperswe 4y agoIn the case of Python (specifically CPython, I'm ignoring PyPy), there is no dynamically generated machine code - CPython interprets without JIT. That's fairly common across most older scripting languages (Perl, Tcl, etc.). The more common instances of JITs are, like you said, Electron/Node.js programs, Java programs, and the odd Ruby program. Anecdotally, I think Rust and Go are very common languages for new CLI apps, and I definitely have significantly more Rust and Go CLI apps than Node.js ones installed at the moment.
- delusional 4y agoBut how many of those need binary translation? Common for two of those is that you just need a good V8 implementation for ARM to completely displace Rosetta 2. For git it's mostly bash which is interpreted and therefore can be one-shot transpiled, or even just compiled for arm natively since it's written in C. The overlap of "is JIT" and "doesn't have a runtime for ARM" is overwhelmingly small. That's probably mostly because runtimes and JITs are opensourced, which mean you can just recompile them for the target. Rosetta 2 is more focused on the proprietary space where they don't often develop proprietary JIT runtimes.
- int_19h 4y agoThis is about old apps that haven't been recompiled for ARM by the maker, no? If such a legacy app was built using Electron, it'll be Electron built for x64 and targeting x64 for its JIT. It doesn't really help if Electron ships an ARM runtime today, unless the OS uses it to replace the one shipped with the app - but then you may be breaking the app, if it relies on some old Electron behavior.
- flohofwoe 4y agox86 (and 8080, Z80) can push/pop the flags register, allowing the flags bit mask to be used as regular data, not just with instructions which explicitly check those flags. So proving that specific flags are not used by the code might actually not be that easy.
- sroussey 4y agoAlso, Rosetta surprisingly does not do much for optimization, opting for correctness and leaving speed to the silicon.