4 ms·
> 3) Some stricter memory model in hardware? Seems like that'd go against most of the stated reason for switching to ARM in the first place. I would assume tha
by andoma 6y ago
> 3) Some stricter memory model in hardware? Seems like that'd go against most of the stated reason for switching to ARM in the first place.
I would assume that a more strict memory model would be enabled only for processes that needs it (ie, Rosetta translated ones). So a cpu-flag is set/cleared when entering/exiting user mode for those processes. Does this require a separate/special cache coherency protocol? A complete L1d flush when entering/leaving these processes (across all CPUs)?
Not and expert in this field and it feels complicated for sure. Is it worth it for just emulating "legacy" applications during a transitional period? Perhaps Apple can pull it off though.
- monocasa 6y ago> I would assume that a more strict memory model would be enabled only for processes that needs it (ie, Rosetta translated ones). So a cpu-flag is set/cleared when entering/exiting user mode for those processes. Does this require a separate/special cache coherency protocol? The benefits you'd get from going to a weaker memory model are by not having that extra coherency in the critical path in the first place. Adding extra muxes in front of it to make it optional would be worse than just having it on all the time. > A complete L1d flush when entering/leaving these processes (across all CPUs)? That wouldn't help because two threads could be running at the same time on different cores against their respective L1 and store buffers.
- andoma 6y ago> The benefits you'd get from going to a weaker memory model are by not having that extra coherency in the critical path in the first place. Adding extra muxes in front of it to make it optional would be worse than just having it on all the time. Indeed true, good point. > That wouldn't help because two threads could be running at the same time on different cores against their respective L1 and store buffers. Of course, this was related to the cost of switching coherency protocol during context switch. But as you say the overhead of just making it switchable is prohibitive in itself.
- phire 6y agoIt can be enabled per instruction. Atomic instructions (and ARMv8.1 added a bunch of new atomic read-modify-write instructions that line up nicely with x86) use the new cache coherency protocol, while the older non-atomic instructions keep the relaxed memory model. Though, I'm not sure if it's worth it to keep two concurrency protocols around. I wouldn't be surprised if the non-atomic instructions get an undocumented improvement to their memory model.