3 ms·
drm likes uncacheable write-combining as an optimization… but that's disabled on arm64 https://patchwork.kernel.org/project/linux-arm-kernel/patch/2019012412065
by floatboth 5y ago
drm likes uncacheable write-combining as an optimization… but that's disabled on arm64 https://patchwork.kernel.org/project/linux-arm-kernel/patch/20190124120658.30288-1-ard.biesheuvel@linaro.org/ https://patchwork.kernel.org/project/linux-arm-kernel/patch/... because it can cause image corruption glitches (saw them myself, had to find and cherry-pick that commit into FreeBSD's drm back then)
Generally I don't think "normal" is necessary? In FreeBSD/aarch64 we interpret most ioremaps (all other than WC and WB) as "device": https://reviews.freebsd.org/D20789 https://reviews.freebsd.org/D20789
and there doesn't seem to be a performance problem. Well, I haven't scientifically tested the performance but SuperTuxKart can do >100fps at 4K on an RX 480 :)
- marcan_42 5y agoWith Device memory you can't do unaligned accesses, so userspace apps that map GPU memory and expect that to work (as it does on x86 and on ARM if you can do a Normal mapping) will break.
- floatboth 5y agoLooking at amdgpu_ttm, for '"On-card" video ram' it sets TTM_PL_FLAG_WC, so that would use ioremap_wc, that's not device, we have it as WRITE_COMBINING which is actually WRITE_THROUGH. So yeah normal mappings might be used somewhere. Where does that limitation on the M1 come from anyway?
- marcan_42 5y agoThe M1 bus fabric is very picky about access modes. We had to add a whole new ioremap_np() to the kernel, because it requires nGnRnE mappings for on-die peripherals, while Linux used nGnRE for everything on ARM64 until now. For PCIe BARs it wants nGnRE instead, and I'll be very surprised if it'll take a normal mapping...