2 ms·
> Except in order memory access is what modern CPUs aim for > Strided access has always been slow, on CPUs, on GPUs, on everything Yes and no. It's not as if
by Sirened 4y ago
> Except in order memory access is what modern CPUs aim for
> Strided access has always been slow, on CPUs, on GPUs, on everything
Yes and no. It's not as if hardware engineers completely stuck our heads in the sand and just let the performance be shitty, they've actually developed very sophisticated prefetchers (in this case, literally called a "stride prefetcher"). Many modern stride prefetchers can detect and launch concurrent lines faults at strides of over a megabyte. Here, we're explicitly exploiting that line faults _do not_ need to happen in order thanks to memory ordering rules provided by the ISA.
While you will still get crumby cache utilization due to hammering a single set, if you are legitimately just going once through a loop and operating on single items (i.e. you're streaming and not returning to items once accessed like in the article), the prefetcher will make a huge difference and get you a decent hit rate.