23 ms·
> make copies for those calculations but this seems inefficient to me I haven't read the whole article, but this "make copies of elements from an array into an
by jonv98 7y ago
> make copies for those calculations but this seems inefficient to me
I haven't read the whole article, but this "make copies of elements from an array into another array for the current frame only" is common in game development.
Remember that on modern CPUs, an L3 miss is about 200x slower than an L1 hit. RAM isn't random access: randomly jumping around is slow, but iterating over an array is fast, both because of the cache and because of pre-fetching.
Say you have a big array of A's, and another big array of B's. For the current frame, some of the A's need to interact with some of the B's. If you go through the entire list of B's, and copy the ones that will definitely need to interact into a new list, call it B2, then maybe (or not) do the same with the A's into A2, then it can often be approximately 30 times faster. Multiply that by 4 (or 8) if you can "zip" through your A2's and B2's with SIMD.
Not only that, but your A2 and B2 lists can be put on a stack allocator (nothing to do with allocating on the stack - it's a special type of O(1) heap allocator whose contents are discarded at the end of each video frame).
- std_throwaway 7y agoCopying is fast if the access pattern is a good fit for the CPU architecture. If you need to copy only every N-th byte from a AoS it might be as inefficient as random access. So, copying could be expensive. The article suggests striping your data in blocks but then you end up with the worst of both worlds in terms of program code complexity.