4 ms·
Clearly memory latency is the bottleneck. I see 2 paths we can take. 1. Memory orientated computers 2. Less memory abstractions We already did the first, the
by pyrolistical 1y ago
Clearly memory latency is the bottleneck. I see 2 paths we can take.
1. Memory orientated computers
2. Less memory abstractions
We already did the first, they are called GPUs.
What I imagine for the second is full control over the memory hierarchy. L1, L2, L3, etc are fully controlled by the compiler. The compiler (which might be just in time) can choose the latency vs throughput. It knows exactly how all the caches are connected to all numa nodes and their latencies.
The compiler can choose to skip a cache level for latency critical operations. The compiler knows exactly the block size and cache capacity. This way it can shuffle data around with pipelining to hide latency, instead of accidentally working right now.