4 ms·
This certainly rings true with my own experiences (worked on both GPUs and CPUs and now doing AI inference at UK startup Fractile). > If there were a hypotheti
by gchadwick 1y ago
This certainly rings true with my own experiences (worked on both GPUs and CPUs and now doing AI inference at UK startup Fractile).
> If there were a hypothetical high-level CPU language that somehow encoded all the information the microarchitecture needs to measure to manage memory access,
I think this is fundamentally impossible because of dynamic behaviours. It's very tempting to assume that if you can be clever enough you can work this all out ahead of time, encode what you need in some static information and the whole computation just runs like clockwork on the hardware, no or little area spent on scheduling, stalling, buffering etc. Though I think history has shown over and over this just doesn't work (for more general computation at least, more feasible in restricted domains). There's always lots of fiddly details in real workloads that surprise you and if you've got an inflexible system you've got no give and you end up 'stopping the world' (or some significant part of it) to deal with it killing your performance.
Notably running transformer models feels like one of those restrictive domains you could do this well in but dig in and there's plenty enough dynamic behaviour in there that you can't escape the problems they cause.
- throwaway31131 1y ago> I think this is fundamentally impossible because of dynamic behaviours. I was thinking that too as it sounds awfully close to a variant of the halting problem. But since I am not sure that it's actually equivalent, and there might be some heuristic that constrains the problem, in theory, to make it solvable "enough" I went with the softer "I don't know how to do it" rather than "it can't be done"