3 ms·
TLDR: let's put compute inside the memory. I'm very leaning towards calling this bullshit. Closer to memory compute is what we have been doing for last 10 year
by extropy 7y ago
TLDR: let's put compute inside the memory.
I'm very leaning towards calling this bullshit.
Closer to memory compute is what we have been doing for last 10 years, and that's why we have ever increasing hierarchical caches.
Yes the ALUs are a tiny part of the power budget because moving / syncing data is the hard problem.
If you want massive parralelism, use a GPU.
The in memory clone and zero could make some sense, but generally you want to start writing to that memory soon after that, meaning you still need to pull it into your cache and the benefit is negated.
- marcosdumay 7y ago> If you want massive parallelism, use a GPU. The GPU doesn't have all of your data stored in it. The clone and zero part have a very weak reasoning, focusing on bad targets because they look better than important ones, but take a look at the next section, where they do a simple database search.
- imtringued 7y agoPIM is already a reality it's just a matter of time until it sees broad adoption. PIM does not suffer from the crippling limitations of GPUs which only perform well with arithmetic bound problems, cannot run conventional non SIMD code efficiently and need a relatively large batch size. For instance try deserializing a million JSON strings with a GPU. The end result is a graph like memory structure which GPUs usually struggle with. GPUs cannot take advantage of sharing the instruction stream here because it will hit branches for almost every single character and diverge quickly. And finally if you only want parse one JSON object then the GPU is worthless. A PIM based solution would not struggle with heterogeneous non-batched workloads with arbitrary layout at all but still offer the same performance advantages that GPUs enjoy compared to CPUs and reduce energy usage at the same time.