4 ms·
I think if you could somehow start computing the resultant matrix elements as soon as you read a row/column from the input ones, you could reach their "physical
by OneDayIGoHome 6y ago
I think if you could somehow start computing the resultant matrix elements as soon as you read a row/column from the input ones, you could reach their "physically possible" limit.
A couch expert on computer architecture here, but a large enough systolic array could be used to achieve their "physically possible" limit? [0]
New CUDA GPUs have been coming with these tensor cores that are just systolic arrays. Google's TPU are the same.
Could someone with more knowledge on systolic arrays comment on whether a large systolic array can achieve this?
[0] https://medium.com/intuitionmachine/googles-ai-processor-is-inspired-by-the-heart-d0f01b72defe https://medium.com/intuitionmachine/googles-ai-processor-is-...
- btilly 6y agoThe answer is no but sort of. You cannot reduce the number of operations by waving a magic architecture want. Hence the no. But you can do operations in parallel rather than in series. Same number of operations, but it takes less time. So if a systolic array is built for the same scale as your problem, your bottleneck does indeed become how quickly you can send data to/from that dedicated system.