3 ms·
The bottleneck for single batch inference is memory bandwidth. The M4 Pro has less memory bandwidth than the P40, so it would be slower. Also, the setup present
by oofbaroomf 2y ago
The bottleneck for single batch inference is memory bandwidth. The M4 Pro has less memory bandwidth than the P40, so it would be slower. Also, the setup presented in the OP has system RAM, allowing you to run models than what fits in 48GB of VRAM (and with good speeds too if you offload with something like ktransformers).
- anthonyskipper 2y ago>>M4 Pro has less memory bandwidth than the P40, so it would be slower Why do you say this? I thought the p40 only had a memory bandwidth of 346 Gbytes/sec. The m4 is 546 GB/s. So the macbook should kick the crap out of the p40.
- oofbaroomf 2y agoThe M4 Max has up to 546 GB/s. The M4 Pro, what GP was talking about, has only 273 GB/s. An M4 Max with that much RAM would most likely exceed OP's budget.