2 ms·Mapping GPUs to LLMs (and back): A bandwidth-based estimator for local inference2 points by apignotti 6mo ago