4 ms·TPI-LLM: Serving 70B-Scale LLMs Efficiently on Low-Resource Edge Devices2 points by CrypticShift 2y ago