3 ms·Loading Llama-2 70B 20x faster with Anyscale Endpoints1 points by fgfm 3y agofgfm 3y agoThe Anyscale team shared how you can achieve considerable speedups for model loading in production with examples on the Llama 2 variants.