3 ms·
We are getting a forward pass time of ~100ms on Meta's original Llama2 70B (float16, batch size 8) PyTorch implementation on 8xA100. Those results are very unde
by agr_nyc 2y ago
We are getting a forward pass time of ~100ms on Meta's original Llama2 70B (float16, batch size 8) PyTorch implementation on 8xA100. Those results are very underwhelming in terms of fully utilizing the GPU flops. If we are doing something wrong, let me know.
The vllm implementation is much faster, I think 50ms or better on either 4 or 8 A100s, forget the exact number.
- yungtriggz 2y agoyea reckon vllm is the way to go. cheers