3 ms·
You could use Microsoft's DeepSpeed to run the model for inference on multiple GPUS, see https://www.deepspeed.ai/tutorials/inference-tutorial/ https://www.deep
by bm-rf 5y ago
You could use Microsoft's DeepSpeed to run the model for inference on multiple GPUS, see https://www.deepspeed.ai/tutorials/inference-tutorial/ https://www.deepspeed.ai/tutorials/inference-tutorial/