2 ms·
Alpa: Auto-parallelizing large model training and inference (by UC Berkeley)
- zhisbug 4y agoServe OPT-175B with Alpa using commodity GPUs: https://alpa-projects.github.io/tutorials/opt_serving.html https://alpa-projects.github.io/tutorials/opt_serving.html If you have difficulties accessing 8x 80GB A100 (or AWS P4) this might be a good use case.