4 ms·
Just curious what your issues with Triton were. We've done OK with it using it to serve LLM models w/ a classifier head via HF Transformers pipeline & Flash Att
by eoskx 2y ago
Just curious what your issues with Triton were. We've done OK with it using it to serve LLM models w/ a classifier head via HF Transformers pipeline & Flash Attention 2, as well as serving text generation models with the vLLM back-end.
- bytesandbits 2y agotriton is not that bad, TensorRT will give you nightmares
- dlewis1788 2y ago100% - probably why vLLM is now the default back-end in Dynamo.