4 ms·
As someone who spent the better part of a year trying to get various Nvidia inference products to work _at all_ even with a direct line to their developers, I w
by Carrok 2y ago
As someone who spent the better part of a year trying to get various Nvidia inference products to work _at all_ even with a direct line to their developers, I will simply say "beware".
- vinni2 2y agoCan you share some of your wisdom on setting up a scalable inference infrastructure?
- Carrok 2y agoUse Ray Serve. https://docs.ray.io/en/latest/serve/index.html https://docs.ray.io/en/latest/serve/index.html
- ipsum2 2y agoAs someone who has run LLMs in production, using Ray is probably the worst idea. It's not optimized for language models, and is extremely slow. There's no KV-caching, model parallelism, and other basic table stakes features that are offered by Dynamo or other open source inference frameworks. Useful only if you have <1 QPS. Use SGLang, vLLM, or text-generation-inference instead.
- Carrok 2y agoThis is probably true, but unlike every Nvidia product we tried, it did, you know, reply to inference requests with actual output. That said, you can serve vLLM with Ray Serve. https://docs.ray.io/en/latest/serve/tutorials/vllm-example.html https://docs.ray.io/en/latest/serve/tutorials/vllm-example.h...
- erulabs 2y agoIt really depends on the task. If you have 1 massive job, Ray sucks and doesn't provide table stakes. If you have 50M tiny jobs, Ray and kuberay is great and serves as the backbone of several billion dollar products. Good for the goose, good for the gander...
- richardliaw 2y ago> If you have 1 massive job, Ray sucks and doesn't provide table stakes. Can you say more?
- islewis 2y agois this in reference to Triton?
- Carrok 2y agoAnd NIM, yes.
- aabhay 2y agoTriton is not that bad at all, considering the wide scope of systems it has to support (tensorrt, onnx, multiple generations of pytorch, cuda, python). It was much nicer than the old Torchserve project which was JVM based.
- eoskx 2y agoJust curious what your issues with Triton were. We've done OK with it using it to serve LLM models w/ a classifier head via HF Transformers pipeline & Flash Attention 2, as well as serving text generation models with the vLLM back-end.
- bytesandbits 2y agotriton is not that bad, TensorRT will give you nightmares
- dlewis1788 2y ago100% - probably why vLLM is now the default back-end in Dynamo.
- raffraffraff 2y agoI've done very little with Nvidia software, but what I have done puts me off ever doing it again. I quit a job partially because it involved trying to get their shit to work. (There were other factors, but that was definitely on the 'GTFO' side)