3 ms·
Use Ray Serve. https://docs.ray.io/en/latest/serve/index.html https://docs.ray.io/en/latest/serve/index.html
by Carrok 2y ago
Use Ray Serve. https://docs.ray.io/en/latest/serve/index.html https://docs.ray.io/en/latest/serve/index.html
- ipsum2 2y agoAs someone who has run LLMs in production, using Ray is probably the worst idea. It's not optimized for language models, and is extremely slow. There's no KV-caching, model parallelism, and other basic table stakes features that are offered by Dynamo or other open source inference frameworks. Useful only if you have <1 QPS. Use SGLang, vLLM, or text-generation-inference instead.
- Carrok 2y agoThis is probably true, but unlike every Nvidia product we tried, it did, you know, reply to inference requests with actual output. That said, you can serve vLLM with Ray Serve. https://docs.ray.io/en/latest/serve/tutorials/vllm-example.html https://docs.ray.io/en/latest/serve/tutorials/vllm-example.h...
- erulabs 2y agoIt really depends on the task. If you have 1 massive job, Ray sucks and doesn't provide table stakes. If you have 50M tiny jobs, Ray and kuberay is great and serves as the backbone of several billion dollar products. Good for the goose, good for the gander...
- richardliaw 2y ago> If you have 1 massive job, Ray sucks and doesn't provide table stakes. Can you say more?