3 ms·
Hi all, I'm one of the authors of Ray, thanks for all the comments and discussion! To add to the discussion, I'll mention a few conceptual things that have chan
by robertnishihara 5y ago
Hi all, I'm one of the authors of Ray, thanks for all the comments and discussion! To add to the discussion, I'll mention a few conceptual things that have changed since we wrote the paper.
*Emphasis on the library ecosystem*
A lot of our focus is on building an ecosystem of libraries on top of Ray (much, but not all, of the focus is on machine learning libraries).
Some of these libraries are built natively on top of Ray such as Ray Tune for scaling hyperparameter search (http://tune.io http://tune.io), RLlib for scaling reinforcement learning (http://rllib.io http://rllib.io), Ray Serve for scaling model serving (http://rayserve.org/ http://rayserve.org/), and RaySGD for scaling training (https://docs.ray.io/en/master/raysgd/raysgd.html https://docs.ray.io/en/master/raysgd/raysgd.html).
Some of the libraries are popular libraries on their own, which now integrate with Ray such as Horovod (https://eng.uber.com/horovod-ray/ https://eng.uber.com/horovod-ray/), XGBoost (https://xgboost.readthedocs.io/en/latest/tutorials/ray.html https://xgboost.readthedocs.io/en/latest/tutorials/ray.html), and Dask for dataframes (https://docs.ray.io/en/master/dask-on-ray.html https://docs.ray.io/en/master/dask-on-ray.html). While Dask itself has similarities to Ray (especially the task part of the Ray API), Dask also has libraries for scaling dataframes and arrays, which can be used as part of the Ray ecosystem (more details at https://www.anyscale.com/blog/analyzing-memory-management-and-performance-in-dask-on-ray https://www.anyscale.com/blog/analyzing-memory-management-an...).
Many Ray users start using Ray for one of the libraries (e.g., to scale training or hyperparameter search) as opposed to just for the core system.
*Emphasis on serverless*
Our goal with Ray is to make distributed computing as easy as possible. To do that, we think the serverless direction, which allows people to just focus on their code and not on infrastructure, is very important. Here, I don't mean serverless purely in the sense of functions as a service, but something that would allow people to run a wide variety of applications (training, data processing, inference, etc) elastically in the cloud without configuring or thinking about infrastructure. There's a lot of ongoing work here (e.g., to improve autoscaling up and down with heterogeneous resource types). More details on the topic https://www.anyscale.com/blog/the-ideal-foundation-for-a-general-purpose-serverless-platform https://www.anyscale.com/blog/the-ideal-foundation-for-a-gen....
If you're interested in this kind of stuff, consider joining us at Anyscale https://jobs.lever.co/anyscale https://jobs.lever.co/anyscale.
- sillysaurusx 5y ago> Our goal with Ray is to make distributed computing as easy as possible. To do that, we think the serverless direction, which allows people to just focus on their code and not on infrastructure, is very important. I watched https://now.sh/ https://now.sh/ deteriorate from a simple, lovely CLI into a dystopian mess due to their push for serverless. They abandoned all other approaches and forced people to use it. Far from making things easier, it became a kafkaesque pipeline of dependencies and configuration settings just to get any small example deployed. Things may be better now, but the experience was so jarring and offputting that I haven't used now.sh for much of anything. Used to use it for everything; somehow https://docs.ycombinator.lol/ https://docs.ycombinator.lol/ is still running, which was a static site deployed back before their serverless stuff. I don't know. You might be right. But just remember, the Ray library -- the actual python lib -- is your bread and butter. It's why everyone loves you. I urge you, never make the mistake of letting it deteriorate. It should be rock solid for everyone forever, with no need to interface with any of your serverless components. The day you try to monetize by trying to sneak in "value adds" by making the code "easy to integrate with your serverless stuff" is the day that you open yourself to bugs, and the temptation to ignore problems in other areas -- because after all, the serverless infra would be where you're making your money, so it makes sense to push everyone in that direction. Ray is so excellent right now that it feels like a sports car. I hope it'll stay excellent for a decade to come. (All I want is the ability to recover from client failures in a way where, if there are tasks in flight, I can tell those tasks how to re-run once all the actors have reconnected. I'm sure there's already a way to do something like this; just haven't looked into the details quite yet.) Best of luck, and thanks for the wonderful lib. EDIT: https://www.anyscale.com/blog/the-ideal-foundation-for-a-general-purpose-serverless-platform https://www.anyscale.com/blog/the-ideal-foundation-for-a-gen... just gives me terrible feelings. Your best bet is to ignore me, because my gut is likely wrong here -- at 33, I'm starting to fall past the hill. But for example: > Ray hides servers Suppose a hacker wants to build an iOS app powered by Ray. They want to create a cluster of servers to process incoming tasks. The tasks are things like "Make memes with AI," artbreeder-style, and then send them back to the iOS client waiting for them. Then the user can enjoy their meme, you throw up a "if you like memes, give me a dollar and you can have all the memes you want," and a million people download your AI meme app and you become the Zuckerbezos of AI. In that context, no one wants to hide servers. No one I know -- anywhere -- thinks it's a good idea. We don't want to rely on your magic solutions. We want to keep our servers running. Because my servers happen to be TPUs, and there's no way that TPUs are ever going to become serverless. But even before I was using TPU VMs, all I wanted to do was to just stick GPUs onto servers and send results around; the serverless stuff gave me the creeps. Perhaps this just means I was ineffective, though, and that everyone I know is also ineffective. I know, I know... you're going to support your non-serverless offerings, and Ray will be wonderful forever, and it'll be roses and rainbows. I hope so. But just don't let the core library become priority #2. It should be priority #1 forever.