4 ms·
I thought Ray was a reinforcement learning platform? Can you elaborate on how it is a replacement for Spark?
by anonymousDan 2y ago
I thought Ray was a reinforcement learning platform? Can you elaborate on how it is a replacement for Spark?
- m_ke 2y agoRay is a distributed computing framework that has a fast scheduler, 0 copy / serialization between tasks (shared memory) and stateful "actors", making it great for RL but it's more general than that. I'd recommend checking out their architecture whitepaper: https://docs.google.com/document/d/1tBw9A4j62ruI5omIJbMxly-la5w4q_TjyJgJL_jN2fI/preview?tab=t.0#heading=h.iyrm5j2gcdoq https://docs.google.com/document/d/1tBw9A4j62ruI5omIJbMxly-l... Imagine spark without the JVM baggage and with no need to spill to disk / serialize between steps unless it's necessary.
- anonymousDan 2y agoThanks. Can you clarify what you mean by your last point? One of the main advantages of spark over Hadoop is that it doesn't need to spill to disk/ serialise between steps, i.e. it is pitched as an in memory big data platform, so I'm a bit confused.
- NeutralCrane 2y agoRay has some really unfortunate nomenclature/branding in my opinion. Ray is, itself, a framework for distributed computing, on top of which they have built numerous applied platforms such as Ray Data, Ray Train, Ray Serve, etc. This includes your reinforcement learning platform, RayRLib. I think they’ve diluted the “brand” a bit with this approach and would be better off sticking with “Ray” for the distributed computing and spinning up the others as something completely separate, but that’s just me.