3 ms·
Is Yelp still using pyleus and Apache Storm? Or have they migrated to Spark and Kafka Streams?
by pixelmonkey 10y ago
Is Yelp still using pyleus and Apache Storm? Or have they migrated to Spark and Kafka Streams?
- wtetzner 10y agoWell, in the article they claim they're using Kafka Streams, and there's a mention of Spark in one of the diagrams. I don't see any mention of pyleus or Apache Storm.
- justinc-yelp 10y agoMost of the stream processing in the Data Pipeline happens inside of an internal project called PaaStorm, which is storm-like. It was built to take advantage of our platform as a service (http://engineeringblog.yelp.com/2015/11/introducing-paasta-an-open-platform-as-a-service.html http://engineeringblog.yelp.com/2015/11/introducing-paasta-a...), which handles process scheduling really well. Architecturally, it's pretty similar to Samza, with distributed processes communicating using Kafka. We do use Spark streaming, and are starting to use Kafka Streams and Data Flow, where they're a better fit. I'm personally most excited about Beam/Flink. We'll probably end up replacing the PaaStorm internals with some other tool, when one with good python support matures. Beam's event-time handling and windowing seem really promising at this point. https://www.oreilly.com/ideas/the-world-beyond-batch-streaming-102 https://www.oreilly.com/ideas/the-world-beyond-batch-streami... is a great overview of the different concerns for stream processing.
- ricardobeat 10y agoHi Justin! Thanks for sharing, very interesting stuff. How do you scale Kafka to handle the massive amount of traffic (and storage) that you seem to generate daily? With services talking among themselves via HTTP there is a lot of resilience built-in. Do you have anything in place to avoid this becoming a single point of failure? It must have become the most critical piece of your infra.
- poooogles 10y agoScaling Kafka is pretty simple, the operations document contains most things you'll need to get started [1]. We push 500k documents a second through over 10 6 core/24gb ram hosts pretty uneventfully. Only real pointer is to size ZK appropriately and make sure you leave lots of memory for the file system cache. 1 - https://kafka.apache.org/documentation.html#operations https://kafka.apache.org/documentation.html#operations
- justinc-yelp 10y agoI'm not actually a good resource on scaling Kafka. Our distributed systems teams do a great job of providing reliable infrastructure and scaling it up, so on the application side we are mostly able to treat it like a black box that just works. In general, I do think poooogles covered it well. Kafka is designed to scale. The one thing we do that you might not expect is splitting data across clusters, depending on what guarantees we want to provide. We also tend to make sure all data is replicated using geographic distribution to avoid SPOF issues. We do use the min ISR settings and different required ack levels, depending on we want to trade off durability and availability for an application.