4 ms·
We are experimenting with workers using the simple python arq library and Redis and I am yet to find a MLOps or Data Engineering use case that is not a good fit
by kfk 4y ago
We are experimenting with workers using the simple python arq library and Redis and I am yet to find a MLOps or Data Engineering use case that is not a good fit for a API+Worker on K8S. For instance, you need to manage ML artifacts? You can just offer an API endpoint so the ML models can automatically update the artifacts. You need data ingestion? You can have a worker running ingestion scripts and kick off the worker via API. We tried pub/sub and Kafka but it can be really wasteful, workers can process work for multiple streams, but Kafka cannot. But of course I wonder if I am missing something, I am not an ML engineer so probably I am?
- alextheparrot 4y agoIt isn’t particularly clear what technical requirements you were working against, but let me give it a shot: A lot of groups start using Kafka because they have high-throughput event streams they want to aggregate over, and then you just use Kafka for everything because managing a 5 TPS topic alongside a 100k TPS topic is trivial. In terms of why Kafka is a good fit for that workflow — database writes are unnecessary and oddly structured for raw events we want to expire in a few days, and dealing with buffering blob file writes can cause data loss, so Kafka can really simplify the producer architecture where the producer is also a consumer from a producer who wants an ack. Combine this with how trivial it is to have multiple readers on a pub sub system, and it is easy to scale from 1 to N consumers of a data source without duplicating the data everywhere. E.g. you could have three aggregation jobs that use the same data, one job that writes the topic’s data to blob-style storage for batch use-cases, and a low latency inference job all running from the same data stream. More or less, Kafka just simplifies scale-out in some cases, maybe not your case, though. If you’re kicking off workers to do ingestion, you might have a system where you are pulling files down at some infrequent cadence (Let’s say every 10 minutes) —- in that case Kafka is likely going to be overkill and feel like a lot of work for an API call that then becomes a tasked tracked in some database.
- kfk 4y agointeresting, in my case I need the events to be inserted in a SQL db, even the real time ones. For instance, I receive Contacts data from Hubspot in realtime, I send those contacts over to Salesforce and I store them in Postgres. Why Postgres? Because we want to keep a history of the contacts, plus we will need to have 1 source of truth for customers data to fulfill various data privacy requirements. How about Kafka for such a use case? Let's say I receive maybe 10,000 contacts per day.
- geoduck14 4y agoNot your OP, but someone who has been thinking about this. Use Kafka as the super fast layer to connect producers and consumers. Have it write to S3/Blob. Then insert into your dB layer for "cold" access. Inserting into a dB takes longer than pub/sub, and you don't want to mess up the publishing of data.