3 ms·
Dislcaimer: I am not fully on bored with k8s and enable clusters that are not kubernetes as well as k8s deployments. Is pachyderm actually used anywhere though
by agibsonccc 9y ago
Dislcaimer: I am not fully on bored with k8s and enable clusters that are not kubernetes as well as k8s deployments.
Is pachyderm actually used anywhere though? Especially compared to the hadoop distros?
Why would I use this over just spinning something up on EMR? If I were on prem, I likely already have hadoop installed. K8s is typically a separate cluster (mainly because it's still a bit error prone for anything with data applications, which is I'm assuming what you're claiming to solve?
Beyond that, re: kubeflow. The problem with just assembling some of these things and calling it "production" is you're still missing a lot of the basic things when people need to go to production:
1. Experiment tracking for collaboration
2. Managed Authentication/connectors for anything outside K8s
3. Connecting with/integrating with existing databases
4. Proper role management: Data Scientists aren't deploying things to prod (at least customer facing prod at scale where money is on the table, "prod" could also mean internal tools for experiments), they typically need to be integrating with external processes and different teams.
Many of these things are left as an exercise for the reader (especially in k8s land). Granted, tutorials and the hosted environments exist, but nothing is a "1 click" deploy that is being promised on any of these "See how easy it is!" blog posts that run something on 1 laptop.
A lot of things being promised here just don't line up for me yet. K8s is maturing quite a bit but when the story is still "Run managed k8s on the cloud or spend time upgrading your cluster every quarter" - I have to say it's far from close to anything a typical data scientist running sk learn on their mac are going to be able to get started with.
The closest I've seen to that that isn't our own product is AWS Sagemaker which actually solves real world problems by actually gluing components together in a fairly seamless way.
Let's just be clear here that we still have a long way to go yet. In practice, we have separate teams that need to collaborate. Data scientists aren't going to download minikube tomorrow and go to production without someone else's approval and going through a huge learning curve yet.
We're moving in the right direction though!