3 ms·
They (Databricks) are not advertising themselves as a best practices way of doing anything, they're just a platform for doing data science and analytics. It's u
by bertomart 6y ago
They (Databricks) are not advertising themselves as a best practices way of doing anything, they're just a platform for doing data science and analytics. It's up to the data scientist to use the tool properly and to have proper methodologies. A platform is simply a set of tools that helps a target audience, and that's what databricks is. Now the issue here is that a lot of datascience either cannot code/do engineering or just don't want to. But in my view people like that are just glorified analysts. Data Science is built on the foundation of software engineering (+ stats, viz, math, etc...); this is why it's so complicated. If you cannot code or don't understand the SDLC or best practices you're basically a carpenter who can't hammer or saw.
- mlthoughts2018 6y agoThe question is not whether Databricks claims to be anything. The question is just what is a best practice, from an ETL / DevOps point of view, and how to enforce ML tooling to adhere to that from first principles. I’ll give a concrete example. In my org we use Dataproc on GCP as a model training task execution paradigm. You define your base environment via some Docker container, put it in GCR, and then define Dataproc jobs in terms of the base environment, the backing compute resources, any GCS bucket connections, and any ML-specific config like hyperparameters. A human being never under any circumstances triggers these jobs. Instead a human user deploys the config as a cronjob or regular job in Kubernetes, and then a scheduler picks them up and runs them. For experimental workloads only, developers can manually trigger a Kubernetes job. Each job consults config, spins up the appropriate Dataproc cluster, runs the job (with visualization tools exposed on ports at the cluster node IPs), and saves artifacts to GCS when done. All of this is controlled via clean and easy internal CLI tools and wrappers to make it simple for any developer. The number one thing this ensures is that no work ever exists in notebook format, beyond tiny scratch work a developer might do strictly to debug code or try a small data proof of concept. The number two thing this ensures is complete reproducibility. Since every possible training task must go through code review, commit all config to version control, get impounded into a container, and execute via a deployed Kubernetes job, it is by definition impossible for someone to execute an ad hoc task that other engineers can’t rerun or have to follow weird setup steps to recreate (it’s all impounded in the container). The third thing it ensures is that all accuracy, monitoring and results artifacts are explicitly tied to the Kubernetes job that controlled the process. It is not possible for some accuracy result to float around untethered from a job ID that uniquely and conclusively ties it to all relevant code for the job. This can be facilitated through MLflow or whatever else. Getting data scientists to “wear corrective shoes” and learn to reorient their way of working to align it with this process has universally paid dividends, both for letting the data scientists experiment faster and more reliably, and for ensuring model training adheres to SRE-related compliance and best practices, so it is pluggable into various tools and constraints those non-ML support teams need in order to do their jobs and offer support to ML teams without getting hit with unstructured notebook spaghetti and bespoke execution paradigms.