4 ms·
The question is not whether Databricks claims to be anything. The question is just what is a best practice, from an ETL / DevOps point of view, and how to enfor
by mlthoughts2018 6y ago
The question is not whether Databricks claims to be anything. The question is just what is a best practice, from an ETL / DevOps point of view, and how to enforce ML tooling to adhere to that from first principles.
I’ll give a concrete example. In my org we use Dataproc on GCP as a model training task execution paradigm. You define your base environment via some Docker container, put it in GCR, and then define Dataproc jobs in terms of the base environment, the backing compute resources, any GCS bucket connections, and any ML-specific config like hyperparameters.
A human being never under any circumstances triggers these jobs. Instead a human user deploys the config as a cronjob or regular job in Kubernetes, and then a scheduler picks them up and runs them. For experimental workloads only, developers can manually trigger a Kubernetes job.
Each job consults config, spins up the appropriate Dataproc cluster, runs the job (with visualization tools exposed on ports at the cluster node IPs), and saves artifacts to GCS when done.
All of this is controlled via clean and easy internal CLI tools and wrappers to make it simple for any developer.
The number one thing this ensures is that no work ever exists in notebook format, beyond tiny scratch work a developer might do strictly to debug code or try a small data proof of concept.
The number two thing this ensures is complete reproducibility. Since every possible training task must go through code review, commit all config to version control, get impounded into a container, and execute via a deployed Kubernetes job, it is by definition impossible for someone to execute an ad hoc task that other engineers can’t rerun or have to follow weird setup steps to recreate (it’s all impounded in the container).
The third thing it ensures is that all accuracy, monitoring and results artifacts are explicitly tied to the Kubernetes job that controlled the process. It is not possible for some accuracy result to float around untethered from a job ID that uniquely and conclusively ties it to all relevant code for the job. This can be facilitated through MLflow or whatever else.
Getting data scientists to “wear corrective shoes” and learn to reorient their way of working to align it with this process has universally paid dividends, both for letting the data scientists experiment faster and more reliably, and for ensuring model training adheres to SRE-related compliance and best practices, so it is pluggable into various tools and constraints those non-ML support teams need in order to do their jobs and offer support to ML teams without getting hit with unstructured notebook spaghetti and bespoke execution paradigms.