6 ms·
How to put machine learning models into production
- simonebrunozzi 6y agoOverall, a well written article. If you're interested in ML Ops, I have a shameless plug to share: on November 19th I host a free online panel, "Rage Against the Machine Learning", with industry experts. [0] [0]: https://cotacapital.zoom.us/webinar/register/8116020076218/WN_DIIptnvUQhi0AkSze_XhAw https://cotacapital.zoom.us/webinar/register/8116020076218/W...
- tixocloud 6y agoThanks for the share. Will check it out.
- bpesquet 6y agoRegistered! BTW, awesome panel title.
- simonebrunozzi 6y agoI'm so glad you like it! Sometimes you're not sure whether people will get from your words the sense that you had in mind :)
- fphhotchips 6y agoThe title doesn't really match the article in my mind. To me, it talks about everything but actually deploying a machine learning model in production. In particular, there are a lot of words around where training data is stored. In my experience, the training data is really more part of the the development process than the actual productionisation of the model. That said, there is a piece here on TFX, which is valuable in this context. I also think the advice about going with proprietary tools that speed up the process is good. Tools like Microsoft's AI tooling, Dataiku and H20 are good in that context. I would have liked to have seen some discussion around when you should deploy a model as an API vs generating batch predictions and storing them - I've done both on a test bench, but I don't really know how well the API scales.
- dcl 6y ago> To me, it talks about everything but actually deploying a machine learning model in production. This seems to be a common theme of a lot of articles about 'how to put ML models in to production'.
- shajznnckfke 6y ago“Deploying a model” is sort of a nebulous concept. You probably have some kind of server, which loads a serialized model and runs data from requests or from batch files through the model to get predictions. Which part are we deploying? The service can be deployed like any other service. The model is probably a file. You deploy it by copying it to s3 I guess? You can go more in depth, and that’s what the article is about.
- shoo 6y agoIf you don't expect to need to tweak the model parameters very often, and it has a simple form (decision trees, linear regression) and forms a sub component of an existing application, another deployment option is just hard coding the model as a library or even an expression amidst the rest of the application code.
- shajznnckfke 6y agoThat’s a good option too. Sometimes the model consists of some training and inference code, which can be loaded as a library, plus a bunch of weights, which may be large enough to warrant a separate file. Either way, I think the deployment of the model isn’t really a hard problem. Validation of model quality, keeping track of what code was used, and what training data, making sure the updated data is where it’s supposed to be before training, tying all these parts together in some kind of comprehensive management system are all harder problems I think.
- laichzeit0 6y agoThere are some subtleties to deployment. If you trained a model with say sklearn version X and the component serving requests deserializes it but it’s using sklearn version Y then things can break badly.
- calebkaiser 6y agoI was expecting this to be more about running inference in production, though the information in the article itself was interesting on its own. There does seem to be a dearth of writing on the actual topic of deploying models as prediction APIs, however. I work on an open source ML deployment platform ( https://github.com/cortexlabs/cortex https://github.com/cortexlabs/cortex ) and the problems we spend the most time on/teams struggle with the most don't seem to be written about very often, at least in depth (e.g. How do you optimize inference costs? When should you use batch vs realtime? How do you integrate retraining, validation, and deployment into a CI/CD pipeline for your ML service?). Not taking anyway from the article of course, it is well written and interesting imo.
- sgt101 6y agoThere seems to be the idea that training an ML model is like compiling code - but every "compile" leaks information into the training pipeline. Repeated testing and choosing (unless it is on a fresh draw) is an optimization step, you are optimizing on the test set. Using a fresh draw is difficult and expensive, especially since the labels may not be available. Using A/B is expensive, multi-armed bandits are more efficient, but again there is an optimisation element there (waits for shouting to start) Additionally surely there is a really significant qualitative judgement step about any model that is going to be used to make real world decisions?
- mlthoughts2018 6y agoYou don’t typically perform optimizations iteratively with feedback from the final test set. Instead you split your training set into validation and training, and you iterate on that, leaving your true hold out test set completely unexamined all along. You would do model comparisons, quality checks, ablation studies, goodness of fit tests and so forth only using the training & validation portions. Finally you test the chosen models (in their fully optimized states) on the test set. If performance is not sufficient to solve the problem, then you do not deploy that solution. If you want to continue work, now you must collect enough data to constitute at minimum a fully new test set.
- gerbler 6y agoThere's a great paper from Google about this "Machine Learning: The High Interest Credit Card of Technical Debt" [0] which discusses why you should use a framework to deploy ML models (the authors are involved in developing TFX). In my experience, spending time explaining results to the business is also a very time consuming element of deploying a model too. 0:https://research.google/pubs/pub43146/ https://research.google/pubs/pub43146/
- sandGorgon 6y agois anyone running TFX in their companies in production ? how has the experience been ? since like everyone is on K8s, im wondering if kubeflow is not the more natural fit
- mlthoughts2018 6y agokubeflow is pretty horrendously bad unfortunately. Most of the installation docs are incomplete and inaccurate, and since the workflow requires building a separate container for each submitted task (instead of separately specifying version control commit) you cannot actually get reproducible results. You’d have to scrape the state of the code out of the identified container tied to a job, since the circumstances under which the container was created for the job can be any arbitrary, out of band changes a developer was working on, such as from a branch they never pushed. This workflow also doesn’t work well in hybrid on-prem + cloud environments because, for example, your model training might run in a cloud Spark task, but your CI pipeline (responsible for building and publishing a container to an on-prem container repo) might run on-prem. kubeflow, for example, has a hard requirement to put containers into cloud container registries, and makes assumptions about what the networking situation is allowing connection between on-prem and cloud container resources. I think industry shifting focus to kubeflow is actually a giant mistake.
- joana035 6y agoAirflow is a good replacement and works well, easy to deploy and to add datasources/steps.
- mlthoughts2018 6y agoAirflow is only a DAG task executor, which hardly scratches the surface of what is needed for managing experiment tracking, telemetry for model training, and ad hoc vs scheduled ML workloads. Airflow is useful as a component of an ML platform, but it is only in principle capable of addressing a really really tiny part of the requirements. You also need to ensure Airflow can easily provision the required execution environment (eg. distributed training, multi-gpu training, heavily custom runtime environments). Overall Airflow isn’t a big part of ML workflows, just a small side tool for a small subset of cases.
- dtjohnnyb 6y agoI've recently come across the MLOps community here https://mlops.community/ https://mlops.community/. The meetups are all on YouTube and have great topics like putting models into production, but also more interesting ones (to me) like ml observability and feature stores. Their slack channel is great too, learned a lot about the reality of using kubeflow vs the medium article hype
- steve_g 6y agoAs a practical detail, I'm wondering if it always makes sense to wrap your predictor in a simple if-then based predictor. If your learned model makes bad predictions in certain specific cases, you can "cheat" with Boolean logic. This could also be useful when the business has a special case that doesn't follow the main patterns. Any thoughts on that?
- mlthoughts2018 6y agoThis is almost always a very bad idea because the if/then condition you describe is business logic - as in, how do you recognize the business situation when you want to decline a prediction, and how does that change? It is very complex because most of the time there is no simple rule such as a threshold on the confidence score of the prediction. In practice it might be more like, “if the user has more than 7 items in their cart and if the user is not a returning customer that filled out personal data and the value of their cart is greater than $100 and they have not put a new item in the cart for 2 minutes, and the confidence score of the predictor is less than 0.4, THEN don’t show the next recommended item, just display a checkout link.” And the number of items, the cart value limit, the time since last item-add, etc., will all be hotly debated by product management and changed 5 times every quarter. In grad school there was a professor who said of machine learning that “parameters are the death of an algorithm” - so you want to avoid coupling extra business logic parameters tightly with the use of machine learning models.
- andersonvieira 6y agoI understand it may not be recommended in complex situations, such as you described, but I think @steve_g's idea may be interesting for some scenarios. I work in automatic train traffic planning, mainly for heavy-haul railways [1]. Recently, we've been working on a regression model to predict train sectional running times based on historical data. As our tool is used during real time operation, we can't risk the model outputting an infeasible value. So we're thinking about defining possible speed intervals, e.g. (0km/h, 80km/h] for ore trains, and falling back to a default value if the predicted running time causes the speed to fall out of this range. [1] https://railmp.com/en/our-solution/ https://railmp.com/en/our-solution/
- londons_explore 6y agoToo many people focus on "properly" putting ML into production... I'd like to propose an alternative... Build a model (once) on your dev machine. Copy it to S3. Do CPU inference in some microservice. Get the production system to query your microservice, and if it doesn't reply in some (very short) timeout, fallback to whatever behaviour your company was using before ML came along. If the results of yor ML can be saved (eg. a per-customer score), save the output values for each customer and don't even run the ML realtime at all! Don't handle retraining the model. Don't bother with high reliability or failover. Don't page anyone if it breaks. By doing this, you get rid of 80% of the effort required to deploy an ML system, yet still get 80% of the gains. Sure, retraining the model hourly might be optimal, but for most businesses the gains simply don't pay for the complexity and ongoing maintenance. Insider knowledge says some very big companies deploy the above strategy very successfully...
- richiecute 6y agoYeah, there was a multiple dev long effort to build "complete" solution but I did what you described to satisfy impatient stakeholders. Since then the fancy solution is incomplete and mothballed while the simple deployment has been running ever since. The next value add was consolidating data in a db with web ui so relevant non-devs can view and help add to it + easily integrate different data sources with automatic validation. I wish their was a nice open source thing here you could spin up without much effort. There's middle ground between complex/paid full solution that can scale to huge data sets and integrate with turking etc and emailing csvs or sharing google sheets links around.
- closed 6y agoAgreed! It seems like there's a lot of power in having "one to beat". As in, let's get some model up before we worry about one that needs daily updates.
- TTPrograms 6y agoYou only get 80% of the gains if your data distribution is stationary and/or you can construct a non-ML solution with near 80% of the ML one. If a non-ML solution is that close to non-ML then I would just stick with that - the advantages in terms of predictability and ease of maintenance would likely outweigh small performance gains. A model that can't be continuously trained inevitably rots due to data, interface or environment changes (and the code is typically very difficult to maintain across team members - if the author leaves it's often a ticking time bomb). If you're OK with the model rotting then it wasn't that important to your business to begin with. This is not true for all businesses.