9 ms·
Why software engineering processes and tools don’t work for machine learning
- rcar 6y agoI found myself wanting a little more content here. Felt mostly like hand waving followed by a plug for their product.
- coddle-hark 6y agoYeah, half the article was commentary on an Andrew Ng keynote and the other half was an ad for their employer.
- coddle-hark 6y agoThis post argues that ML is experimental in nature which makes it harder to plan for. That’s true of any R&D project though! Academia is typically not run on kanban boards, for example.
- ska 6y agoNeither are early stage R&D projects in industry. It's a fallacy to believe that there is a bright line distinction though, and that no engineering practices or tools will improve your research productivity. Traditionally e.g. academic labs traditionally eschew all such approaches not because they wouldn't benefit, but because they don't understand what they are missing and have little incentive to learn, and few opportunities for it.
- lumost 6y agoMost processes have benefit proportional to the number of participants. As an individual little process is required, and spending time maintaining the process is extraneous overhead. Academics tend to want maximal freedom ( in exchange for minimal compensation ) and adding process that constrains their work is unlikely to be productive. Most ML projects in industry are leveraging understood techniques in novel ways to build real customer facing products within reasonable investment timelines. To say that one can't provide any scope or estimates on these seems weak.
- ska 6y agoFair point, but it's not so clear cut. A lot of process is about effective communication over largish groups and/or over largish times. Neither of these really apply in academic research work most of the time. On the other hand, there are practices that can save you a ton of time even if you are working alone. In an academic setting this can often mean you just learn the ineffectual techniques of whomever came before you. In this case a small amount of education and change of practice (i.e. process) absolutely can dramatically improve overall lab productiveness. Industry isn't immune at all - especially when treating your "R&D" group quite separately from your "production engineering". Another thing to remember is that there is no such thing as no process; you have one even if you don't understand it. It can be lightweight and almost invariably you can improve it (thought that's not always worth the time spent)
- NordSteve 6y agoI stopped reading the post once I got to the drawing of the waterfall diagram. There's not a single service I'm aware of that runs today on a waterfall process. Not sure how you can trust anything else in the post with that obvious of a strawman argument. I'm not going to spend the time trying.
- sjg007 6y agoPlenty of waterfalls out there.
- fizixer 6y agoThis is a nitpick on "... not a single ..." claim by GP. GP is otherwise right. First figure in the article is a joke. And that makes the article a joke. Plenty of waterfalls out there. Even more non-waterfalls, aware of iterative, modern process, who will kick you out of the discussion if you try to sell them waterfall. ML processes need innovation over those of traditional s/w dev. This article is the worst place on the internet to look for that innovation.
- sjg007 6y agoPlenty of waterfalls in ML pipelines as well.
- fuzzfactor 6y agoThe circular chart example given for Machine Learning is virtually identical to a cycle of continuous improvement in a service or manufacturing operation, except for one little detail, with Continuous Improvement you don't want the minimum viable product, you want the Maximum Viable Product. Seems like a machine learning outfit would really benefit from a leader who is smarter to begin with, a better learner, and already more skilled in the task that you want the machine to do.
- talal7860 6y agoI kind of did the same.
- tootie 6y agoI've found the same to be true of any sort of creative engineering project. It's common to have an ill-defined end goal "must be cool" or "is it art?" which is just inherently subjective.
- legerdemain 6y agoI've seen a lot of the software engineers on teams I've worked with complain about the way data science and machine learning professionals work with code, and They. Just. Don't. Get. It. Stop trying to push the tools and processes that you developed for yourselves onto us! People I work with don't care that our notebook tools have rudimentary syntax highlighting or type-checking. We don't rely on our tools to make sure our code is written correctly. We just memorize the APIs and rely on experimentation to get things working. Please don't make us use git. Checking out and checking in files is too slow! Oftentimes, Slacking code to a colleague is the best way to collaborate, and the fastest. Don't make us use confusing and broken authentication schemes to access our data! I work best if I can just download live data to my machine and work locally. If it's too big or it's something like streaming, just give me an admin password that I can save in my notebook and use to hit some database clusters. Nothing wrong with that!
- squarerootof-1 6y agoIs this sarcasm?
- Areading314 6y agoI'm hoping this isn't even sarcasm, just GPT-3
- lumberjack 6y agoYes.
- Jugurtha 6y agoIt most definitely is, but there's a large part of truth in it. You can try and train people to write good code or follow best practices but from what I have seen, you're swimming against strong currents. What we have here is the: "I don't get it, why don't you read this 500 pages manual and type in those 10 commands to print your document" era of computing. People wondered why people had trouble "just" programming something. That is not the job to be done. There are huge frictions and broken interfaces between roles: developers, engineers, data scientists, domain experts with everyone expecting the other to "just" do X. If data scientists just wrote better code. If developers just were more familiar with machine learning. If ML engineers just could be more CICD aware. If domain experts just could get it right. This is not working.
- miemo23 6y agoNot a great article. The lead-graphic misrepresents software development processes, the machine learning process is just an agile process which as a "software engineering process" directly counters the article's title, it clearly does work for machine learning.
- ccvannorman 6y agoYes, a quick Google search for Agile Vs Waterfall yields a comparison that looks exactly the same, so indeed Agile processes do work with Machine Learning (at least at a high level). Nothing really novel here, except "waterfall doesn't work for ML". https://www.digite.com/blog/waterfall-to-agile-with-kanban/ https://www.digite.com/blog/waterfall-to-agile-with-kanban/
- fizixer 6y agoFirst figure is a joke. As many in the comments have pointed out about the actual situation, the two process models in the figure need to be swapped (better yet: replace ML model with the words "no rhyme or reason ad-hoc BS"). As it stands, ML, data-science, computational-science, acadamics are the worst offenders when it comes to following a sane s/w engineering model for whatever they're doing. (for crying out loud, some even refuse to use anything other than FORTRAN-77 ... !!!) Every ML, data-science, academic team needs to hire at least one traditional s/w engineer experienced in proper modern s/w development. And they need to take that person seriously instead of being jackasses thinking they have the PhD degrees, or are PhD candidates, and that person doesn't.
- adolph 6y agoThe article seems to be getting hugged to death, maybe try reading Hidden Technical Debt in Machine Learning Systems https://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-systems.pdf https://papers.nips.cc/paper/5656-hidden-technical-debt-in-m... The Abstract: Machine learning offers a fantastically powerful toolkit for building useful complex prediction systems quickly. This paper argues it is dangerous to think of these quick wins as coming for free. Using the software engineering framework of technical debt, we find it is common to incur massive ongoing maintenance costs in real-world ML systems. We explore several ML-specific risk factors to account for in system design. These include boundary erosion, entanglement, hidden feedback loops, undeclared consumers, data dependencies, configuration issues, changes in the external world, and a variety of system-level anti-patterns.
- jpgvm 6y agoBecause software totally doesn't require prototyping, experimentation and measuring of intermediate progress. /s Dealing with data teams that just won't learn reasonable software engineering practices but still want to run software in production is incredibly painful. Eventually you just have to sit down with them and lay down the ground rules on what is and isn't acceptable. Once they get over the initial stage of complaining they eventually convert and can't understand why they ever did it wrong in the first place... just like junior software engineers. The snowflake part where they feel they are somehow different and entitled to use poor processes and force poor tools and ecosystems on everyone else though really does irk me. Yes your field is quickly expanding, yes your salary is probably over-inflated vs your competence and experience but that doesn't make you special.
- kelnos 6y agoMost of the data science / ML code I've seen reminds me of academic researcher code. Hack and experiment until it kinda works, never write tests because what's the point, it works on my machine and that's good enough. For a ML-heavy project I ended up being the "productionize this stuff" guy. My favorite was a bunch of python code that depended on a library that under the hood would spin up a new JVM instance to do some NLP work every time it was called (or whatever incorrect way they were using it caused this to happen). And they wanted to do this in the context of a synchronous HTTP request/response cycle. My overall impression was that these were incredibly smart people, but they had no interest in understanding the context in which their code was going to run, or in understanding business requirements around reliability or maintainability. They just wanted to hack on cool projects that they found interesting. Given that, it doesn't really surprise me that a lot of ML/data science people don't "get" accepted software development processes. There were some who made the leap and learned to understand why some of these practices are useful and tend to give you a better product, but they seem to unfortunately be in the minority. This is a pretty young field/subfield, and I think it's to be expected that practices are pretty raw and underdeveloped.
- sbpayne 6y agoI often see comparisons of traditional software engineering vs machine learning like this where the former is a linear sequence of stages without cycles and the latter is some kind of tight iterative loop. But my work in both building backend services and building machine learning models has always looked like how the author describes machine learning models. In my experience the "process" generally isn't that different; but the cycle times are long enough for machine learning that it can feel different.
- kthejoker2 6y agoIt's sad to see such "to hell with it" naivete (and the comments here with broad supporting evidence that data scientists are poor software developers. There are also a lot of DS people wrestling with the deeper challenges at the heart of MLOps and beyond just advocating sane team-based workflows, are looking at ways to harden and structure ML specific processes in the planning, design, build, and deploy phases. Sean Taylor at Lyft on choosing metrics - https://medium.com/@seanjtaylor/designing-and-evaluating-metrics-5902ad6873bf https://medium.com/@seanjtaylor/designing-and-evaluating-met... Test coverage of ML models and a good discussion on why and how ML code testing isn't the same as traditional testing - https://www.jeremyjordan.me/testing-ml/ https://www.jeremyjordan.me/testing-ml/ A tool for Behavioral testing of NLP models - https://github.com/marcotcr/checklist https://github.com/marcotcr/checklist How to be a 10x data scientist (spoiler: have a professional end to end workflow) - https://towardsdatascience.com/how-to-be-a-10x-data-scientist-4718accf7d3f https://towardsdatascience.com/how-to-be-a-10x-data-scientis... Ml-specific technical debt, how to measure it, some ideas on how to handle it, this paper has spawned a lot of great tools - http://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-systems.pdf http://papers.nips.cc/paper/5656-hidden-technical-debt-in-ma... Adversarial instance and outlier detection - https://github.com/SeldonIO/alibi-detect https://github.com/SeldonIO/alibi-detect And of course a whole suite of tools and research beingn done on data, pipeline, and model management, optimizing deployment and serving, explainability and interpretability, all in the service of applying enterprise-ready, engineering-grade rigor to data science. And then you read something like this and want to tear your hair out.
- blackbrokkoli 6y agoThe nicest possible explanation is that this article is an excellent piece of sarcasm. The very first thing in the article is a comparison of Waterfall and Agile but it is labelled "Software Engineering" and "Machine Learning"?! Like, I am so confused. Is the implication that AI researches invented Agile or something? Literally everybody does (or pretends to do) some kind of Agile these days. And I don't mean this in a "my enlightened startup does software dev like this" way, big corporations have the actual job title SCRUM master for years now. I did mom-and-pop store Wordpress projects with "machine learning methods" apparently - nice! After wording this out I brought myself to read the whole article, and, well, it is a marketing piece for their productivity tool Comet which solves all the problem you have by building AI, which does not have "The provable correctness of software engineering" (I wish software would be provably correct :| ).
- Jeff_Brown 6y agoI read so many of those words and learned so little.
- ganti_r 6y agoThe fundamental difference in ML of you work with data, where as in s/w you work with rules. It's a subtle difference but I believe is very profound. It's very hard to develop intuition around data and most s/w devs treat data as rules. It's not entirely the fault is DS folks who happen to come from non s/w background and vice-versa. In my mind the management should understand that their is a gap and seek to help close the same.