4 ms·
In my experience the use of notebooks (exclusively) goes hand in hand with not knowing things such as version control, tests, deployment process or dependency m
by HuwFulcher 6y ago
In my experience the use of notebooks (exclusively) goes hand in hand with not knowing things such as version control, tests, deployment process or dependency management exist.
I don't mean to sound harsh about other Data Scientists from a non software engineering background but the standard workflow is to fiddle around with a notebook until you can get a result. That's as far as it goes, no real robustness to it.
That's a pretty big generalisation but in organisations where they "home grow" their Data Science capability many of the online courses don't cover production level Data Science.
- MrPowers 6y agoYour experience aligns with what I've seen. All the notebooks are in one place. Some are for important production jobs, other are for data exploration. It's easy to make a little edit in a notebook and accidentally break production jobs. Comparatively harder to make an edit in a git repo and do a deploy that'll break production jobs (e.g. if the JAR doesn't compile or the CI errors out cause the tests don't pass). Notebook based production jobs get even more dangerous when NotebookA depends on NotebookB and so on.
- alexott 6y agoTreat notebooks like other code - separate onto staging and production, with defined promotions between them - it’s possible. You can run tests in CI/CD pipelines, etc. You can set permissions so nobody can update production notebooks manually, ...