4 ms·
I personally agree that this does not seem too useful for any DS team that needs to deploy a model to production. But there are a whole horde of DS teams whose
by jpau 6y ago
I personally agree that this does not seem too useful for any DS team that needs to deploy a model to production. But there are a whole horde of DS teams whose outputs are basically PPT presentations.
Think like pricing forecasts, decision modelling, marketing segmentation for product design, ....
I think a common thread to these teams -- at least those that I've seen -- is that they consist of stats->DS backgrounds, and no eng->DS backgrounds.
Many of these teams are orchestrating everything within the notebook. I've seen notebooks that contain complex workflows that extend to 10K LOC. I've lost sleep over such things.
- FridgeSeal 6y agoI agree, and so many data science and data engineering tools all seem to revolve around using notebooks, much to my frustration. I’ve worked in places whose data pipelines were built around seemingly infinite notebooks, all containing consistently poor software engineering. It’s been enough to make me vow to not let people write notebooks that go into prod under my watch lol. I’m constantly on the watch for software engineering focused tools that solve the issues, rather than data science/engineering focused tools. So many are inextricably linked into python as well, which doesn’t gel nicely with anywhere that has multiple languages in the codebase.
- ricklamers 6y agoWe'll allow your team to move parts of your code base from notebooks to scripts (.py/.sh/.R) to alleviate your frustration. That way you can keep using notebooks only for those parts where it makes most sense. They come together in pipelines that are JSON defined and git-versioned.
- calebkaiser 6y agoShameless plug, but I help maintain Cortex, and "software engineering focused tools that solve (ML) issues" is a neat summary of our entire philosophy. For example, instead of notebooks, our model serving platform (https://github.com/cortexlabs/cortex https://github.com/cortexlabs/cortex) uses YAML to structure deployments, and Python scripts to write inference APIs. It's still inextricably linked to Python, but only for writing your API. It's agnostic as to how the model itself is developed, so long as it can generate predictions.
- ricklamers 6y agoCurrently, we are not focused on helping DS teams putting models into production. It's those 'messy' projects with 10K LoC notebook that could easily be broken up into multiple steps (perhaps some of them notebooks, some of them Python scripts with library function usage) that we feel are a great match for pipelines in Orchest today. When a team is still experimenting with what models to go with (i.e. trying neural networks v.s. decision trees) it can be helpful to have a more structured prototyping environment with reproducability and easier scalability. Which is also where Orchest shines. If you want to test some of these models in production, you could easily push artificats to endpoints for serving, in the final steps of a pipeline.