22 ms·
Whenever I can I strongly recommend not using Jupyter for anything more than the most transient tasks. I don’t know whether it’s the Data Science culture or Ju
by HuwFulcher 4y ago
Whenever I can I strongly recommend not using Jupyter for anything more than the most transient tasks.
I don’t know whether it’s the Data Science culture or Jupyter but there is a big lack of discipline in writing maintainable code in DS and non-existent git support is part of that.
I always strongly discouraged developing models using notebooks, instead advocating for using .py files and then using notebooks for sanity checking data.
I don’t have any clever ideas for how we can move past Jupyter but the sooner we do the better.
- maegul 4y agoYes, the Data Science culture around maintainable code does seem to be reaching a critical level of toxicity (in some environments at least). In line with a nephew comment of mine, I feel that bringing the immediate interactivity or iteration cycle of notebooks to the development experience would help a lot, and not be too bad a thing for common development either. I've heard of the related nbdev project, which seems like an interesting and compelling idea. But it'd be nice to see the reverse: something that makes ordinary python development more immediate than using a debugger/vanilla REPL.
- Leo_Germond 4y agoI think that improving the shell experience and allowing e.g. multimedia content to be displayed and manipulated directly into the shell, would help a lot with interactivity. Maybe some specific terminal emulator (like kitty) with ipython would constitute a good starting point...
- tetris11 4y agoIt depends how you use it. When you're still new to the data, using a freeflow text-and-codeblock workflow like jupyter or org-mode really speeds up the exploration phase. Once you have a consistent set of questions, and methods to answer them, then yes, copy off the relevant chunks into their own scripts and source these when using similar data to bring you up to speed, and modify them to your tastes. The issue with starting off with an external script initially is the distracting temptation to refine your code so it can be better used with future data, despite not yet having seen or not knowing what that future data is like. The initial "play and explore" phase of an analysis is very important imo, and notebooks really facilitate that.
- HuwFulcher 4y agoI agree, Jupyter has its place in helping so exploration and learning. A problem that Data Science faces is that the majority of courses don’t show Data Scientists how to progress on from notebooks to write robust training pipelines that are reproducible and safe.
- pplonski86 4y agoLow quality Data Science code is not a fault of Jupyter. The Jupyter allow you to load big chunk of data or some large model only once, and then use it for experiments in other cells. It is hard to replace this feature with plain `*.py` file. For me, this is the killer feature.
- immibis 4y agoBelieve it or not, this also applies to IRC bots. You can edit the code without breaking the connection.
- qsort 4y agoOne of my most upvoted comments says something to the effect of "notebooks bad", so you're preaching to the choir here -- however: - I work with several people who are purely data scientists, and I lean on "culture" rather than "Jupyter". In some circles, probably influenced by academia, programming is considered to be low status work. You are not going to solve the problem by switching to .py files, even though for most tasks literally anything is better than Jupyter. That they don't use git, or that git wasn't originally even a concern, is a consequence of that low-status perception. If you pitched something to the developer community and told them "oh, by the way, you can't use git", they'd synthesize tomatoes out of thin air to throw at you. - I'd carve an exception for stuff that satisfies ALL of the following: (a) is self-contained in a single notebook, (b) has no dependencies on anything non-standard, (c) is demonstrative in nature, or a personal exercise rather than production software. For example, I wrote my solutions to the Advent of Code problems in a notebook and I liked the experience, especially how you could mix math and code.
- HuwFulcher 4y agoI would also lean more towards culture and the environment that Jupyter provides only perpetuates it. I think what made me leave Data Science in the end was that I wasn’t driven by work outside of the notebooks (i.e. coming up with a mathematically superior solution) but driven by building ML driven systems as a whole. I think notebooks are a great way of presenting findings and showing your workings at the same time. If that was their main use then I wouldn’t have any issues
- gradschoolfail 4y agoI would say that in academia, explorability and immediacy is way prioritized over reproducibility and maintainability.. Two kinds of human beings?
- qsort 4y agoThe incentives in academia are different. The objective is to publish, code is important only insofar as it allows you to achieve that goal. This is not to say that academics can't code, but even if you are a professor who cares passionately about making high-quality software, you're fighting uphill, because that's not what you're being evaluated on. If you want to make the argument about "two kinds of people", I think it's more about A-type/B-type data scientists in the industry. I'm really mostly a developer and not a data scientist, but when I assist in DS tasks I wear a distinct B-type hat, and that informs my perspective. A-type people have different priorities and that's fine; my gripe is when you try to import A-type practices in a B-type scenario.
- montebicyclelo 4y agoThe big benefit of Jupyter in the context of machine learning, is that you are often dealing with models that take quite a few seconds to load. You can put big, slow loading, things into memory in the top cells, then try a bunch of logic with them below. Whereas when working with just '.py' scripts, you'd have to reload the model every time, which can make for slow and uncomfortable iteration.
- HuwFulcher 4y agoYes that’s a big plus of notebooks. Hopefully a solution can be found for .py files in future where you can earmark the top part of the script to be cached so the interpreter skips over it
- akx 4y agoWhere would an interpreter cache things if it's not running anymore? The disk? You're back to loading data from disk.
- HuwFulcher 4y agoYep, I don't know enough about the interpreter under the hood but an interactive mode like a debugger where you can go back to a previous line, etc might be the solution. I doubt that's high on the priorities of the Python team though.
- nidnogg 4y agoOne alternative to loading models in .py scripts is making use of joblib's dump() and load() methods for pipelines. https://joblib.readthedocs.io/en/latest/generated/joblib.dump.html https://joblib.readthedocs.io/en/latest/generated/joblib.dum.... That way, if you put your classifiers in joblib pipelines, once you're done with fitting steps you can just export your trained classifier with: joblib.dump(pipe, "trained_classifier.dump") And resume your work with: joblib.load("trained_classifier.dump") Considering this works for any Python object, a lot of heavy lifting can be exported for later (swift) use this way.
- 4y ago
- medo-bear 4y agoi strongly agree with what you are saying about Jupyter, however i strongly disagree about using netobooks in general (literal programming) one of the key things that a good notebook system must allow you to do is to mix something like markup format + LaTeX + source code. writing math-heavy documentation and explanations is simply impractical and limited (readability suffers) if done in comments. jupyter however is severely limited as it is unreadable in its raw format and therefore does not play well with a version control system such as git instead there is a solution that allows one to do everything jupyter does good with the additional benefit that it plays with version control really well - ie org-mode [1]. the only difference is that instead of using a browser to interact with it, you use emacs. the added benefit to this is that you can also use full-featured key bindings (emacs / vim) and even integrate a language server for auto-completion [2] EDIT: moreover the list of supported languages in orgmode far exceeds that of jupyter [3] (or did the last time i made this comparison) [1] https://orgmode.org/ https://orgmode.org/ [2] https://emacs-lsp.github.io/lsp-mode/manual-language-docs/lsp-org/ https://emacs-lsp.github.io/lsp-mode/manual-language-docs/ls... [3] https://orgmode.org/worg/org-contrib/babel/languages/index.html https://orgmode.org/worg/org-contrib/babel/languages/index.h...
- Grumbledour 4y agoThere is always a lot of org-mode promotion on here when the topic is interactive notebooks. And I get it, people love it and it solved many of the problems other systems have. But org-mode users need to understand that the one thing holding org-mode back is simply emacs. I know you probably all love it, but everyone else is not interested in breaking of their fingers by learning obscure key command chains just to use org-mode. Sorry, but that is just the reality. If someone can implement the majority of org-mode in a better editor, there might be more users interested. But as it stands, it's just to much of a hassle.
- medo-bear 4y ago> I know you probably all love it, but everyone else is not interested in breaking of their fingers by learning obscure key command chains just to use org-mode. Sorry, but that is just the reality i'm sorry to burst your strong held convictions but you can choose any of the following a) use any key-bindings you like including emacs, vim, cua, or combination of b) use org-mode without any knowledge of more advanced emacs commands (except basic knowledge of using an editor) c) drink some milk (gotta have strong bones) and learn how to use the emacs system including emacs lisp and have one of the most advanced computing environments in existence at your service sorry, but that is just the reality
- wasimlorgat 4y agoI'm always surprised when people advocate for .py files over notebooks because of poor software practice. (Genuine question) have you found that it improves the situation at all?
- HuwFulcher 4y agoI’ve found varied success. In general, I’ve encouraged the move across to being teaching source control. That has been in contexts where notebooks are being used for critical outputs rather than exploration. When you get into MLOps as well, having .py templates actually makes the Data Scientist’s job easier as they can plug and play their models into a system that tracks inputs, outputs and changes for them
- cantagi 4y agoYes, people writing unmaintainable code in Jupyter notebooks is a problem. Personally, I start every notebook with %load_ext autoreload %autoreload 2 then develop production quality code in .py files.
- shapefrog 4y agoWell that has improved my life - thanks!
- etrautmann 4y agoI didn't realize anyone didn't do this. Totally essential, great point!
- jstx1 4y agoMy workflow is: 1. Experiments in notebooks. Notebooks are saved under git but mostly as a backup, I don't care how nicely they play together. I don't get why you would discourage notebooks for running experiments, doing it with .py files sounds kind of miserable. 2. Services and library code in .py files, under version control, just like any other software we write.
- HuwFulcher 4y agoExperiments using notebooks are fine as long as they are well documented. Having your services and library code as .py files you can import in is great. The issue comes with how to move from experimentation to deployment. If you already have services/library code as .py files you make your life a lot easier. The issue comes when everything is spread across multiple, poorly documented notebooks. If you're working with an MLOps team it makes their life a nightmare to take those notebooks and conform them into something usable. Jupyter is great when it is used in the right way.
- targafarian 4y ago100% agreed. People use a great tool in a poor way and then broadly condemn the tool. And any tool that is sufficiently flexible to be broadly useful can be used in very poor ways. Jupyter is great, it gets me over the barrier potential for starting a task every time. I build and prove out an algorithm/task piece by piece. Once I'm happy, I move the meat of it to a function in a .py file, and move the code I used to test the algorithm to a unit test function. Delete the duplicated bits and replace with imports, and then what remains is a tutorial/demonstrator notebook using the function I wrote and maybe some nice plots to go along with that, that I wouldn't put in a unit test (nor that show up in docstrings). This can be converted to sphinx docs if the code gets big enough. What a great tool for incrementally building software! In my world, I build brick by brick, not all at once. Jupyter is a key to that process.
- mFixman 4y agoJupyter is not great for collaboration with multiple people editing, but with a little bit of order it's perfect for in-person working and presenting that work. Notebooks can be clean if you follow some rules: 1. Code flow always goes down: holding Option+Enter should execute all fields without any errors. Don't do `x += 1` if `x` is defined underneath. 2. All blocks are idempotent: running any block 5 times should produce the same result as running it 1 time. Don't do `x += 1` unless `x` is defined in that block. 3. Keep block-local variables short and block-global variables long. Don't do `x += 1` unless you are not using `x` anywhere else. Also, the Table of Contents extension [1] is a life-saver for making long analyses workable. [1] https://jupyter-contrib-nbextensions.readthedocs.io/en/latest/nbextensions/toc2/README.html https://jupyter-contrib-nbextensions.readthedocs.io/en/lates...
- z3c0 4y agoAgreed on all points. Notebooks really aren't that hard to maintain. They just require some slightly different rules from standard scripts. Personally, I like to label block-global variables in capital case (like PEP8 constants), so as to make them easy to spot. Being formatted like constants also causes me to think twice about altering it after instantiation.
- cinntaile 4y agoIt would be great to have tools available that force these rules on you.
- da39a3ee 4y ago> with a little bit of order it's perfect for in-person working It's not perfect for in-person working because a single person should always keep their work under version control, and they should be able to view meaningful diffs to understand the history.
- analog31 4y agoI have a rule that helps with hidden state and out-of-order execution. Once in a while I do a "restart kernel and run all cells." If doing that breaks anything, then I have to fix it. But it also ensures that a notebook is reproducible later on. Of course I don't have things that take hours to run. It would be nice if there were something that would make out-of-order problems light up, the way that code editors can highlight errors while you're editing. A limitation of "browser as editor" is that it misses out on some of the powerful things that code editors do today. Another thing is to put things in functions, so temporary variables are disposed of. That's a halfway step to putting things in .py files. A benefit if .py files is not always that jupyter is bad, but that variable scoping is good hygiene.
- scombridae 4y agonot using Jupyter for anything more than the most transient tasks While most programmers have reached this conclusion, they're generally not day-in day-out jupyter users. They need to understand *everything* is transient for scientists who optimize for proof-of-concept and publish-and-forget-it paper writing.
- frumiousirc 4y ago> *everything* is transient for scientists who optimize for proof-of-concept and publish-and-forget-it paper writing.* Which itself is a huge problem. Happily this mindset is changing, at least in some scientific all fields. For example, in particle physics proposals a document ("data management plan") much be written describing how that unconscionable attitude will not be taken with the experiment's data and software. That said, this transient mindset and derision of real software skills is still fairly prevalent in this field.
- scombridae 4y agoWhich itself is a huge problem More "nature of the beast" in my opinion. Science measures itself by how many alluring women it can date; engineering, by how long it can keep the wife happy.
- frumiousirc 4y agoExcept for the fact that some experiments are taking decades themselves or are one part of a long progression of related experiments. So continuity of software and data through generations of students, postdocs and even professor types is needed. Even for short-lived experiments reproducibility is important. So much of today's experiments ultimately rely on complex stacks of software to get their results and on data which humans can afford to acquire only once. Preserving both is necessary for future re-validation or reuse. This problem should not be excused.
- LeanderK 4y agoIf you work with something visual, interactive then this workflow is so super awkward that I never end up doing it. For data-driven workflow you have to analyse the data, note down your thoughts, analyse a bit more and then come to a conclusion. Your conclusion might be code living in .py files, or another type of data then consumed by something else. But this will result in a significant part of the "thought-process" and relevant code living in those notebooks, with all their problems. I can't just switch to some .py files because I want to change the axis for some plot, or look at it in log-scale. But then where do you draw the line? A .py file for only 10 lines of code generating the resulting .csv? That's also a pain to maintain because you have all those disconnected files. We need those notebooks, they have to get better.
- kriro 4y agoIt all depends on the context. In academia it is a great tool. I can set up a couple of notebooks on our GPU server and give many students access to powerful GPUs without having to worry abbout shell access etc. Aditionally they are ready to go and do interesting things immediately and don't have to install the environment on their laptops (which might be win/linux/mac but at least these days that's easier but still extra work for them). I also use it a lot for experimenting, parameter tuning etc. It's not too bad to have it explicitly distinct from production level code. Run/tune/experiment in notebook, once you're happy with the model -> code it up in .py file(s). Also great for quick presentations :) However, the fast.ai team is actually doing a pretty solid job running everything off notebooks. So if I wanted to go that direction (and skip the .py files) it's that project I'd look at for how to do it.
- carderne 4y ago> without having to worry about shell access etc Do you mean you _don't_ want to give the students shell access? By default you can run shell commands from within a Jupyter notebook be prefixing them with `!`.
- rovr138 4y agoThey might mean without having to setup individual user accounts for the server. https://jupyter.org/hub https://jupyter.org/hub and if it's just the professor's lab, just a jupyter lab instance with one password works too.
- jhrmnn 4y agoI think of Jupyter Notebooks as scratch paper on my desk. It's not to archive things, it's for developing ideas. Once ideas are developed, I transfer them to a long-term medium (LaTeX or Markdown document, Python source file, etc).
- moonshotideas 4y agoSame, it’s perfect for “work in progress code”, and working out a problem step by step. I’ve always wanted this environment in other languages
- ajford 4y agoYep. I worked in scientific applications, and when developing some new data cleaning and processing pipelines for our hydrology data, Jupyter was phenomenal. It was easy to use as a presentation, with figures and plots embedded. With controls enabled, you could demonstrate what varying certain parameters would do and pitch proposed cleaning profiles. I was rather easily able to send a directory and it's notebooks/data sources to colleagues in the water sciences team so they could validate my results on their own (they were luckily also familiar with Python and Jupyter), and caught some minor bugs in the pipeline. This was all much more collaborative and concise, and I feel Jupyter played a huge part in it. Once it was done, it and a "final draft" pdf were added to the Docs in the repo and the pipeline was written out into a full application of it's own.
- el_oni 4y agoI like to use vscodes code blocks. Add in # %% and it makes the code executable in isolation. Like a cell in jupyter. This means i can write a py file. Execute bits indepentantly and when im ready to check it in just remove the block comments
- lake_vincent 4y agoOh man, thank you. I grew up with C++ and Java as my main languages, so I always feel more at home with a py file. Notebooks never caught on with me.
- fifilura 4y agoIt can be very useful to run recurring jobs (e.g jobs that run once per day) in a notebook to add the output as kind of advanced logging. And then serve the results as a static page under my.logs.intranet/my-job/2022-02-14/my_recurring_task_notebook.ipynb.html You can get so much more context regarding what went well or wrong compared to browsing through log lines in some more or less user friendly tool.
- tuukkah 4y ago> non-existent git support From the beginning of the article: "With nbdev2, the Jupyter+git problem has been totally solved. It provides a set of hooks which provide clean git diffs, solve most git conflicts automatically, and ensure that any remaining conflicts can be resolved entirely within the standard Jupyter notebook environment."
- slewis 4y agonbdev2, which this article is about, is a solution to this problem. It makes notebooks testable, composable, versionable, and more.
- mistrial9 4y agodude - more than 1 million undergraduate computer science students worldwide will learn Jupyter this Fall, and you are getting contrarian votes among a bunch of average-of-masters+industry CS people here "we" have to learn and teach the next sets of people new to computer science
- HuwFulcher 4y agoTotally agree with you. I’ve been teaching people over my career so far, with varying degrees of success
- operator-name 4y agoOthers have mentioned the usefulness of literate programming so I won't reiterate that. Partially the lack of discipline comes from the implicit data dependancies between cells. Variables are all globally scoped and unless you ensure the notebook can be ran top to bottom its easy to introduce subtle bugs. I believe Julia's https://github.com/fonsp/Pluto.jl https://github.com/fonsp/Pluto.jl solves this issue quite well. Another part comes from cells that should really be functions. In my opinion this is because functions are 2nd class citizens compared to cells, and could be improved with UI (function cells? node based programming?). Programming is more than just manipulating text, so why shouldn't tools move in a direction of just being fancy text editors?
- bobbruno 4y agoYou're thinking it the wrong way. Notebooks don't do well in software development, but they are extremely useful on exploratory data analysis and quick iteration when searching for a suitable modeling approach. These two tasks use code, but for completely different purposes. A DS is working on the data, understanding it and trying to identify what information it may have. Then they try to find a model that will leverage that information to deliver whatever inference solves the business need. This is extremely interactive and iterative, and everything from the actual business problem to the ML approach may change at each iteration. Imposing software development practices at this point is disruptive to the train of thought, which is very burdened already by the level of uncertainty and all the mathematics required to understand the data results. The goal is to find a viable approach, not write production code. Once this approach is found, a good clean-up/refactor is strongly recommended, to then start a proper software development that will create a live product from the found approach. I call this the switch between research mode and development mode, and it has strong parallels to the way R&D is done in many industries. I believe a lack of understanding of this dual nature of ML is what causes many of the problems in MLOps: plans that don't take into account the research time and risk, mixed teams where engineers don't understand the initial nature of DS work, attempts to put notebooks containing research code in production, etc. Even planning for the refactor doesn't solve it all - what will happen when the next generation of a model has to be created? Will the refactor Ed code be forced on the DS and ruin their research productivity? Will they start from scratch again and not only lose all the refactor/dev cost but also make this a recurring cost? I have been looking for answers for this for years now, and found none so far. Source: I've been working with data for 27 years, as a data engineer, data architect and data scientist. When I do DE, my code is considered high quality by my peers, but when I'm doing DS research, I know I write bad code - and I won't change that. It's more productive to work this way and do the big refactor (possibly leaving the notebook env behind along the way) than the alternative.
- OkayPhysicist 4y agoWriting maintainable code is basically what defines the role of software developer. Pretty a lot of technical roles --engineers, scientists, hell, even economists-- pick up enough programming experience to hack stuff together. And those hacked together solutions are universally a nightmare to work on afterwards.
- prepend 4y agoThis is similar to my approach with all logic in a .py module and then the notebook is only for presentation, comments, and formatting. This works out fine with everything in git and it’s pretty rare to actually have a conflict in the .ipynb.