13 ms·
Version control for Jupyter notebooks was one of the biggest complaint I had. Specifically, diff and merge with the JSON files (.ipynb) is ugly. I built Review
by amirathi 8y ago
Version control for Jupyter notebooks was one of the biggest complaint I had. Specifically, diff and merge with the JSON files (.ipynb) is ugly.
I built ReviewNb[1] to solve one of those problems (diff). Note that, there is nbdime[2] which works well for local diff/merge. The idea for ReviewNb is to have much tighter integration with GitHub etc.
[1] https://reviewnb.com https://reviewnb.com
[2] https://nbdime.readthedocs.io/en/latest/ https://nbdime.readthedocs.io/en/latest/
- stephengillie 8y agoDiffing JSON as text must be painful. Diffing JSON as data should be somewhat simple.
- justin66 8y agoAre there any merge tools that offer features for this a lot more sophisticated than basic text comparison?
- stephengillie 8y agoPowershell has Compare-Object, which will diff .NET objects. It has the convenient alias "diff". JSON can be converted to .NET objects by ConvertFrom-JSON. So you can import 2 JSON files and diff them in Powershell.
- avip 8y agoNot merge, but jq can diff, and it can also do a consistent dump (jq -cS) for you. Edit: to clarify, jq -S does deep keys sorting. $ echo '{"z":{"b": "second", "a": "first"}, "x": 4, "y": 7}' | jq -S { "x": 4, "y": 7, "z": { "a": "first", "b": "second" } }
- SiempreViernes 8y agoTighter integration with git is very interesting, but this is sadly just integration with github. I think coupling to github makes sense if you are a building a dev-support service, but for a end user it makes little sense to wed the vcs to a specific website.
- MurrayHill1980 8y agoThe RCloud project covered some of this ground https://cscheid.net/2015/08/17/collaborative-visual-analysis-with-rcloud.html https://cscheid.net/2015/08/17/collaborative-visual-analysis... It takes the view that everything should be saved and versioned. In hindsight it seems obvious that this can overwhelm people with dead ends and scratch work and in general the flat workbook space doesn't provide enough help with organizing results. There are some other ideas mentioned in the conclusion of the RCloud paper.
- dev_dull 8y agoI’m glad I’m not the only one. When I inherited some “production notebooks” (if that’s a thing) I couldn’t believe it was nearly impossible to do basic things such as test and review changes (via version control).
- i-am-charmander 8y ago"Production notebooks" should not be a thing... Unless prefaced by a bold RUN ALL.
- bitL 8y agoYou don't use Jupyter notebooks in production; they are super useful for pitching ideas to clients/bosses and doing some early prototyping. I feel sorry for anyone that has to work with "pure data scientists" that have no clue about software engineering practices...
- cwyers 8y agoIt depends on what you're doing, yeah? In RMarkdown notebooks... yeah, I wouldn't write models in one. But if the focus is on embedding some visualizations and tables into a document, and then refreshing the document to every so often pull in new data, I can see that as a production use for a notebook. TL;DR: Can be useful for reporting, wouldn't use it anywhere else in the pipeline.
- brylie 8y agoFWIW, Netflix uses Jupyter notebooks in production, using nteract UI: https://medium.com/netflix-techblog/notebook-innovation-591ee3221233 https://medium.com/netflix-techblog/notebook-innovation-591e... https://nteract.io/ https://nteract.io/ This approach seems promising, particularly as it facilitates cross-disciplinary collaboration.
- entee 8y agoAt our company if it's in a notebook it's not considered ready for production, it must run as a script before being considered for Eng to take over from DS. It's actually not that hard to write a notebook in such a way that it converts easily to a script. Just check and make sure that your variables/functions/whatever are initialized above the cell(s) they're used in, declare all imports in the top cell, and periodically move cells to fix any inconsistencies with these rules (checking that you didn't break anything of course). I've always said that Data Scientist doesn't mean, "I don't do engineering," good basic eng practice helps make more productive data science and brings it into production more robustly. How do you know your models work well if the code that generated them is inscrutable? I wonder how much of the "3 engineers for 1 data scientist" ratio I hear all the time is due to Data Engineering being assigned the role of cleanup to code that should be better in the first place.
- ssivark 8y agoThe hard part is that introducing a tool like git (which requires you to choose moments to take a snapshot of the file, and then add some commit message) breaks the flow of interactive experimentation that notebooks are so good for. And then we need to find a way to make those commits useful, because the time ordering of commits could be different from the time order in which cells were run! That is what is crucial to making computations reproducible — viewers should be able to replay the history of how a notebook result came to be. (EDIT: Note that this is the case only for stateful computations -- if a notebook interface was used to construct a dataflow graph (like spreadsheets) with values updating live, then this wouldn't be so much of a problem. More fundamentally, it is not at all obvious that thinking of notebook contents as akin to code is the best way to use version control) I wonder whether there is a solution along the lines of auto-committing each cell before it’s executed and the results just after the cell is executed. Otherwise a user has to do too much manual organizing, which is a problem the notebook should ideally solve. When a user is happy with the experiments and the provenance of their results, they should be able to use an interactive rebase to create a cleaner version to share/archive.
- ISL 8y agoI'm not a Jupyter user, but I solve the reproducibility problem with Make. As a project moves from exploration toward production, the entire thing is wrapped into a Makefile that can flow from raw data to publication in a single call to make.
- ontouchstart 8y agoTo have reproducible prototypes, I use Make to wrap the whole workflow in docker. Then I push the code to gist and forget about it. Although GitHub gist doesn't allow binary file, images embedded in .ipynb (JSON), on the other hand works in gist. Here is an example. https://gist.github.com/ontouchstart/854a3c280b81f530d3ae9cb4c1bcbdc4 https://gist.github.com/ontouchstart/854a3c280b81f530d3ae9cb... The notebook generated by nbconvert (see the instruction in the Makefile) is too big to display in GitHub gist Web UI but works fine in nbviewer. https://nbviewer.jupyter.org/gist/ontouchstart/854a3c280b81f530d3ae9cb4c1bcbdc4/raf.nbconvert.ipynb https://nbviewer.jupyter.org/gist/ontouchstart/854a3c280b81f...
- astral303 8y agoRStudio’s Markdown notebooks do not suffer from this and save a separate output file that can be gitignored.
- andrestan 8y agoRMarkdown and Knitr are dramatic improvements in terms of final outputs and VC relative to notebooks. Notebook believers (Satan worshippers, imho) would suggest that notebooks are best for developing in and not primarily made for use as final outputs.
- wodenokoto 8y agoAnd they pay for this on other accounts: No inline rendering of markdown. Opening an .Rmd file is a lottery to see if rendered graphs and tables still exists. Tables render completely differently in editor, HTML and pdf
- gbrown 8y agoFor me, markdown is meant to be readable even when not rendered. I could see how not having persistent graphs and tables might be an issue, but my own philosophy is to start fresh each time - I treat it like a templating language with some convenient rendering features for prototyping, rather than like an IDE. Your last point also has an upside - it's using different engines (Rmarkdown vs. Sweave). I can write whatever HTML or LaTeX code I want, depending on what's appropriate. I wouldn't want to have to make web documents with LaTeX, nor would I want to make PDFs with HTML.
- mike_ivanov 8y ago> No inline rendering of markdown. That's incorrect, take a look here -> https://blog.rstudio.com/2016/10/05/r-notebooks https://blog.rstudio.com/2016/10/05/r-notebooks
- solomatov 8y agoIf you are interested in workbooks which are collaborative and versioned, take a look at http://datalore.io/ http://datalore.io/ Version control is transparent and integrated and it's possible to work with workbooks collaboratively.
- kprybol 8y agoAny chance of this being offered for on-prem install in the future? Looks interesting but cloud only makes it a no go for my team.
- rgardaphe 8y agoWe're building something just like that at qri (https://qri.io https://qri.io) a free and open source dataset version control system. Right now all datasets on qri are public by default, but we're working toward supporting. encryption and private networks.
- solomatov 8y agoWe are seriously considering such a possibility. Do you have any specific requirements for on prem installation?
- kprybol 8y agoBasically just the ability to run on Linux.
- benjaminjackman 8y agoI think they fundamentally json is just the wrong format for these files. Speaking from (ancient and limited) experience I made a little notebook-style interpreter for learning scala back in 2009 or so called scalide. It saved its files ("scalapads") to XML. XML actually worked better in some ways since most of the code could live between the tags unescaped (sans < > &) so it merged / diffed the user code well. The meta-level stuff (cell boundaries etc) needed by the notebook ... not so much. In json the code has to be escaped into strings, and json is really finicky about syntax (e.g. no trailing commas). So it doesn't work well. I never got the chance to redo it, however the solution I was leaning to for my post "I won the lottery, I can work on fun stuff" attempt was to store the meta-code in a version of the host language(s), with some simple syntax that could live comfortably in the comments of various different languages to do things like encode the cell divisions and so on. Basically something like: #notebook[lang=python] #cell[lang=python] def add(x,y): return x + y #endcell //notebook[lang=scala] //cell[lang=python] def add(x: Int, y: Int) = x + y //endcell This I think would be beneficial for a couple of reasons. 1. Better diffing / merging. 2. One click toggle between show source and view as notebook mode, which would really allow this to work in an IDE like vscode pretty seamlessly. The cells become something akin to //#regions in the IDE. But at the end of the day you are still editing a source code file, so you can edit the whole file easily. 3. The keyboard shortcuts for executing and jumping between cells would generally work in raw code mode, so you could just edit there continuously and manually writing out //cell //endcell. Also the executing results could appear in block comments inline in the editor, off to the side, or in a popup above, the code you are editing. 4. The IDEs could uprender the comment-syntax into cells as they gained better support for the paradigm (similar to how they do for code folding / syntax higlighting already). 5. Eventually, perhaps a cross language, metasyntax could be established to make things a bit more concrete than magic comments (get ready for some serious bikeshed painting though!) The closest I have seen anything come in this regard is Quokka however it's not quite all the way there.
- TeMPOraL 8y agoThe closest thing I've seen to what you described would be... Emacs. It actually uses the "metadata in file-specific comments" paradigm. You can put file-local values for Emacs variables in comments at the top or bottom of your file, like described in [0]. Your example could be rewritten as: # -*- notebook-lang: python -*- or // -*- notebook-lang: scala -*- Still, the usual way of using Emacs for "interactive notebooks" is via org-mode, which is a better Markdown with support for (among other things) executing code blocks straight in the org document you're writing. This way, Emacs support all your points 1 to 5, and is generally more powerful than Jupyter or other similar things, but it also means you can kiss any kind of collaboration goodbye. For some weird reason, the more powerful a tool, the less likely it is other people will be using it. -- [0] - https://www.gnu.org/software/emacs/manual/html_node/emacs/Specifying-File-Variables.html#Specifying-File-Variables https://www.gnu.org/software/emacs/manual/html_node/emacs/Sp...
- zimablue 8y agoJupytext linked elsewher eont he thread seems like a step in the right direction. Instead of changing the whole tool, accept that you're always going to be married to github and change the serialization-layer to be source control friendly. Basically, split the input from the output+metadata and flatten it all to text. Then you can source control it fine and if you need use the output+metadata fold them back in.
- Desustorm 8y agoI have used jupytext (https://github.com/mwouts/jupytext https://github.com/mwouts/jupytext) for this and it seems to work great - it outputs a separate .py file which is easily diff-able.
- jimhefferon 8y agoThanks for the note; this looks good.
- autokad 8y agoI'm a huge fan of databricks, its got github sync built into it
- deleted 8y ago[deleted]
- Fomite 8y agoI nicknamed one we used during the Ebola epidemic (tight deadlines, lots of people working, etc.) "The Wall of Madness". There were tons of ## JOHN: DONT RUN PAST HERE, EVERYTHING BROKEN comments.
- deleted 8y ago[deleted]