3 ms·
I'd love to hear some thoughts about keeping a "lab notebook" for ML experiments. I use Jupyter Notebooks when playing around with different ML models, and I fi
by heynk 9y ago
I'd love to hear some thoughts about keeping a "lab notebook" for ML experiments. I use Jupyter Notebooks when playing around with different ML models, and I find that it really helps to document my thought process with notes and comments. It also seems that the ML workflow is very 'experiment' driven. I'm always thinking "Hm, I think if I tweak this hyperparameter this way, or adjust this layer this way, then I'll get a better result because X". Thus, I have a bit of a hypothesis and proposed experiment. I run that model, and see if it improved or not.
Then, I run into an issue where I can either: 1. overwrite the original model with my new hyperparameters/design and re-run and analyze or 2. keep adding to the same notebook "page" with a new hypothesis/test/analysis loop, thus making the notebook pretty large. With number 1, I often want to backtrack and re-reference how a previous experiment went, but I lose that history. With number 2, it seems to get big pretty quickly, and coming back to the same notebook requires more setup, and "searching" the history gets more cumbersome.
Does anyone try using a separate notebook page for each experiment, maybe with a timestamp or "version"? Or is there a better way to do this in a single notebook? I am thinking that something like "chapters" could help me here, and it seems like this extension might help me: https://github.com/minrk/ipython_extensions#table-of-contents https://github.com/minrk/ipython_extensions#table-of-content...
- detaro 9y agoThere is http://sacred.readthedocs.io/en/latest/collected_information.html http://sacred.readthedocs.io/en/latest/collected_information..., but I haven't used it and it doesn't integrate with Jupyter as far as I know
- paulgb 9y agoMy process is to separate numbered "scratch" notebooks in which I do one-off analysis from "pipeline" notebooks/code that process data for use in later analysis/modeling. Pipeline notebooks are made to take input/output locations from environment variables and run with nbconvert. The pipeline code is run by a DAG-like runner (sort of like Make) that is aware of the parameters and what level of the pipeline they are used at so that I can trace back outputs to the parameters and files used at every stage of the model. For stages that are notebooks I can also see the HTML generated by that run of the notebook. Unfortunately I've looked far and wide and there's no real open source version of what I've described, but I'm hoping to open source it once I've worked through the kinks.
- akhilcacharya 9y agoI'm interested in something like this more generically - like a task runner that creates maybe creates a branch for a particular job/training run on a remote server that you can monitor/trace source for.
- sillysaurus3 9y agoDon't overwrite. It's so valuable to go back when you run into an evolutionary dead end.
- BucketSort 9y agoYou could export the results of the models to a file and reference those. With tensorflow you can do this quite easily. I think notebooks are crucial for ML research, especially in teams. It's how we've been sharing our research with each other(i.e. https://github.com/DanburyAI/SG_DLB_2017 https://github.com/DanburyAI/SG_DLB_2017 ). Wolfram calls these computational essays ( http://blog.stephenwolfram.com/2017/11/what-is-a-computational-essay/ http://blog.stephenwolfram.com/2017/11/what-is-a-computation...). He was actually one of the pioneers in these types of notebooks. Distil.pub has a similar ethos. I think the whole notebook way of doing research is central to reducing research debt ( https://distill.pub/2017/research-debt/ https://distill.pub/2017/research-debt/ ). In the past, it is common to throw away the ladder in mathematical research and write nice lean formal papers. This often produces a type of debt. In a team, the ladder is vitally important. Notebooks help retain the ladder.
- RobertoG 9y ago", I often want to backtrack and re-reference how a previous experiment went, but I lose that history" Have you considered using Git or some other control version system? I'm not sure about the practicality in your case, but it seems like something to try.
- Derbasti 9y agoI have the same problem. In time, I found notebooks to be rather useless as lab notebooks for this reason. I now use notebooks only transiently, for running a few experiments, but not as lasting documentation. Instead, I copy a summary of my experiments and findings into a journal, and delete the notebook once it has done its job. The journal is then my true research notebook. It contains experiments, results, code examples, future tasks, ideas, and a daily summary of how I spent my time. Particularly, I use org-journal on Emacs, but any old journal would do.
- westurner 9y agoThese are ASCII-sortable: 0001_Introduction.ipynb 0010_Chapter-1.ipynb ISO8601 w/ UTC is also ASCII sortable. # Jupyter notebooks as lab notebooks ## Disadvantages ### Mutability With a lab notebook, you can cross things out but they're still there. - [ ] ENH: Copy cell and mark as don't execute (or wrap with ```language\n``` and change the cell type to markdown) - [ ] ENH: add a 'Save and {git,} Commit' shortcut CoCalc (was: SageMathCloud) has (somewhat?) complete notebook replay with a time slider; and multi-user collaborative editing. ("Time-travel is a detailed history of all your edits and everything is backed up in consistent snapshots.") ### Timestamps You must add timestamps by hand; i.e. as #comments or markdown cells. - [ ] ENH: add a markdown cell with a timestamp (from a configurable template) (with a keyboard shortcut) ### Project files You must manage the non-.ipynb sources separately. (You can create a new file or folder. You can just drag and drop to upload. You can open a shell tab to `git status diff commit` and `git push`, if the Jupyter/JupyterHub/CoCalc instance has network access to e.g. GitLab or GitHub) ## Advantages ### Reproducibility Executable I/O cells The version_information and/or watermark extensions will inline the software versions that were installed when the notebook was last run Dockerfile for OS config Conda environment.yml (and/or pip requirements.txt and/or pipenv Pipfile) for further software dependencies BinderHub can rebuild a docker image on receipt of a webhook from a got repo, push the built image to a docker image repository, and then host prepared Jupyter instances (with Kubernetes) which contain (and reproducibly archive) all of the preinstalled prerequisites. Diff: `git diff`, `nbdime` ### Publishing You can generate static HTML, HTML slides with RevealJS, interactive HTML slides with RISE, executable source with comments (e.g. a .py file), LaTeX, and PDF with 'Save as' or `jupyter-convert --to`. You can also create slides with nbpresent. MyBinder.org and Azure Notebooks have badges for e.g. a README.md or README.rst which launch a project executably in a docker instance hosted in a cloud. CoCalc and Anaconda Cloud also provide hosted Jupyter Notebook projects. You can template a gradable notebook with nbgrader. GitHub renders .ipynb notebooks as HTML. Nbviewer renders .ipynb notebooks as HTML. There are more than 90 Jupyter Kernels for languages other than Python. https://github.com/quobit/awesome-python-in-education#jupyter https://github.com/quobit/awesome-python-in-education#jupyte...