5 ms·
How we made Jupyter notebooks load faster
- dkga 2y ago[flagged]
- deleted 2y ago[deleted]
- davidgomes 2y agoI was on this team at SingleStore and I can vouch for how hard this team worked on this project. I just opened a couple notebooks in production and they loaded *instantly*, so kudos to the team for seeing this project through. (If you're not familiar with SingleStore's Jupyter Notebooks, they're similar to Databricks Notebooks[1] or Azure Synapse Notebooks[2]). [1]: https://docs.databricks.com/en/notebooks/index.html https://docs.databricks.com/en/notebooks/index.html [2]: https://learn.microsoft.com/en-us/azure/synapse-analytics/spark/apache-spark-development-using-notebooks https://learn.microsoft.com/en-us/azure/synapse-analytics/sp...
- yunohn 2y agoI appreciate the write-up, it's very insightful. But I was quite concerned about the amount of response mocking being used as a solution. Did the team ever consider instead making upstream changes to jupyter-lab (which is FOSS), so that these requests are either deferred or can be configured to not run? That seems like it would benefit everyone, including your company - and might even uncover further optimizations.
- tgrine 2y agoHi, I'm one of the co-authors of the blog post. You raise a valid point and, in large part, I agree with you. To offer some explanation, there are essentially 2 main reasons for us taking this approach: 1. Contributing to an open source project, especially one the size and complexity of jupyter-lab, is generally going to be a slower process than finding a solution "in house". Improving the load times became a priority once we realized our notebooks were bringing value to users, and we wanted to deliver a better experience as soon as possible; 2. It's not always apparent if the changes you are looking for from an open sourced project are useful to a more general audience or if it's very specific to the way you are using the project. A lot of the requests that were mocked could only be so because we either didn't use them in our implementation (for example, users and workspaces) or because we know the response won't change (for example, some extension settings which we don't allow users to change). Is this a common situation for others or is it a niche circumstance of how we are using jupyter-lab? If it's not common, then adding these options in jupyter-lab itself could just increase its complexity while not bringing that much benefit (not saying this is necessarily the case here); To your point though, a good example of this is the checkpoints feature. There is an open issue requesting the option to disable checkpoints[1] as it is not always useful for people. We had the same issue, since we are not using checkpoints, but the requests were always being made. Ultimately, we just mocked the checkpoints requests, but it's probably the case that making the changes to jupyter-lab to disable this would benefit us and other people as well. [1]: https://github.com/jupyterlab/jupyterlab/issues/11826 https://github.com/jupyterlab/jupyterlab/issues/11826
- yunohn 2y agoThanks for the clarification, that's fair enough. But I hope you do look into upstreaming, maybe even just by opening an Issue to guage community interest. Like the pre-existing checkpoints Issue you linked to, opening some for your functionality might show others what is possible.
- paddy_m 2y agoImpressive work. I love jupyter, but it's a bear to work on. What mix of JS packages do you see? Could that be built into one uber package?
- lneves12 2y agoWe are actually bundling everything inside one big main.js file, compared to jupyter-lab app that loads each extension from a different file using webpack federated modules. We did some benchmarking and it was actually faster than having one file per extension. There is definitely still some room for improvement here, but we have some other places we would like to optimize first, like, optimising the fetching of the notebooks content. You can take a look at the notebooks entrypoint network request: https://portal-notebooks.singlestore.com/ https://portal-notebooks.singlestore.com/
- paddy_m 2y agoI take it that's supposed to be the pre-loading page to just look at the requests, not a full working UI? After cursory googling I couldn't find one, do you have a public notebook gallery?
- lneves12 2y agoYeah, sorry should have made that more clear. This doesn’t really load anything, it’s just the entry point for our iframe. To try our notebooks you can create an account at portal.singlestore.com (we have a gallery of notebooks there)
- canucker2016 2y agoI don't see a Content-Encoding header on the response for the JS and HTML files, which suggests the 11.5MB JS and the HTML files aren't compressed. Not much of a worry on the tiny HTML file, but the 11.5MB JS file should compress to a much smaller file on the wire.
- lneves12 2y ago
- spiralk 2y agoI dislike how Jupyter notebooks have become normalized. Yes, the interactive execution and visuals are nice for more academic workflows where the priority is quick results over code organization. However, when it comes to sharing code with others for the sake of doing reproducible science, jupyter notebooks cause more trouble than they are worth. Using cell based execution with python is so elegant with '# %%' lines in regular .py files (though it requires using VSCode or fiddling with vim plugins which not all scientists want to do I suppose). No .ipynb is necessary, .py files can be version controlled and shared like normal code while sill retaining the ability to use interactively, cell by cell. Its much easier to organize .py files into a proper python module, and then share and collaborate with others. Instead, groups will collect jumbles of slightly different versions of the same jupyter notebooks that progressively become more complex and less manageable over time. It's not a hypothetical unfortunately, I've seen this happen at major university labs. I'm not blaming anyone because I understand -- the funding is there to do science and not rewrite code to build convenient software libraries. Yet, I can't help but wish jupyter notebooks could be removed from academic workflows.
- luplex 2y agoIn the end, usability wins. In a Jupyter notebook, you have a much better idea of state between cells, you can iterate much faster, you can write documentation in readable markdown. Often, Jupiter notebooks are more like interactive markdown than they are like python scripts.
- dxbydt 2y ago> In a Jupyter notebook... > Often, Jupiter notebooks... Everytime I search my Slack, I have to run two searches because DS can't agree on how to spell the damn thing.
- Twirrim 2y agoI use jupyter notebooks at work, not so much for academic stuff, but often to help build and show a narrative to folks, including executives (where I have any even remotely technical leadership). It's great for narrative stuff, especially being able to emit PDFs and what not. I've been in a number of meetings where I've got the code up in Jupyter, sharing the screen, and leadership want us to tweak numbers and see the consequences. It's great for exploring code and data too, especially situations where I'm really trying to feel my way towards a solution. I get to merrily intermingle rich text narrative and code so I explain how I got to where I got to and can walk people through it (I did that with some experimenting with an SMT solver several months ago, meant that people that had no experience with an SMT solver could understand the model I built). I'd never use it to share code though. If we get to that stage, it's time to export from jupyter (which it natively supports), and then tidy up the code and productionise it. There's no way jupyter should be the deployed thing.
- pizza 2y agoWhile we're on the topic of jupyter enhancements, would really love to be able to pop a cell off the pending execution stack if I realized running it would be a mistake and still have time before it gets there.. :^)
- sa-code 2y agoAnd also preserving cell output if you happen to reload the page