10 ms·
Voila – From notebooks to standalone web applications and dashboards
- adwww 5y agoCan't wait to be asked to deploy 5gb Anaconda based blog articles using this.
- mg 5y agoIt seems this is doing round trips to the server to do the calculations? I often think that with the current state of Javascript in the browser, we could build an awesome, super fast, Jupyter style notebook software that runs completely in the browser. With the modules implemented as native Javascript modules which are dynamically loaded. Is anybody working on this? I have built a rough version of this idea for myself and been using it for my own statistic needs for a few months now. It is far from being polished/flexible enough to be useful as a general purpose notebook though.
- buro9 5y agoOnly you know your data volumes but for the scenarios in which I've reached for Jupyter they've often involved very large amounts of data and calculating on the server is what I wanted and needed. Agreed that for some things, it would be great to be able to explicitly offload to the browser.
- mg 5y agoYes, my dataset is tiny. I mainly use the JS notebook to analyze my selftracking log [1] which is about 10k lines of data at the moment. I have not yet tried to load a lot of data into it. Would be interesting to see when the load time starts to outweight the benefits of instant calculations. Maybe at something like 10 million datasets? Hard to say. 1: https://www.gibney.org/a_syntax_for_self-tracking https://www.gibney.org/a_syntax_for_self-tracking
- eldudebrother 5y agoHave you seen https://nomie.app https://nomie.app? You can use a CouchDB backend to store the logs as semi-structured data.
- dallathee 5y agoCheckout https://github.com/jupyterlite/jupyterlite https://github.com/jupyterlite/jupyterlite, WASM'd Jupyter in the browser with Pyodide
- dgb23 5y agoNice, my instinct reading the above was: The (future) target of this could be WASM instead of JS. It's a very promising technology IMO. Not necessarily because it is faster (which is only 0.3x-2x according to my findings) but because it is a nice, simple compilation target.
- deleted 5y ago[deleted]
- gibbonsrcool 5y agoI believe Google Colab provides what you described.
- rogue7 5y agoOn the javascript side there is https://observablehq.com https://observablehq.com
- BiteCode_dev 5y agoYou can, with pyodine: https://github.com/pyodide/pyodide https://github.com/pyodide/pyodide Here is an example instance: https://notebook.basthon.fr/ https://notebook.basthon.fr/ The thing is, creating a whole stack in pure JS would be very hard, since the current scientific stacks uses a lot of fortran, assembly and C with python to bind them all. Or julia. It's millions of man hours we are talking about. So compiling Python into WASM is probably the best deal for such an app. For the regular web, it would be a deal breaker: you don't want to load 15 mo of runtime before being able to interact with a web page. But for such a scientific app, it's not a problem. Besides, were you to write it entirely in JS, the size would be huge as well.
- goatlover 5y agoObservable notebooks exist. The issue for JS replacing Jupyter is that Python already has a large number of scientific libraries where the performant ones are based on wrapping C++ or Fortran code. You can handle more data and processing on the server. Also, because Jupyter gives you shell access and other languages like R can make use of it. Language-wise, the other issue JS has is that it lacks operator overloading, which makes array handling a lot nicer. Python has magic methods for doing that. Julia, R and Matlab have that built into their languages. Julia also has it's own version of reactive notebooks similar to Observable. And then you have a lot of scientists who already know and use Python, R or Julia. I can't really imagine a statistician who's well versed in R finding much value in JS. Javascript just isn't made for complex statistics the way R is. JS now feels like Java back at the turn of the millennium, when people were thinking Java could just be used for everything, and it sort of was. Or any language could run on the JVM (the web has largely replaced that idea). But there's a reason for the various different programming languages. Some are just better at doing certain things. And JS is not a scientific computing language.
- protoduction 5y agoCheck out Starboard (https://starboard.gg https://starboard.gg, I'm the creator). I think it's exactly what you're looking for (you can use plain dom, html, css, javascript, python).
- matsemann 5y agoMy usecase for jupyter notebooks often is to use the server to calculate stuff, though. I rent some beefy 16 cpu, 128gb ram, tesla p100 machine so that my laptop won't melt.
- marianoguerra 5y agoI'm working on a version of this: https://instadeq.com/ https://instadeq.com/
- pavlovskyi 5y agoI really like Jupyter notebooks to build a simple concept and then move to .py files. But what I observe, especially at the entry level or junior level jobs in data science is that people spend huge amount of its work on jupyter, which did not focus on how to plan flow properly. What I meant is that there is very short path from usefullness to overkill.
- BiteCode_dev 5y agoThat's because they are not programmers. One should not expect them to be experts in 2 fields. They do their work with the tool provided, and if we want a better output, we need to provide better tooling or accept what comes out. I want my physicists to spend their mental effort on physics, not on software architecture.
- pavlovskyi 5y agoI agree that they are not programmers but in my opinion it cannot become an explanation to write a code which become non repeatable, especially if they have big impact on how the flow will look on production environment.
- BiteCode_dev 5y agoScientists are not meant to create code that ends up on production. Once they have a working concept, they should team up with a programmer to make it live. Again, if it's not possible, then you accept the imperfection of the result, or provide better tooling. There is no blame to put on them whatsoever.
- throwaway210222 5y agoI disagree, many, many scientists hire a professional statistician to do the stats for their papers. Similarly, they should hire experienced, qualified software engineers to write/check the software in their papers. They don't because 'everyone can code - its just logic'.
- dennisy 5y agoThis is cool, looks very similar to Streamlit. https://streamlit.io/ https://streamlit.io/
- deleted 5y ago[deleted]
- fxtentacle 5y agoTried their example link: "Problem: package xeus-cling-0.12.0-h5a79028_0 requires xtl >=0.7.0,<0.8.0a0, but none of the providers can be installed" sigh I don't know when it started, but it seems to be a recent trend to add dependencies for anything and to package everything on demand. It's probably for security or something. But I do miss the days when people would link a static binary that "just works" even without internet and that'll keep working a week later, because it includes all of its dependencies as opposed to downloading and updating 500 packages on-demand.
- raziel2p 5y agoPython projects around jupyter, pandas etc. seem especially bad at making reproducible environments. They lock to versions that don't work very well, only work with specific versions of Python (without documenting it)...
- SylvainCorlay 5y agoThis has nothing to do with Python. The faulty package was xeus-cling, a C++ kernel for Jupyter.
- qayxc 5y agoAmen to that. That's why tools like Anaconda and Docker were created and now even a simple utility can use gigabytes of disk space...
- matsemann 5y agoPython is especially bad at getting things up and running. Often ends up with a mix of conda and pip stuff, different python installs through pyenv or similar, some global installed version of cuda, jupyterlab etc., wheels not created for your particular platform and version of python making you have to install a whole c++ toolchain etc. It's bonkers.
- SylvainCorlay 5y agoVoilà author here. This is fixed. I was requiring a wrong version of cling in the example. FYI, xeus-cling is a Jupyter kernel for the C++ programming language. https://github.com/jupyter-xeus/xeus-cling https://github.com/jupyter-xeus/xeus-cling
- vladev 5y agoNote that there's also streamlit [1]. It uses regular python files, rather than notebooks, so they can be easily version controlled. And it has more UI tools. [1]: https://streamlit.io/ https://streamlit.io/
- pavlovskyi 5y agoIn my daily routine we are using streamlit and it is pretty decent, mainly because you do not have to care much about backend. And, what was mention by you it has impressive amount of UI tools and relatively active community.
- vatican_banker 5y agoThis is great. I've been looking for the equivalent of RShiny in the python world and never heard of streamlit before
- IanCal 5y agoExceptionally strong recommendation for streamlit from me. I can create a GUI for a tool that looks nice faster than I can make a CLI. I've built useful production systems (ok, sure, for internal use) in literally minutes. You're a bit limited in what kinds of apps you can make but the tradeoffs it makes here means that it's astoundingly easy to make a wide range of very useful tools.
- byteface 5y agoI wasn't aware of this so thanks for sharing. I've been setting up a repo that utilises github actions to build exe/app files as noted in this guys blog... https://data-dive.com/multi-os-deployment-in-cloud-using-pyinstaller-and-github-actions https://data-dive.com/multi-os-deployment-in-cloud-using-pyi... It uses pyinstaller to build and even pushes the build as a zip into your release page on github and appears to be working quite well.
- d4rkp4ttern 5y agoStreamlit is great for demos but not for building a product.
- d_rc 5y agoDeepnote recently introduced interactivity when you publish your notebook as well: https://docs.deepnote.com/collaboration/publishing-a-notebook#interactivity-of-published-projects https://docs.deepnote.com/collaboration/publishing-a-noteboo...
- rogue7 5y agoI recently used this to do a POC at my day job. I was able to demo a machine learning tool quite smoothly to executives. Later it was implemented in production with a regular stack (Flask + Vue). Voila is really empowering for e.g. data scientists that are comfortable in a jupyter environment but aren't js wizards. Running locally, I just love the reactivity it provides: you don't worry about sync between front-end and back-end, everything is propagated through websockets I believe (Jupyter is Tornado-based). However, for production you might want to use another tool, since it (currently) executes every session in isolation, so every time an user connects it re-runs everything from scratch. Moreover, the round-trips to the server can be slow if you are e.g. in a different continent so this degrades the UX. Here is an example of a small ML app I built with Voilà (this will probably crash due to HN hug of death™), and JAX on the backend: http://grad-descent.herokuapp.com/ http://grad-descent.herokuapp.com/
- nxpnsv 5y agoWould it be crazy to add Viola: to the title of this sub?
- abdullahkhalids 5y agoI have a question about hosting costs - not a SW. Suppose I write some educational Jupyter notebooks, which are not particularly resource intensive, say 100 seconds of compute time per notebook. I host them on some cloud server, using something like OP, and get a 1000 people to learn from it. Maybe they end up using say, 1000 people x 5 notebooks x 100 seconds/run x 20 runs of each notebook = 10 million seconds of compute time. How much would such a server cost to host, where "many" of these people are working on the notebooks together? Just need a rough estimate.
- qayxc 5y agoThis could cost anywhere from nothing (e.g. free) to 3-figures (in USD) depending on the specifics. How many users are accessing the notebooks concurrently (e.g. all 1000 or only a dozen at a given time)? Is there any downtime, i.e. do the users come from the same time zone, so that app can have inactive hours (say it's OK to be unreachable during the night)? Depending on the specifics, free hosting may be available (e.g. via Heroku, Google Colab, AWS Free Tier etc.). As far as paid offers go, this is way too unspecific to be answered in a meaningful way. The answer depends on the actual resource requirements (RAM, storage, data transfer, CPU cores), estimated usage patterns (concurrent users), and your location. TBH, if no commercial interest is involved, just hosting the notebooks on Github or making them accessible via Google Colab would be the easiest option.
- abdullahkhalids 5y agoWe run a non-profit minimal-budget workshop where currently hundreds of people work together at the same time, but they run code on their own computer. But making people install Jupyter and other python packages on their computer is difficult. So we are exploring the possibility of the hosting the notebooks ourselves. We don't want options like Google Colab, because we want the experience to be tightly integrated (there are also issues around GDPR). So we want to run our own server. If we can run a ten-day workshop of 500-1000 people, where most people work everyday at the same time in a 5 hour slot, and keep costs under 50-100 usd, we would make the switch. But I understand that it is difficult to make estimates without trying how much resource usage there actually is.
- lysecret 5y agoI love jupyter notebooks but I think the way they have be to be used is in a "throw away" fashion. E.g. use it to explore data, develop some algorithm, then put it to PY files. It is there to develop something which is worth versioning. I think in this way this seems like a cool addition: To evaluate the worthiness of the algo you might need to show it some people, this is where this comes into play.
- rhizome31 5y agoLast time I checked, Voilà was very slow for anything but the simplest dashboards. This may or may not be a problem depending on the context as instant page load isn't always needed but it's a an aspect to take into account. So my friendly advice is to do your benchmarks with real-world use cases before investing time in this solution.
- throwaway2568 5y agoVoila is quite nice but I find panel [1] is the best option these days. It has plenty of widgets, including those from Voila which can be used as a backend, a few different ways of defining callbacks and has added nice features lately like autoreload if you are using scripts instead of notebooks [2] and new fast HTML elements so it's super easy to define custom widgets straight from the web. They have a discussion page comparing the project to the standard alternatives (dash, streamlit, voila etc) [3]. The docs could do with improving but their discourse is very active [4]. [1] https://panel.holoviz.org/getting_started/index.html https://panel.holoviz.org/getting_started/index.html [2] https://github.com/holoviz/panel/pull/1983 https://github.com/holoviz/panel/pull/1983 [3] https://panel.holoviz.org/about/comparisons.html https://panel.holoviz.org/about/comparisons.html [4] https://discourse.holoviz.org/c/panel/5 https://discourse.holoviz.org/c/panel/5
- lmeyerov 5y agoWe find that design decisions like forcing data scientists to code UI callbacks are big limiters to adoption, which is intuitive as that's pretty close to telling them to write JavaScript in Python. Same thing for styling ("CSS in Python".) They can in theory, but rather spend time on other things. So far, the only low-code PyData framework we saw that avoids most "JS in Python" is StreamLit. However, even there, it is still awkward in practice, so we still see limited adoption by folks who are fine with notebooks, so rarely goes beyond a champion. So there is room to grow.
- throwaway2568 5y agoI would still recommend panel, it is perfectly straightforward to make a clean UI in pure python and the "depends" approach to interactivity works just by adding decorators to functions. You can prototype in either notebooks or scripts, particularly with the auto reload feature which I believe is inspired by streamlit. Here is an example https://panel.holoviz.org/gallery/layout/distribution_tabs.html#layout-gallery-distribution-tabs https://panel.holoviz.org/gallery/layout/distribution_tabs.h...
- 5y ago
- otobrglez 5y agoOP here. I'm sorry that I didn't put the name "Voilà" in the title. I was so very excited that I've mistyped myself.