4 ms·
Mean of me to say, but you're just better off using Jupyter as a local notebook sandbox, for one, the relevant development Docker image does bundle Spark[1], ma
by forgetfulness 2y ago
Mean of me to say, but you're just better off using Jupyter as a local notebook sandbox, for one, the relevant development Docker image does bundle Spark[1], making it more convenient to fire up, and more importantly, it's used way more than Zeppelin, as orgs not using Jupyter are probably using Databricks notebooks instead, and it's split between those two.
Zeppelin does make it easier to run Scala Spark, I find, but Scala Spark usage has declined rapidly.
1. https://hub.docker.com/r/jupyter/pyspark-notebook https://hub.docker.com/r/jupyter/pyspark-notebook
- Moto7451 2y agoI worked at a non Databricks using organization and sharing Jupyter notebooks hosted on Kubernetes ended up being such a difficult endeavor that an ops team was hired for it. I don’t think we really got positive ROI on this but some people felt really cool (we had too much of a bias towards self hosting). We did need some sort of sharing and collaboration mechanism and at least for that job this checks a lot of the boxes, especially since our Spark SQL jobs couldn’t be visualized in Jupyter while I worked there.
- appplication 2y agoWe have found Zeppelin to largely be frustrating, bug riddled, and overly restrictive for normal notebook use cases. I agree that Jupyter for PySpark makes more sense in almost every use case. We made the switch as an org about 2 years ago and haven’t looked back. Jupyter has its own issues but does feel much usable by just about every metric.
- deleted 2y ago[deleted]