5 ms·
The other benefit of the JVM is the multiple languages that run on it - did you try Clojure or JRuby already? Or do you mean Java as in JVM and you don't want t
by dm3 11y ago
The other benefit of the JVM is the multiple languages that run on it - did you try Clojure or JRuby already? Or do you mean Java as in JVM and you don't want that in your stack for some reason?
- vegabook 11y agoI am not a fan of any compatibility layer technology as I am a bit of a bare-metal purist, and I believe Linux does this already. The original post waxes lyrical about the Unix philosophy, but part of the Unix philosophy is supreme efficiency and this is violated by the JVM. I know I know - performance is comparable but then coding is a bit like art - you have to love what you are doing. As for using other languages - I played with Python on Storm but I very quickly found that I had to basically use Java because the entire community and documentation was java-centric (last I checked - a few months ago). I may be being obtuse here - in fact I know I am - but my point is that there is a large community out there which is not Java friendly (logically or not) and so I lament the fact that all the advancement in this new, very exciting, field, is JVM-based. My dislike for Java was sealed by the Bloomberg terminal API, some of which's basic functionality is 5 to 6 dots deep. Seriously I had 80+ character function calls. This.that.that.this.this.finally()! Perhaps we'll just have to kick in, leave our biases behind and start running with the JVM. FWIW my use case is streaming fixed income (bond pricing) data analysis. Python/Numpy is running out of steam fast for me and we're using quite a bit of C. I basically come from the scientific computing set. We're applying advanced statistical analysis to pricing (not HFT - we're operating in the 5 minute thru 1-week horizon - not 5ns).
- pixelmonkey 11y agoIf you're interested in Python + Storm, we are making it happen via our open source project streamparse. See the project on Github[1] and also my PyCon US 2015 talk[2]. Your point about documentation is well-taken. We are trying to document streamparse + Storm usage from a Pythonista standpoint via our online documentation[3], e.g. here is our detailed Python API documentation[4]. [1]: https://github.com/Parsely/streamparse https://github.com/Parsely/streamparse [2]: https://www.youtube.com/watch?v=ja4Qj9-l6WQ https://www.youtube.com/watch?v=ja4Qj9-l6WQ [3]: http://streamparse.readthedocs.org/en/stable/ http://streamparse.readthedocs.org/en/stable/ [4]: http://streamparse.readthedocs.org/en/stable/api.html http://streamparse.readthedocs.org/en/stable/api.html
- teacup50 11y agoIn one breath "I am a bit of a bare-metal purist" and in the next "As for using other languages - I played with Python on Storm". So which one is it? :-)
- srean 11y agoI am not the OP but it is not necessarily a contradiction. The underlying goal in Pythonic data-science is that one pushes as much of the computation into pre-packaged loops that have already been written in a close to the metal language (C, C++, Fortran). Well, that's as far as the intent goes, practice deviates from it by degrees. This does not work quite as well in Java (or JVM) because JNI is just supremely god-awful. JVM semantics are overly strict, this over-specification kills optimization opportunities. Is heavy on memory, I would rather use the memory for loading more data than fill it with overhead. Finally there is this OOP culture that gets in the way. That said, JVM is one of the most mature, and well engineered VMs out there (Java is another story), but not very well suited for number crunching, because there is more to number crunching than calling BLAS APIs.
- vegabook 11y agolast time I checked it was Java Scala or Python. None is bare metal. We use Python as glue. Any heavy lifting is done using Numpy ecosystem or C.
- infinite8s 11y agoWhat do you find missing in the python/numpy space? Inability to handle analysis over streaming data?
- vegabook 11y agoInability to scale easily to clusters, unless the problem is embarrassingly parallel (ours isn't always), and no equivalent to the conceptual beauty and efficiency of distributed directed acyclic graphs that Storm amongst others uses. We could hack Python up to do this, but if the problem became TRULY huge, hundreds or thousands of nodes, we'd probably end up spending more time building a processing infrastructure, and repaying technical debt, than doing the actual algorithms, while Spark / Storm etc already have massive scaling ability built in. Still love Python and especially the Numpy ecosystem though and we stream thousands of ticks per second through it no problem, with a bit of help from multiprocessing, MKL libs and Numbapro. But as we get towards tens of thousands we're hitting the buffers, and most importanly, problem complexity is rising so the DAG approach is looking very attractive.
- infinite8s 11y agoI'm looking to build something like Naiad on top of python, disco and numba. Mind if I pick your brain on the kind of tooling you'd like to see around big data python? My email is in my profile.
- parasubvert 11y agoClojure is a fascinating language. Rich Hickey gives a lot of his rationale for why he picked the JVM for it [1], which boils down to VMs being the platforms of future, not OSes like Linux. [1] http://clojure.org/rationale http://clojure.org/rationale