5 ms·
First class language support in Apache Spark: * Scala * Python * Java * R All these languages are equal, but Scala tends to be more equal than others
by nchammas 10y ago
First class language support in Apache Spark:
* Scala
* Python
* Java
* R
All these languages are equal, but Scala tends to be more equal than others in some areas of the API. I also believe R is mostly restricted to the DataFrame API.
Third-party language support:
* Clojure [0]
To develop on Spark in a new, non-JVM language, you'd need a bridge to Java. That's how PySpark works [1], and I believe R follows a similar pattern.
[0] https://github.com/yieldbot/flambo https://github.com/yieldbot/flambo
[1] https://cwiki.apache.org/confluence/display/SPARK/PySpark+Internals https://cwiki.apache.org/confluence/display/SPARK/PySpark+In...
- baldfat 10y agoI am thinking that R is the reason why this is happening. Now that Microsoft has supported and invested in R (Smart thing on their part) R will get more Spark features. Having better R connections to Spark's API would be amazing.
- _dark_matter_ 10y agoFirst class support for running dplyr pipelines on Spark would be killer, imho.
- pjmlp 10y agoInteresting how the whole R thing happened, about three years ago I wouldn't have any idea what R is about and now see it everywhere, even at customer sites.
- koloron 10y agoHaskell: https://github.com/tweag/sparkle https://github.com/tweag/sparkle
- bunderbunder 10y agoThere's also third-party support for .NET via https://github.com/Microsoft/Mobius https://github.com/Microsoft/Mobius
- cwyers 10y agoThird-party from Spark's perspective, but first-party from C#'s perspective.
- henridf 10y agoThird-party Node.js bindings: https://github.com/henridf/apache-spark-node https://github.com/henridf/apache-spark-node https://github.com/EclairJS/eclairjs-node https://github.com/EclairJS/eclairjs-node
- jbooth 10y agoMore like 3 tiers of support: Scala -- native API Java -- java wrappers for Scala API, slightly clunky but same implementation Python+R -- janky process involving forking a process and feeding strings back and forth between JVM and python/R via pipes
- nchammas 10y agoYou're right about Python and R having to pass data back and forth to the JVM for certain operations, but also keep in mind that native code still runs in the native interpreter. That means you have access to the full ecosystem of the native language. For example, if I want to convert an RDD of JSON strings into Python dictionaries: import json rdd_dict = rdd.map(lambda x: json.loads(x)) Same goes for any external Python libraries I install on the cluster and want to use in my Spark job. You can even run your Python code on PyPy [4]! For me, working in Python generally feels like a first class experience on Spark. There are areas -- like GraphX [0], certain niche features [1] -- where Scala is definitely easier to work with, but with time that is becoming less [2] and less [3] true thanks to the DataFrame API. [0] https://spark.apache.org/graphx/ https://spark.apache.org/graphx/ [1] http://stackoverflow.com/q/23995040/877069 http://stackoverflow.com/q/23995040/877069 [2] https://github.com/graphframes/graphframes https://github.com/graphframes/graphframes [3] http://stackoverflow.com/a/37150604/877069 http://stackoverflow.com/a/37150604/877069 [4] https://github.com/apache/spark/pull/2144 https://github.com/apache/spark/pull/2144
- twic 10y agoHas anyone tried running Jython on top of the native API? Would that make any sense?