7 ms·
Announcing Apache Spark 1.4
- minimaxir 11y agoI'm excited about SparkR, even though R is shunned in the field of big data. Between that and dplyr (which inspired the SparkR syntax) for data manipulation and sanitation, it should be much easier to write sane, reproducible code and visualizations for big data analysis. (the Python/Scala tutorials for Spark gave me a headache) SparkR appears to have strong integration into Rstudio, which is big news: http://blog.rstudio.org/2015/05/28/sparkr-preview-by-vincent-warmerdam/ http://blog.rstudio.org/2015/05/28/sparkr-preview-by-vincent...
- IndianAstronaut 11y agoIt will be interesting to see how all the R libraries play with Spark. There are bound to be some hiccups there.
- minimaxir 11y agoMy interpretation is that it will convert DataFrames to normal data.frames when necessary. Unfortunately, this removes the performance efficiency of Spark. Since currently SparkR only supports aggregation, it limits the usability of SparkR slightly. Future versions will apparently have MLib support which should alleviate that.
- mwexler 11y agoNot sure that R is _shunned_ in big data, as much as there are better solutions once you get to a certain level of big.
- eranation 11y agoR on Spark is great, but the biggest issue in my view is R's runtime licensing, isn't it GPL? Am I worried for nothing?
- zmmmmm 11y agoI too have been mystified by R's licensing. I actually don't see how anyone can ship a commercial product using R in its current form. At very least you're in a legal gray area, at worst you are involuntarily open sourcing your product. Not that there's anything wrong with open sourcing a product, but I think there's an enormous potential issue that could foul up a lot of people down the track. The best discussion I have seen about this pretty much ends up with uninformed speculation. For now, I take the policy of "explore and prototype in R, build the real system in something else". Fortunately the flaws and limitations of R as a language make this a sensible choice for a host of other reasons as well.
- threeseed 11y agoR is absolutely not shunned in big data. It is very popular. There is a reason Microsoft acquired Revolution Analytics.
- fleeno 11y agoAs someone who doesn't know what Apache Spark is, this article reads like it could have been randomly generated.
- sixdimensional 11y agoApache Spark is a general purpose distributed data processing and caching engine. It is an evolution of MapReduce concepts into more general "directed acylic graph" processing, which is very flexible for defining and executing data processing work on a distributed cluster. It's got some similarities to PrestoDb, Apache Drill and or Apache Storm (although not quite the same). It also has some nice data mining libraries, a library for handling streaming data, some connectivity to external data sources and a library for accessing data stored in its generic "data frames" via SQL. "Data frames" are just an abstraction for a dataset, but they are distributed, and in-memory and/or persistent. Personally, I like to think of as an engine for data analysis/processing and queries, but different in that it is not really a "database" like you would traditionally consider. It's almost like if you took the SQL data processing engine out of your database and made it really flexible. Edit: Also, all the functionality of Apache Spark is programmatically accessible in Java, Scala and Python, or through SQL with their Hive/thrift interface.
- mdellabitta 11y agoThere's an About page on that same site, available in the top navigation: http://databricks.com/spark/about http://databricks.com/spark/about
- DannoHung 11y agoDoes anyone know if there's a guide to integrating Spark between a realtime write only database and a historical database? I've looked into using Spark Streaming, but I can't work out how you could seamlessly transition data from a streaming batch to the historical db in a reasonably tight time period. I'd be willing to pay for training if it came to it, but I don't think I'm using the right search terms.
- sixdimensional 11y agoMay I ask, why do you want to integrate Spark in the middle of the two? I am seeing Spark used more for distributed processing/caching data rather than being a conduit for data movement from one system to another. You have a realtime write only database and you want to update a historical database from that write only database? Or do you just want to join data across the two sources on the fly? Those are two pretty different use cases. Based on what you're asking, you might find these two articles interesting: - http://blog.confluent.io/2015/03/04/turning-the-database-inside-out-with-apache-samza/ http://blog.confluent.io/2015/03/04/turning-the-database-ins... - http://lambda-architecture.net/ http://lambda-architecture.net/
- DannoHung 11y agoWell, maybe I'm totally off on this, but it's more that I'd like to be able to run analytics which include real-time data without having any notable pauses. I'm willing to look at anything in terms of getting the data from the real time capture into the historical database as long as the spark queries "just work". Sorry, I think maybe "integrating between" was the wrong way to phrase it. On the other hand, I mean there's clean up and preprocessing I want to do on data that goes into the historical dataset, so hey, why not do that clean up/processing with Spark? I've seen Lambda Architecture before, but it seems like it's kinda gone dark and unless I just totally overlooked it, I don't think there was a "Hey, this is the way to do it guys!"
- threeseed 11y agoNot sure if you have used it but Spark is exceptionally good at data movement. In fact that is what a lot of people initially started using it for (as a replacement for Hive/Pig). You can write SQL against HCatalog tables, do some transformation work then write the results out to a different system. We have hundreds of jobs that do just this.
- eranation 11y agoAnyone who wants to pick up Spark basics - Berkeley (Spark was developed at Berkeley's AMPLab) in collaboration with DataBricks (Commercial company started by Spark creators) just started a free MOOC on edx: https://www.edx.org/course/introduction-big-data-apache-spark-uc-berkeleyx-cs100-1x https://www.edx.org/course/introduction-big-data-apache-spar... (If you wonder what is Spark, in a very unofficial nutshell - it is a computation / big data / analytics / machine learning / graph processing engine on top of Hadoop that usually performs much better and has arguably a much easier API in Python, Scala, Java and now R) It has more than 5000 students so far, and the Professor seems to answer every single Piazza question (a popular student / teacher message board). So far it looks really good (It started a week ago, so you can still catch up, 2nd lab is due only Friday 6/12 EOD, but you have 3 days "grace" period... and there is not too much to catch up) I use Spark for work (Scala API) and still learned one or two new things. It uses the PySpark API so no need to learn Scala. All homework labs are done in a iPython notebook. Very high quality so far IMHO. It is followed by a more advanced spark course (Scalable Machine Learning) also by Berkeley & Databricks. https://www.edx.org/course/scalable-machine-learning-uc-berkeleyx-cs190-1x https://www.edx.org/course/scalable-machine-learning-uc-berk... (not affiliated with edx, Berkeley or databricks, just thought it's a good place for a PSA to those interested) The Spark originating academic paper by Matei Zaharia (Creator of Spark) got him a PHd dissertation award in 2014 by the ACM (http://www.acm.org/press-room/news-releases/2015/dissertation-award-14/ http://www.acm.org/press-room/news-releases/2015/dissertatio...) Spark also set a new record in large scale sorting (Beating Hadoop by far): https://databricks.com/blog/2014/11/05/spark-officially-sets-a-new-record-in-large-scale-sorting.html https://databricks.com/blog/2014/11/05/spark-officially-sets... * EDIT: typo in "Berkeley", thanks gboss for noticing :)
- deleted 11y ago[deleted]
- spacko 11y ago> It is followed by a more advanced spark course (Scalable Machine Learning) Is it really more advanced regarding Spark? The requirements state explicitely that no prior Spark knowledge is required.
- chiachun 11y agoThe release notes: https://spark.apache.org/releases/spark-release-1-4-0.html https://spark.apache.org/releases/spark-release-1-4-0.html Another major change is that it supports Python 3 now. https://issues.apache.org/jira/browse/SPARK-4897 https://issues.apache.org/jira/browse/SPARK-4897
- choppaface 11y agoThey've integrated Tungsten / native sorting into shuffle and observed some decent speedups: * https://issues.apache.org/jira/browse/SPARK-7081 https://issues.apache.org/jira/browse/SPARK-7081 * https://github.com/apache/spark/pull/5868#issuecomment-101837095 https://github.com/apache/spark/pull/5868#issuecomment-10183... However, I guess reduceByKey (and friends) don't benefit yet. Their SGD implementation still uses TreeAggregate ( https://github.com/apache/spark/blob/e3e9c70384028cc0c322ccea14f19d3b6d6b39eb/mllib/src/main/scala/org/apache/spark/mllib/optimization/GradientDescent.scala#L189 https://github.com/apache/spark/blob/e3e9c70384028cc0c322cce... ) so I wonder when they're planning to add some of the "Parameter Server" stuff (e.g. perhaps butterfly mixing or Kylix http://www.cs.berkeley.edu/~jfc/papers/14/Kylix.pdf http://www.cs.berkeley.edu/~jfc/papers/14/Kylix.pdf )
- Tepix 11y agoToo bad the website is so hard to read. Time for that site to join contrastrebellion.com
- krat0sprakhar 11y agoI, somehow, always keep getting confused between Spark and Storm! Can someone explain the difference between the two (usecases etc.) as if explaining to a five year-old? Thanks!
- nl 11y agoStorm = Streaming data processing, written in Clojure and previously used at Twitter until they replaced it with Heron. Spark = Streaming (technically micro-batch) and batch data processing, written in Scala and used very widely.
- lazzlazzlazz 11y agoIs support for User Defined Aggregation Functions (regarding DataFrames) slated for 1.5?