12 ms·
2016 Spark Summit East Keynote
- josep2 11y agoStarted using Spark 1.6 a few months ago. Excited for the Kafka Connector feature.
- mydpy 11y agoMe too. Matei said that Spark 2.0 should be released (stable) late-April or early May. But you can always download and build the source!
- kod 11y agoWhat exactly was the mention of Kafka referring to? Spark has had decent kafka integration for a while now.
- peterstjohn 11y agoI think it's support for Kafka Connect, new in Kafka 0.9: http://kafka.apache.org/090/documentation.html#connect http://kafka.apache.org/090/documentation.html#connect
- kod 11y agoIt doesn't have anything to do with Kafka Connect. Matei was actually talking about the existing Spark Kafka direct stream implementation, which has been available since Spark 1.3 The video of the talk is available here: http://livestream.com/fourstream/sparksummiteast2016-tracka/videos/112612459 http://livestream.com/fourstream/sparksummiteast2016-tracka/...
- peterstjohn 11y agoAh, thanks! I have unfortunately been too busy to watch the streams this week, and assumed it was Connect because I think Confluent is/was doing a talk with Kafka Connect and Spark at the summit.
- azth 11y agoSlide 10: > CPU speeds have not kept up with I/O in the past 5 years. I presume he means the other way around? Also, what does he mean by native memory management? Does he mean off-heap allocation? And what's he referring to regarding code generation?
- mydpy 11y agoHe means the other way around: I/O improvements are outpacing CPU improvements. Native refers to (I think) the following issues: https://issues.apache.org/jira/browse/SPARK-12785 https://issues.apache.org/jira/browse/SPARK-12785 https://issues.apache.org/jira/browse/SPARK-8641 https://issues.apache.org/jira/browse/SPARK-8641 Code generation is enabled by SPARK-8641, but not sure exactly what it entails. I think it is related to some of the RDD transformation/action merging they do to optimize runtime operations in 2.0. Your thoughts?
- vvanders 11y agoNot overly familiar with Spark but usually the bottleneck for CPUs is access to DRAM. Unless you're doing completely linear reads or prefetching appropriately(which almost no one does right) you'll be cache-miss bound and your CPUs will be idle waiting for data.
- azth 11y agoThat's what I was thinking. The slide seems to imply that workloads are becoming CPU bound due to IO speeds increasing a lot.
- zzalpha 11y agoThat's precisely what he's saying. This is the thesis of those who have been watching the explosion in solid state disks. The claim is that bulk I/O is becoming so damn fast that the pendulum is swinging toward computing power being the new limiting factor.
- 11y ago
- mydpy 11y agoI think the more exciting announcement was Databricks community edition, which allows you to use 2.0: https://news.ycombinator.com/item?id=11126179 https://news.ycombinator.com/item?id=11126179
- eranation 11y agoVery excited to hear the plans for GraphFrames - finally GraphX getting some attention! https://spark-summit.org/east-2016/events/graphframes-graph-queries-in-spark-sql/ https://spark-summit.org/east-2016/events/graphframes-graph-...
- mydpy 11y agoAt Spark Summit East there are a lot of people evangelizing GraphX and trying to convince people to think about their problems using graphs.
- TheGuyWhoCodes 11y agoHas it become easier to run ad hoc queries with spark? I remember a year ago that the only available solution was the job server by ooyala. Which seems to be a missing feature of core Spark, and isn't something I was willing to bet my product on. Datastax evangelized people to use Spark to run queries over Cassandra but it looks so awkward and time consuming to copy jars around to the master, basically you need a dev ops team to this and even more scriptology for production.
- mydpy 11y agoInfinitely. Yes. The addition of Dataframes in 1.3 and numerous enhancements to SparkSQL have made it as easy as: val myDF = sqlContext.read.format("com.databricks.spark.csv") //allows you to read a csv file (for simplicity) .option("header", "true") .option("delimiter", "\\t") .option("mode", "PERMISSIVE") .option("inferSchema", "true") .load(filename) .where(filter_query) .cache() myDF.registerTempTable("myDF") sqlContext.sql("SELECT COUNT(*) FROM myDF").show() sqlContext.sql("SELECT COUNT(*) FROM myDF where filter_query").show()
- eip 11y ago>to run queries over Cassandra Why not just use Presto? It gives you basically full SQL capability for Cassandra with minimal effort.
- TheGuyWhoCodes 11y agoPresto seems like a good product, haven't really test it too much tho. When we started to look at Cassandra 1.5 years ago, Spark integration was all the rage, and the premise it was the missing link for doing analytic on Cassandra. We tested it thoroughly and came to the conclusion that is wasn't a mature enough solution.
- DannoHung 11y agoAre Spark streams ever going to reach a point where you can just have a table sitting in memory aggregating data and then you run queries on the whole thing without having to worry about windowing or anything?
- krcz 11y agoHow advanced is the Structured Streaming functionality? Looking at the JIRA [1] I cannot find even design prototype there, which is kind of strange if they want to have it ready by end of April. But as there was a presentation on the topic at the summit [2], I hope it's just developing it without discussion on JIRA. [1] https://issues.apache.org/jira/browse/SPARK-8360 https://issues.apache.org/jira/browse/SPARK-8360 [2] https://spark-summit.org/east-2016/events/keynote-day-3/ https://spark-summit.org/east-2016/events/keynote-day-3/
- mydpy 11y agoDid you read the proposal? https://issues.apache.org/jira/secure/attachment/12775265/StreamingDataFrameProposal.pdf https://issues.apache.org/jira/secure/attachment/12775265/St...
- mziel 11y agoLast Spark Summit the videos were up on Youtube 1-2h after each talk. Anybody knows where to find the ones from this summit?