Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rxin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
rxin
10y ago
Please keep the contributions coming!
62.
▲
Spark as a Compiler: Joining a Billion Rows per Second on a Laptop
(databricks.com)
251 points
by
rxin
10y ago
|
53 comments
63.
▲
Approximate Algorithms in Apache Spark: HyperLogLog and Quantiles
(databricks.com)
2 points
by
rxin
10y ago
|
0 comments
64.
▲
by
rxin
10y ago
Yup was replying to you. Clicked the wrong button :)
65.
▲
by
rxin
10y ago
Yup similar to that. Our next blog post in the pipeline is going to reference this paper.
66.
▲
by
rxin
10y ago
Some of the performance gain was coming from the use of Unsafe in earlier versions of Spark (e.g. Spark 1.5). However, the massive gain you are seeing in Spark 2.0 are not coming from Unsafe. It is coming from this idea we call "whole-
67.
▲
by
rxin
10y ago
tl;dr is yes (to your last question). The longer answer is that this is about how to logically think about the semantics of computation using a declarative API, and the actual physical execution (e.g. incrementalization, record at a time pr
68.
▲
Spark 2.0 Technical Preview
(databricks.com)
259 points
by
rxin
10y ago
|
36 comments
69.
▲
The Unreasonable Effectiveness of Deep Learning on Spark
(databricks.com)
32 points
by
rxin
11y ago
|
14 comments
70.
▲
by
rxin
11y ago
Did it?
71.
▲
by
rxin
11y ago
Let me know if you have any questions .... :)
72.
▲
by
rxin
11y ago
Wechat had intense competition in China. And putting competition aside, I find wechat much more convenient and feature rich than other international messaging apps (eg whatsapp). Last time I visited China I paid for cab rides and bought ele
73.
▲
by
rxin
11y ago
As part of Spark 2.0, we are introducing some new neat optimizations to make a general engine as efficient as specialized code. I just tried on Spark master branch (i.e. the work-in-progress code for Spark 2.0). It takes about 1.5 secs to s
74.
▲
On-Time Flight Performance with GraphFrames for Apache Spark
(databricks.com)
3 points
by
rxin
11y ago
|
0 comments
75.
▲
by
rxin
11y ago
You guys should submit a talk to Spark Summit. Look forward to it.
76.
▲
by
rxin
11y ago
It's also not 100X faster than "Spark alone". This article is plain wrong.
77.
▲
Introducing Databricks Community Edition: Apache Spark for All
(databricks.com)
4 points
by
rxin
11y ago
|
0 comments
78.
▲
by
rxin
11y ago
The blog post actually provides code to reproduce all the steps and the chart. See http://go.databricks.com/hubfs/notebooks/TensorFlow/Distribu... http://go.databricks.com/hubfs/notebooks
79.
▲
by
rxin
11y ago
Did you actually read the article? It was using Spark to parallelize hyperparameter tuning, which is embarrassingly parallel.
80.
▲
by
rxin
11y ago
The "broadcast" is pretty cheap because often you already have the data in some distributed file system, or if on a single node the network bandwidth is pretty high. The problem with a lot of the deep learning workloads is that it
81.
▲
MLlib Highlights in Spark 1.6
(databricks.com)
2 points
by
rxin
11y ago
|
0 comments
82.
▲
by
rxin
11y ago
Author of the blog post here. Feel free to ask me anything about Spark.
83.
▲
Spark 2015 Year in Review
(databricks.com)
5 points
by
rxin
11y ago
|
2 comments
84.
▲
Spark 2015 Year in Review
(databricks.com)
2 points
by
rxin
11y ago
|
0 comments
85.
▲
by
rxin
11y ago
I just created a JIRA ticket tracking this: https://issues.apache.org/jira/browse/SPARK-12635 Thanks for the reminder!
86.
▲
by
rxin
11y ago
Despite our attempts at warning people, a lot of users still use groupByKey in RDDs. Hopefully over time this won't be a problem as the engine should be able to figure out more intelligently and do the proper rewrite (of course, we won
87.
▲
by
rxin
11y ago
This is true when you compare the performance vs Java/Scala, but if you compare it with other tools that are native in Python, it is not really much worse. For examples, Pandas operations that use custom UDFs are substantially slower t
88.
▲
Introducing Spark Datasets
(databricks.com)
2 points
by
rxin
11y ago
|
0 comments
89.
▲
by
rxin
11y ago
I think most of the improvements are indeed available in Python, including better memory management, improved Parquet performance, and the many algorithms. The main two things that are not yet available are the Dataset API and the streaming
90.
▲
Announcing Spark 1.6
(databricks.com)
104 points
by
rxin
11y ago
|
21 comments
More ›