Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rxin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
121.
▲
by
rxin
12y ago
GraphX actually graduated from Alpha in Spark 1.2. We have a few important improvements and changes to GraphX planned for 1.4, including Java API, and possibly a Python API.
122.
▲
by
rxin
12y ago
We are aware of at least 500 production use cases. Howeve,r since it is is open source software, there are a lot more that we don't know about. You can find some public ones here: http://spark-summit.org/east/2015&
123.
▲
Announcing Spark 1.3
(databricks.com)
103 points
by
rxin
12y ago
|
23 comments
124.
▲
by
rxin
12y ago
Actually I just realized the main difference -- if you load the data from memory, it is much faster, which the case of R.
125.
▲
by
rxin
12y ago
The example was mostly a toy example. The power really comes when you get interactivity for small data and big data. Using this, you could scale up to TBs of data on a cluster and still get results relatively fast, which is not something yo
126.
▲
by
rxin
12y ago
Indeed, DataFrames give Spark more semantic information about the data transformations, and thus can be better optimized. We envision this to become the primary API users use. You can still fall back to the vanilla RDD API (afterall DataFra
127.
▲
by
rxin
12y ago
In a way yes. It is a little bit more than that because DataFrames internally are actually "logical plans". Before execution, they are optimized by an optimizer called Catalyst and turn into physical plans.
128.
▲
by
rxin
12y ago
I'm one of the authors of the blog post as well as this new API. Feel free to ask me anything.
129.
▲
by
rxin
12y ago
There are some ways you can integrate the two. E.g. Streaming allows you to apply arbitrary RDD transformations, and thus you can pass a physical plan generated by DataFrame into streaming. We will work on better integration in the future t
130.
▲
by
rxin
12y ago
Somehow my comment was removed ... :(
131.
▲
Introducing DataFrames in Spark for Large Scale Data Science
(databricks.com)
138 points
by
rxin
12y ago
|
41 comments
132.
▲
by
rxin
12y ago
Maybe using BP-Means: http://arxiv.org/abs/1212.2126 K-Means without preset K.
133.
▲
by
rxin
12y ago
This is a cool feature, and is one of the prime example of what Spark's tight integration of various libraries can enable (in this case Spark Streaming and MLlib). It was originally designed by Jeremy Freeman to handle workloads in neu
134.
▲
Introducing Streaming K-Means in Spark MLlib 1.2
(databricks.com)
69 points
by
rxin
12y ago
|
10 comments
135.
▲
Spark officially sets a new record in large-scale sorting
(databricks.com)
3 points
by
rxin
12y ago
|
0 comments
136.
▲
by
rxin
12y ago
I believe it was for Jim Gray instead. http://archive.wired.com/techbiz/people/magazine/15-08/ff_ji...
137.
▲
by
rxin
12y ago
Yes, but the final data is sequential only. We were discussing about random access, which only applies to the intermediate shuffle file. Maybe you can email me offline. I can tell you more about the setup and how Spark / MapReduce work
138.
▲
by
rxin
12y ago
Yes absolutely. If I can get a single machine with 200TB of SSDs, that'd have been great :) But as soon as we have more than 1 node, then having more nodes is better. We can actually demonstrate this quantitatively. We are required to
139.
▲
by
rxin
12y ago
No it doesn't. The old record used 2100 nodes so the entire data actually fit in memory. There shouldn't be much seek happening even in the MR 2100 case. In Spark's case, the data actually doesn't fit in memory. Also thi
140.
▲
by
rxin
12y ago
Hi Todd, Except in the case of MR 2100 nodes the entire dataset fit in memory :)
141.
▲
by
rxin
12y ago
Actually Doug Cutting himself (who created Hadoop) tweeted about this. I guess Spark gets some of his blessing :) As pointed out in the article multiple times, we are comparing with MR here. We are not comparing with Hadoop as an ecosystem.
142.
▲
by
rxin
12y ago
It would help a little bit (maybe a few percent), but not much because the scheduling latency was relatively low for these tasks (the largest scheduling delay was ~10 secs, whereas each task takes minutes).
143.
▲
by
rxin
12y ago
Not sure why you mentioned seek time. In large scale, distributed sorting, I/O is mostly sequential.
144.
▲
by
rxin
12y ago
Hi Austin, We haven't tested Spark 0.8 at this scale. In general Spark is advancing at a rapid rate that 1.1 is very very different from 0.8.
145.
▲
by
rxin
12y ago
It was part of the enhanced networking. Without enhanced networking, we were getting about 600MB/s, vs 1.1GB/s with.
146.
▲
by
rxin
12y ago
The job is actually very linearly scalable. i.e. running it on 200 nodes roughly doubles the throughput of 100 nodes.
147.
▲
by
rxin
12y ago
Thanks for sharing this. I'm the author of this blog post. Free free to ask me anything.
148.
▲
by
rxin
12y ago
It is mainly the cost of getting nodes from EC2 at that point. It becomes hard to get a huge number of i2.8xl instances. Spark runs fine on thousands of nodes.
149.
▲
by
rxin
12y ago
The old entry had 10Gb/s <full-duplex> (40 nodes/rack 160Gbps rack to spine. 2.5:1 subscription), 64GB of RAM, and 12 x 3TB SATA. The network part is probably the most important one here, and both have comparable network.
150.
▲
by
rxin
12y ago
There are several use cases that I know of. One public one is Taobao http://databricks.com/blog/2014/08/14/mining-graph-data-with...
More ›