Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rxin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
91.
▲
by
rxin
11y ago
You can already use Spark SQL to query files (e.g. CSV, JSON, Parquet) directly on various storage systems without ETL. It doesn't have as much research novelty as NoDB, but it is used in a lot of production data pipelines and ad-hoc a
92.
▲
by
rxin
11y ago
If you use Spark SQL to query ElasticSearch as a data source, it already has better SQL support with joins and you can push predicates down into ES.
93.
▲
by
rxin
11y ago
Actually Spark SQL's data source API has a very expressive predicate pushdown interface and most data sources implement them. id = 1234 should not do a full scan.
94.
▲
by
rxin
11y ago
While you might find it hard to believe given all the negative sentiments echoed by Western media, in China there are lots of stories like this. I also know a lot of people that have done it this way. Not necessarily multi billionaire, but
95.
▲
Announcing Spark 1.5
(databricks.com)
2 points
by
rxin
11y ago
|
0 comments
96.
▲
by
rxin
11y ago
This is great work, but certain part of the comparison is not accurate, probably due to their lack of understanding of Spark. First and foremost, it would make more sense to compare against the DataFrame API of Spark, which is very Pandas l
97.
▲
by
rxin
11y ago
"tree allreduce or bittorrent"
98.
▲
by
rxin
11y ago
I don't think he said anything about Spark inventing allreduce. Spark did use torrent broadcast, which I believe is pretty unique.
99.
▲
When One App Rules Them All: The Case of WeChat and Mobile in China
(a16z.com)
165 points
by
rxin
11y ago
|
58 comments
100.
▲
Spark and Mesos–shared History and Future [Mesosphere Hackweek]
(mesosphere.com)
1 points
by
rxin
11y ago
|
0 comments
101.
▲
Announcing Apache Spark 1.4
(databricks.com)
154 points
by
rxin
11y ago
|
45 comments
102.
▲
by
rxin
11y ago
Coincidentally, R on Spark (SparkR) is also announced today: http://databricks.com/blog/2015/06/09/announcing-sparkr-r-on... It will appear in Spark 1.4 to use R on a cluster of machines, or a single mac
103.
▲
by
rxin
11y ago
That part is actually coming: https://github.com/apache/spark/pull/5713
104.
▲
by
rxin
11y ago
It is actually not possible to enforce type checks at compile time, due to the dynamic nature of data (e.g. you can generate a DataFrame from JSON files whose schema are automatically inferred by Spark, or generate a DataFrame by loading a
105.
▲
by
rxin
11y ago
Actually Spark SQL doesn't load everything into memory. Its data source API supports pushing predicates down, and if the data sources implement it, it even supports running aggregations and joins in the data sources!
106.
▲
by
rxin
11y ago
Here's the JIRA ticket: https://issues.apache.org/jira/browse/SPARK-7075
107.
▲
by
rxin
11y ago
And on a related note, congratulations to Matei Zaharia for winning the ACM Best Dissertation Award for his work on Spark: http://awards.acm.org/doctoral_dissertation/
108.
▲
by
rxin
11y ago
Author of the blog post here. Kay already pointed out that increasing cluster size usually reduces network utilization. In addition to that, there a few challenges with just getting bigger cluster size with "on-demand" resources:
109.
▲
Project Tungsten: Bringing Spark Closer to Bare Metal
(databricks.com)
3 points
by
rxin
11y ago
|
0 comments
110.
▲
by
rxin
11y ago
I couldn't find any pricing information for printing a figurine on the website. Anything you can disclose more? Thanks.
111.
▲
Recent Performance Improvements in Spark: SQL, Python, DataFrames, and Mor
(databricks.com)
9 points
by
rxin
11y ago
|
0 comments
112.
▲
Deep Dive into Spark SQL’s Catalyst Optimizer
(databricks.com)
3 points
by
rxin
11y ago
|
0 comments
113.
▲
by
rxin
12y ago
Sorry I don't think the last paragraph you said is true. I dont see any special data types or models that can be modeled by MPP architecture but cannot be modeled in Spark. In short, I don't believe there is much difference at the
114.
▲
by
rxin
12y ago
Spark 2.0: Rearchitecting Spark for Mobile Platforms https://databricks.com/blog/2015/04/01/spark-2-rearchitectin... This is probably among the most technical, nerdy April fool's.
115.
▲
Spark Turns Five Years Old
(databricks.com)
5 points
by
rxin
12y ago
|
0 comments
116.
▲
Improvements to Kafka Integration of Spark Streaming
(databricks.com)
5 points
by
rxin
12y ago
|
0 comments
117.
▲
by
rxin
12y ago
Glad you liked it! I left another comment about for-yield in a separate comment.
118.
▲
by
rxin
12y ago
Thanks -- internally we are having a lot of debate about for comprehension. In particular, the use of it might simplify certain use cases (especially async with futures). That said, it also often leads to very high-entropy, dense code and t
119.
▲
by
rxin
12y ago
Thanks - fixed!
120.
▲
Databricks Scala Style Guide
(github.com)
49 points
by
rxin
12y ago
|
29 comments
More ›