Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
epdlxjmonad
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
TPC-DS Benchmark: Trino 476, Spark 4.0.0, and Hive 4 on MR3 2.1
(mr3docs.datamonad.com)
1 points
by
epdlxjmonad
1y ago
|
1 comments
2.
▲
by
epdlxjmonad
1y ago
In this article, we report the results of evaluating the performance of the latest releases of Trino, Spark, Hive-MR3 using 10TB TPC-DS benchmark. Trino 476 (released in June 2025) Spark 4.0.0 (released in May 2025) Hive 4.0.0 on MR3 2.1 (r
3.
▲
Performance Evaluation of Trino 468, Spark 4.0.0-RC, and Hive 4.0.0
(mr3docs.datamonad.com)
1 points
by
epdlxjmonad
1y ago
|
1 comments
4.
▲
by
epdlxjmonad
1y ago
Performance Evaluation of Trino 468, Spark 4.0.0-RC2, and Hive 4 on MR3 2.0 using the TPC-DS Benchmark In this article, we report the results of evaluating the performance of the following systems using the 10TB TPC-DS Benchmark. Trino 468
5.
▲
by
epdlxjmonad
3y ago
This article evaluates the performance of the following systems. Trino 418 (released on May 17, 2023) Spark 3.4.0 (released on Apr 13, 2023) Hive 3.1.3 on MR3 1.7 (released on May 15, 2023)
6.
▲
by
epdlxjmonad
4y ago
Diablo where you control a large party -- I tried this idea myself (using a bit of hack) and found it to be such great fun and extremely addictive. The only bad thing about it is that you can't go back to a single-player Diablo any mor
7.
▲
Spark on MR3 – A New Way to Run Apache Spark
(datamonad.com)
2 points
by
epdlxjmonad
5y ago
|
0 comments
8.
▲
Show HN: Fault Tolerance in Hive on MR3 on Kubernetes
(youtu.be)
1 points
by
epdlxjmonad
6y ago
|
0 comments
9.
▲
Run Hive on Kubernetes, even in a Hadoop cluster
(datamonad.com)
3 points
by
epdlxjmonad
6y ago
|
0 comments
10.
▲
Show HN: Hive on MR3 on Amazon EKS with Autoscaling
(mr3docs.datamonad.com)
1 points
by
epdlxjmonad
6y ago
|
0 comments
11.
▲
by
epdlxjmonad
6y ago
While I cannot give a definitive answer because I am not an expert on Spark internals, my opinion is that the discrepancy results mainly from query and runtime optimization. Apart from adding new features (e.g., ACID support), a lot of effo
12.
▲
by
epdlxjmonad
6y ago
Query plans are heavily optimized, and map-side joins are used extensively. The use of optimizations exploiting memory makes the so-called in-memory computing of Spark no longer relevant because Hive also uses memory efficiently. Hive commu
13.
▲
by
epdlxjmonad
6y ago
There is a common belief that SparkSQL is better than Hive because SparkSQL uses in-memory computing while Hive is disk-based. Another common belief is that Presto is better than Hive because it is based on MPP design and was invented for t
14.
▲
by
epdlxjmonad
6y ago
I agree that Spark on Kubernetes will have a hard time fixing the problem of shuffling. If they choose to use local disks for per-node shuffle service, a performance issue arises because disk-caching is container-local. If they choose to us
15.
▲
Correctness of Hive on MR3, Presto, and Impala
(datamonad.com)
1 points
by
epdlxjmonad
6y ago
|
0 comments
16.
▲
by
epdlxjmonad
6y ago
Using a single thread to simulate everything is cool (as stated in my previous comment on FoundationDB at https://news.ycombinator.com/item?id=22382066 ). Especially if the overhead of building the test framework is small. I
17.
▲
by
epdlxjmonad
6y ago
For debugging a distributed system, it may be just okay to use the traditional way consisting of testing, log analysis, and visualizing that everyone is familiar with. Yes, there are advanced techniques such as formal verification and model
18.
▲
by
epdlxjmonad
6y ago
Reading this book AND trying to follow its key lessons makes a huge difference in productivity, which I can testify from my own experience. This book is often compared to the bible of software engineering, suggesting that everyone knows abo
19.
▲
by
epdlxjmonad
6y ago
Yes, closely related, but invariants can appear anywhere in the code (like loop invariants), and are less restrictive than pre-conditions and post-conditions which must appear in the beginning and end of methods. So, invariants are more abo
20.
▲
by
epdlxjmonad
6y ago
When writing code, we often think "according to the design of our system, this condition must be true at this point of execution.” Examples are: 1. The argument x must not be 0. 2. The variable x must smaller than the variable y. 3. Th
21.
▲
by
epdlxjmonad
7y ago
Testing a distributed system using a single machine may look like an unorthodox approach. From our experience, however, when building a test framework for a distributed system, everyone would be automatically led to think about building it