Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
monstrado
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
61.
▲
by
monstrado
8y ago
Great questions! > what is the volume your system is operating at? This varies, as our workload is dynamic in that anyone at any time can inject a query for the data stream, but for this sake lets say 5k. > Also how does it work for s
62.
▲
by
monstrado
8y ago
As a resident of Raleigh, NC..I can say the feedback has been unexpectedly positive. Although, many people approve of the scooter situation, there has been a couple shortcomings: * Riding on sidewalks!! Not only does the app discourage you
63.
▲
by
monstrado
8y ago
Both the HLL (Algebird) and TDigest implementations we're using have a simple way to serialize a compressed representation. So basically just reading the row, merging the value currently stored, and writing the merged value back. Depen
64.
▲
by
monstrado
8y ago
Interfacing with the Java client using Scala
65.
▲
by
monstrado
8y ago
Real-time aggregations over a stream of data where we may have multiple servers writing a partial aggregation to the same row. With FDB I can safely read data, merge it with my in-memory copy, and then write the final result back. That'
66.
▲
by
monstrado
8y ago
I've been having enormous success using FDB for my POC. It's ability to do atomic mutations is honestly game changing for our use case. Mandatory transactions are also a lifesaver, as our previous implementation required careful O
67.
▲
by
monstrado
8y ago
Two come to mind. Apache NiFi: http://nifi.apache.org/ StreamSets: https://streamsets.com/opensource/ Side note, I'm an Apache NiFi committer and have had great success using it for enterprise ETL
68.
▲
by
monstrado
9y ago
You don't need to do this in real time, instead you could log the data (e.g. mouse clicks, key presses) to something like a database, Kafka, or Kenesis. Now you've unloaded the anti cheat logic to other servers. You don't nec
69.
▲
by
monstrado
9y ago
Why not just go to their documentation? http://kafka.apache.org/documentation/#introduction
70.
▲
by
monstrado
9y ago
Perhaps you haven't spent enough time reading into Kafka? It solves very different problems than the RabbitMQ/Clery "worker queue" pattern you're referring to. That doesn't necessarily mean people don't us
71.
▲
by
monstrado
11y ago
Great paper. This has a lot of similar goals to Apache Impala, mainly around improving raw data access to be as efficient as possible. http://pandis.net/resources/cidr15impala.pdf
72.
▲
by
monstrado
11y ago
I've had a chance to actually test this device and I can say it was comfortable, and the VR experience felt smooth and responsive. I was skeptical since it was powered by a phone, but I was definitely surprised.
73.
▲
by
monstrado
12y ago
Just because people use Hadoop and MapReduce interchangeably doesn't make it correct. I would love to hear how you think Spark can replace Hadoop, because that is an astonishingly inaccurate statement. Which part of Spark reliably dist
74.
▲
by
monstrado
12y ago
Saying things like "Going beyond Hadoop" is very misleading. Virtually all of the Hadoop vendors out there, whether it be Cloudera, Hortonworks, MapR commercially support Spark as a computation framework for Hadoop, and some alrea
75.
▲
by
monstrado
12y ago
Very few cluster sizes in enterprises reach this limit, and by limit I mean 1k -> 2k nodes. In reality, there's very little demand for namenode scalability at the current moment, but once this changes, the community will implement i
76.
▲
by
monstrado
12y ago
> The NameNode is a single point of failure. The NameNode service supports high availability out of the box, and uses a QJM (Quorum Journal Manager) to share edits between the active/standby node. To say the NameNode is a SPOF is no
77.
▲
by
monstrado
12y ago
Yes, the HBase scanners in Impala are not very fast, and we know that. This is an area that needs improvement to maximize parallelism, but as of right now there are a bunch of things on the Impala roadmap that takes priority (disk-based agg
78.
▲
by
monstrado
12y ago
Agreed, and also LLAMA doesn't support high-availability at the moment (soon to be fixed). We rely heavily on up to date table/column statistics in order to accurately determine resource consumption, and unfortunately Impala doesn
79.
▲
by
monstrado
12y ago
Although Impala is still a fairly new product, my team has been using it internally at Cloudera in production for over a year for real-time log analysis to our support engineers ( http://bit.ly/USFQdh ), among other ad-hoc BI
80.
▲
by
monstrado
12y ago
You pay a significant resource penalty when using Serdes, and since performance is one of the biggest priorities to the Impala team, we decided to leave this out for now. A very common workaround is to use Hive to generate Parquet data from
81.
▲
by
monstrado
12y ago
Although I have a lot of respect for the amplab, they did not do their due diligence with that benchmark. Mainly for a few reasons, they didn't test using columnar storage in Hadoop (ORC / Parquet), which is what Redshift is using
82.
▲
by
monstrado
12y ago
These type of articles baffle me, you're comparing a high-performance analytical database to a batch-orientated SQL engine. The whole point behind these query engines on Hadoop (Hive, Presto, Impala, etc) is to separate the database fr
83.
▲
by
monstrado
13y ago
Any particular reason why ORC was chosen as the columnar store format over Parquet ( https://github.com/Parquet/parquet-format )? Reason I ask is because Parquet seems to have its own development cycle, roadmap, and is p
84.
▲
by
monstrado
13y ago
Which areas would you say HDFS needs improvement the most? Just so you know, HDFS is still very actively developed, and keeps introducing features / improving functionality (e.g Native NFS, In-Memory Caching, Short Circuit Reads, Hig
85.
▲
Writing MapReduce Jobs with Clojure
(blog.cloudera.com)
1 points
by
monstrado
13y ago
|
0 comments
86.
▲
by
monstrado
13y ago
What are you using on the back-end to perform the queries? Are you using MapReduce? What is the average latency expectations when using the application?
87.
▲
by
monstrado
13y ago
Impala isn't technically Cloudera only, it's open source ( https://github.com/cloudera/impala ), and other people have gotten it to run on their Hadoop distribution, but since it's developed by Cloudera, i
88.
▲
by
monstrado
13y ago
The benchmark above is testing Impala with SequenceFiles compressed with GZIP, against RedShift, which is not a fair comparison. In the "What's next?" section, they say they want to re-do the Impala tests using Parquet, which
89.
▲
by
monstrado
13y ago
Although a little out of date, there is a website dedicated to this: https://amplab.cs.berkeley.edu/benchmark/
90.
▲
by
monstrado
13y ago
Is this similar to Myrrix? http://myrrix.com/
More ›