Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
monstrado
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
91.
▲
by
monstrado
13y ago
These are all very subjective statements, things like "I choose ... because". If you asked 20 programmers why they like statically typed languages, I am willing to bet you will get quite the variety of answers. Personally, I think
92.
▲
Building a business around open source. What works, and what doesn't.
(linkedin.com)
1 points
by
monstrado
13y ago
|
0 comments
93.
▲
by
monstrado
13y ago
> Salesforce is a frigging CRM tool Not really, they do a whole lot more than that, especially in the SDK land, I hear they even have their own IDE for coding the Salesforce way. > SAP is a frigging ERP system Another one of those com
94.
▲
by
monstrado
13y ago
Comparing apples to oranges really, but I suppose it's more accurate if you take into account all the people who are (or could be using) a relational database without much issue and then switching over. The article is right though, I&#
95.
▲
by
monstrado
13y ago
As one of of the team members that built this, I can try to answer any questions that you guys might have.
96.
▲
Analyzing billions of log lines in seconds, How our support team uses Impala.
(blog.cloudera.com)
2 points
by
monstrado
13y ago
|
1 comments
97.
▲
by
monstrado
13y ago
Thanks for the clarifications, I wasn't not meaning to discredit your article in anyway. I am just trying to help people understand that these databases are very different from eachother, and were created to solve different use cases.
98.
▲
by
monstrado
13y ago
These are very different databases, PostgreSQL is a transactional database, while Redshift (aka: ParAccel) is an analytical database. Each of these databases have implemented much different design decisions, which improve queries on certain
99.
▲
by
monstrado
13y ago
True, but there's no way to avoid people interpreting NoSQL as a relational database alternative. To your point regarding Google's F1, try looking at Impala ( https://github.com/cloudera/impala )...
100.
▲
by
monstrado
13y ago
HBase doesn't claim to be some NoSQL database or equivalent to a relational database, anyone who thinks otherwise hasn't actually read into HBase. Just look at HBase's website which describe it exactly as > Apache HBase is
101.
▲
by
monstrado
13y ago
Sorry, I should have clarified. When a node goes down, there isn't any sort of "minute" interruption...When there's a full scale outage and you need to perform an actual recovery, that can take minutes.
102.
▲
by
monstrado
13y ago
No problem, glad I could help. Your use case sounds pretty interesting, HDFS should fit the bill for sure. You should take a look at Parquet ( http://parquet.io/ ) for storing your data. It's an open source columnar form
103.
▲
by
monstrado
13y ago
Which version of HBase are you running? If we have a node or two go down in our cluster, HBase is completely unaffected. Recoveries can take up to a few minutes, in the old ages it took hours. I know that MTTR (mean time to recovery) is bei
104.
▲
by
monstrado
13y ago
I don't think that's something they are trying to solve right now. Although HBase guarantees write consistency, I think http://research.google.com/pubs/pub36726.html is the closest paper on how to go about it
105.
▲
by
monstrado
13y ago
Disclaimer: I work at Cloudera as a Tools Developer What do you mean by unstructured? Do you mean the data has yet to be parsed into a format which could be logically grouped into columns? Or do you mean that it's deeply nested? Since
106.
▲
by
monstrado
13y ago
If you're comparing this to your riak, redis, tokyo cabinet, ... database than most likely. With HBase, one of it's bread and butter operations is the start and stop row scan (otherwise known as a range scan). The only thing equiv
107.
▲
by
monstrado
13y ago
:) sometimes I don't know what to believe on the internet
108.
▲
by
monstrado
13y ago
Agreed, and some of these databases don't even embrace it. For example, HBase's main website makes no mention of being a "NoSQL" database ( http://hbase.apache.org/ ). Of course, given the right circumsta
109.
▲
by
monstrado
13y ago
Could you please elaborate? It totally depends on use case...
110.
▲
by
monstrado
13y ago
> impala is supposed to kill sql too I'm not sure what you mean by this, Impala is in no no way trying to replace "SQL". Impala is a general purpose distributed query engine that currently translates SQL to a query plan.
111.
▲
by
monstrado
13y ago
You might find this interesting, a database query engine using LLVM IR to speed up performance. http://blog.cloudera.com/blog/2013/02/inside-cloudera-impala...
112.
▲
by
monstrado
13y ago
Cloudera's Impala SQL query engine uses LLVM to maximize performance, here's an article published recently talking about how it's used. http://blog.cloudera.com/blog/2013/02/inside-cloudera-impa
113.
▲
LinkedIn Announces This Year's Top 10 Tech Startups
(blog.linkedin.com)
2 points
by
monstrado
13y ago
|
0 comments
114.
▲
by
monstrado
13y ago
I guess it depends on your use case, most of the use cases I've seen have been time series using range iterations which is incredibly fast, but I understand your concern if you're only using it for random gets.
115.
▲
by
monstrado
13y ago
Searching is not the same thing as a range iteration, normally if you want to search a Key/Value than you'll need to scan the entire dataset. A range iteration allows you to scan a subset of the data, and as long as they didn't fundamentall
116.
▲
by
monstrado
13y ago
Sorted keys are extremely useful for time series data. For example, if a key has a timestamp in it and you'd like to do an aggregation over a few days of "keys"...it's very simple to do an iteration from a STARTKEY to a STOPKEY (the keys in
117.
▲
by
monstrado
13y ago
Check into Google's LevelDB library.
118.
▲
by
monstrado
13y ago
It depends, if you plan to scan our entire data set it could take 30-40 seconds (roughly ~2.8TB), but we have our data partitioned based on a key that makes sense for the kind of data you'd need to populate a web page and these queries are
119.
▲
by
monstrado
13y ago
We have a 14 node cluster, the nodes have anywhere between 4-6 disks. Performance has been pretty amazing, we can do ad-hoc queries on this 4.5B row table. Each node has read throughput at about ~1.3GB/s for full table scans (data is snappy
120.
▲
by
monstrado
13y ago
We're currently running Impala in production with a table that currently has over 4.5B rows which powers an internal log analysis website. We don't have any hard limitations for concurrent queries, and no vacuuming since the data lives with
More ›