5 ms·
I started using Impala this summer on AWS Elastic Mapreduce. Previously I'd been using Hive. We had at least a 100x speedup in most of the work we do because of
by gallamine 12y ago
I started using Impala this summer on AWS Elastic Mapreduce. Previously I'd been using Hive. We had at least a 100x speedup in most of the work we do because of it. Things that would have been impossible before (near realtime classification) are now feasible with mini-batches. My only (major) gripe is that AWS only supports several versions back. This makes it difficult to troubleshoot problems and find documentation. That said, the benefits far far outweight the cost.
- supergirl 12y agoimpala is good but it's no match for the best analytics databases out there
- anonymousDan 12y agoIn terms of performance or features? They had a paper at CIDR this year that claims they are competitive with a well known analytics DB (presumably oracle although it wasn't named explicitly).
- supergirl 12y agoafaik impala doesn't have multicore joins while almost any other serious db has. so if it uses 1 core out of 10, 20 to do the join then it might be 10x, 20x slower than the best
- eropple 12y agoMost of those analytics databases are pretty temperamental in cloud environments, though. I haven't used Impala, but something that can run on EMR (especially if it can get better turnaround times than Redshift) is pretty interesting. (Vertica in AWS is a tire fire. Avoid avoid avoid.)