4 ms·
What are people actually using these days for analytics? What is missing? (i.e. what would the dream startup do for you?) There's a lot of hype around certain
by amit_m 11y ago
What are people actually using these days for analytics? What is missing? (i.e. what would the dream startup do for you?)
There's a lot of hype around certain technologies, and I suspect it does not reflect reality. I've used hadoop+pig a few years ago and it was absolutely terrible.
- bagels 11y agoI'm working on moving some analytics from pig + hadoop streaming over to Redshift. The debug cycles are so much faster with sql queries.
- blumkvist 11y agoYou can use sql on hadoop...
- threeseed 11y agoI am assuming you tried Hive, Impala, Spark SQL ? Redshift seems like a great fit for Adhoc analytics. Just a shame you can't run it in your own data centre (some of us can't use the cloud).
- ddorian43 11y agoYou can use citusdb (postgresql, columnar). Not the same, but close.
- alecco 11y agoIf you are using SQL... why didn't you try a columnar engine?
- saosebastiao 11y agoRedshift is a columnar engine
- bagels 11y agoTo all those who have asked, yes, hive is also in use.
- threeseed 11y agoThe majority of big data analytics is using the standard Hadoop stack and testing Spark. Dream startup would be to make the day to day life of running a cluster more manageable. If I had an app that could tell me how to optimize queries, tune the parameters and make the what/why/where/how questions easier I would buy it instantly.
- virmundi 11y agoLook no further. Netezza by IBM dies that all with SQL. Seriously, I'm impressed by that product. It also supports Hadoop.
- IndianAstronaut 11y agoI absolutely love Spark and really look forward to SparkR joining the main project this summer.
- alecco 11y agoIf the data is structured, columnar databases still rule. But if you are stuck in SV's echo chamber it's all MapReduce/Hadoop/Spark/etc.
- pacala 11y agoHow about Spark + Parquet via SQLContext / SchemaRDD?
- alecco 11y agoThe overhead of that must be ridiculous.
- pacala 11y agoWhat kind of overheads should we be aware of? Spark can run on a single machine, it could be an interesting benchmark to compare against Vertica or similar products. I've seen similar systems tuned for under 1s response times, even when running in distributed mode, though they were proprietary.