4 ms·
No, actually Java is a bane to the database world. Cassandra doesn't work, and Hadoop is a complete waste of hosts for most companies (hence the move to Spark.
by throwaway_xx9 10y ago
No, actually Java is a bane to the database world.
Cassandra doesn't work, and Hadoop is a complete waste of hosts for most companies (hence the move to Spark.)
- slantedview 10y ago> hence the move to Spark ...which also runs on the Java Virtual Machine and is subject to the same pros and cons.
- sgt101 10y agoheh - you beat me too it!
- sgt101 10y agoDisagree on two fronts : - By hadoop I assume you mean Map Reduce ? There are other Engines like SAMSA and FLINK and Kafka makes a great event store. Anyway, MR is super for massive throughput batch jobs, for example huge HIVE queries or Pig jobs and if you are reading, breaking the heap size, doing one thing and then writing there is no bonus from doing it in SPARK. - SPARK is written in Scala, which runs on the JVM. And it has a nice Java API as well !
- growse 10y agoI'd be interested to know how you can assert that Hadoop is a 'complete waste of hosts for most companies'. Also, don't underestimate the many, many people successfully running Spark on YARN at scale. Hadoop is actually quite helpful to some workloads.
- dkersten 10y agoMost companies simply don't have the data volume to make Hadoop worthwhile. You can process tens of TB in an RDBMS on a beefy machine cheaper than a Hadoop cluster. Hadoop is slow, but on huge data volume the overheads are dwarfed by the parallelism gained. Most companies don't have huge volume though. For example recently I saw someone propose using Hadoop for a sub-TB dataset...
- nicobn 10y agoBlanket statements like "Cassandra doesn't work" and "Hadoop is a complete waste of hosts for most companies" are unproductive and contribute nothing, unless you can back them with data and real world examples. So, what data do you base these assertions on ? Also, not to burst your bubble but a lot of businesses (if not the majority) run Spark on YARN. And Spark is built on the JVM.
- pc86 10y agoIf they had data and examples they would almost certainly have enough experience not to say things like "______ doesn't work" and "______ is a complete waste of hosts."
- btym 10y agoSpark, Kafka, Flink, Storm, YARN, Samza, etc... Good luck staying out of the JVM. "The database world" is a bane to the big data processing world.
- markbnj 10y ago... logstash, kibana, elasticsearch, lucene, solr ... yeah pretty hard to not run java if you're doing distributed, scalable systems.
- frugalmail 10y agowtf?
- dang 10y agoPlease keep programming language (and other technology) flamewars off HN. You're welcome to make a substantive critique. We detached this subthread from https://news.ycombinator.com/item?id=11669189 https://news.ycombinator.com/item?id=11669189 and marked it off-topic.