6 ms·
I think you may be conflating Hadoop the mapreduce framework with Hadoop the ecosystem, which includes hdfs, Hive, Spark and others. To the best of my knowledge
by monkeyfacebag 6y ago
I think you may be conflating Hadoop the mapreduce framework with Hadoop the ecosystem, which includes hdfs, Hive, Spark and others. To the best of my knowledge, former is waning in popularity (supplanted by tools like Spark), but the latter remains in wide use.
- throwaway_pdp09 6y agoStraight up, what's people's views on all this big data stuff? I see it in so many job ads and I really can't believe it's necessary. Sure, at the far end of one side of the bell curve are companies like uber but otherwise, are people using it to process a few terabytes that could be done better on one multicore server? How many companies have enough data to justify it? Personal opinions welcome.
- royjacobs 6y agoI've found that a lot of big data solutions in use by companies are only in use because business requirements were "we want all the data and we'll figure out what to do with it later", but then it turns out only a fraction of the data is used or only in a heavily aggregated form. That makes sense, of course, nobody is going to manually go through a few terabytes worth of records manually.
- trollied 6y agoTraditional SQL-based data warehouses still rule the world, and probably will for the foreseeable. HN reports on a small number of businesses that use bleeding-edge technologies (often just for the sake of it). The vast majority of businesses stick to the "traditional" way of doing things. At the end of the day, most places just want something to point PowerBI at etc.
- cyberdrunk 6y agoWhat I've seen is that thousands of jobs for dozens of team are run on company's Hadoop cluster. Sure, each team could provision their own custom infra and run the job there, but having a centralized way to do it, with all the extra niceties (scalable capacity, good monitoring, logging, HA), can provide some company-wide efficiencies. Plus, some of the jobs can in fact be huge and you may need dozens of nodes to process them (we have such jobs in a bank, where we don't really have big data) - doing it without a cluster would be problematic.
- throwaway_pdp09 6y agoSounds like you're one who might actually need it.
- cmrdporcupine 6y ago"Big data is like teenage sex: everyone talks about it, nobody really knows how to do it, everyone thinks everyone else is doing it, so everyone claims they are doing it. " https://www.quora.com/Big-data-is-like-teenage-sex-everyone-talks-about-it-nobody-really-knows-how-to-do-it-everyone-thinks-everyone-else-is-doing-it-so-everyone-claims-they-are-doing-it-What-is-your-opinion-about-this-statement https://www.quora.com/Big-data-is-like-teenage-sex-everyone-...
- devonkim 6y agoIt’s the same with almost all hyped up technology or movements trying to push stuff into the C suite for conversations.
- zten 6y agoAnswer the question about how to process a few TB on a multicore server, and you might still find yourself using Spark, or something like it. If you start from the assumption that you've been ingesting data and storing it in a compressed columnar format like Parquet or ORC, then you're already locked into a solution that exists in the Hadoop ecosystem. This turns out to be an effective way to deal with terabytes of data because depending on the query you want to execute, the file format (and a smart partitioning structure) helps turn your problem of reading terabytes into one of reading gigabytes. And, everything in Hadoop land is generally a query, so you're going to need a query engine like Hive, Impala, Spark, etc. AND, you probably need something like Hive metadata so that you're not crawling some directory structure for the schema and input files every time you start up a new process to run a query. You _could_ write something on your own that just forks out a bunch of threads in a single process to rip through the data, but why? Think about what you're effectively implementing - a bespoke query engine that runs one query plan. Spark has already written an API and a query engine (multi-process distributed, unlike whatever you're likely to hand-roll) and lots of input/output code. You can spark-submit --master local[*] and use up all of the cores and ram everything into one monster JVM if you really wanted to. Finally, you could stuff all of this in memory, but at what cost, and what are you going to use to do it? (Worse, what happens if the VM goes down - those terabytes are going to take their sweet time reloading over a network link.)
- throwaway_pdp09 6y agoHow; plenty of disks to give you the IO. It's usually about the IO. If you want memory, go to a server site and configure to max out a server with lots of DRAM slots. You might be surprised. But really getting enough mem is less important than IO, so stick with disks, SSDs I suppose. You can get single chip with dozens of cores for not too much. I guess that's how I'd do it. > This turns out to be an effective way to deal with terabytes of data A few terabytes don't need cluster, typically. > You _could_ write something on your own that just forks out a bunch of threads in a single process to rip through the data, but why? Because it's simple and easy. I wrote one in a few days. Not much code. > You can spark-submit --master local[] and use up all of the cores and ram everything into one monster JVM if you really wanted to. Point is, do you need to? > those terabytes are going to take their sweet time reloading over a network link You do not run big IO over a network like that if you can avoid it. With a single server with plenty of SSDs, you trivially don't. My take anyway.
- StreamBright 6y agoExactly. Btw. the other end of companies do not use Hadoop either. S3 + Presto is super popular nowadays in case of cloud.
- teraku 6y agoSpark being an example of a tool that (also) exists inside the Hadoop ecosystem and is still gaining acceptance due to popularity, but at the same time losing traction as better option are arising.
- rory_isAdonk 6y agoWhat better options? Can you drop some links? Thanks :)
- StreamBright 6y agoDepending on cloud vs on-prem.
- teraku 6y agoNot really, just use Apache Beam and set the runner to whatever you have / prefer
- StreamBright 6y agoNo thanks. I do not see any value of the n+1 unnecessary abstraction over things that I am already familiar with. An average customer does not want to use these things either.
- teraku 6y agoSo the main issue with Spark is that the streaming is not that great. In general a Spark cluster persé is not bad, but I see more and more people using Apache Beam and then just use a runner that either they already have or that fits them best. https://beam.apache.org/ https://beam.apache.org/
- eunos 6y agoApache Flink
- fongitosous 6y agois spark that widely used? arent people moving away from it for deep learning frameworks like tf?
- chrisjc 6y agoSpark is more than just a ML framework. It's extensively used for ETL, stream and batch processing, etc...