11 ms·
Don't use Hadoop – your data isn't that big (2013)
- isp 11y agoPrevious comments (2013): https://news.ycombinator.com/item?id=6398650 https://news.ycombinator.com/item?id=6398650 Also worth a read, "Command-line tools can be 235x faster than your Hadoop cluster (2014)": http://aadrake.com/command-line-tools-can-be-235x-faster-than-your-hadoop-cluster.html http://aadrake.com/command-line-tools-can-be-235x-faster-tha... | https://news.ycombinator.com/item?id=8908462 https://news.ycombinator.com/item?id=8908462
- fs111 11y agoPick a technology that can grow with your dataset. Take a look at Cascading or if you like Scala, scalding. They work on your laptop and on a 2000 node cluster the same way.
- kod 11y agoI honestly don't know why anyone would use Cascading or Scalding when Spark exists.
- saryant 11y agoBecause they have legacy Hadoop to support?
- kod 11y agoSpark interoperates with hdfs.
- evancasey 11y agoHuge spark fan here. Love the execution model, API, supporting libs etc. Unfortunately, Spark doesn't scale well on large datasets (10TB+). Sure, it's possible (and has been done), but right now there are too many rough edges to make it a better choice than Scalding/Cascading for data processing at scale. Most of this boils down to fine tuning certain Spark parameters, which is a pain when you're dealing with long-running, resource intensive workflows.
- fs111 11y agoBecause Cascading does not break your apps by changing the API every minor release, that would be one. Also, being production grade software for many years is another.
- baldfat 11y agoLook at R and realize your data can fit in RAM. Almost every data set examples for clustered could really fit inside of a computers memory.
- bagels 11y agoWhat servers support, say, 64tb of ram, let alone the petabytes that a much larger company may have to deal with?
- jasode 11y ago>If you have a single table containing many terabytes of data, Hadoop might be a good option for running full table scans on it. The author mentions this at the end almost as a footnote but in my experience, this is usually the motivating reason for IT to use Hadoop. It's about processing throughput and not just absolute data size. A lot of processing jobs can't use indexes (except for lookup tables) so a PostgreSQL 4TB db with a schema for OLTP workloads doesn't solve the issue. A 4tb harddrive might be cheap for $150 but at 100MB/sec, it takes 11+ hours to read. However, a big data cluster with 10 nodes can each scan their local 400GB shard and the batch job gets done in 1 hour. Even newer 10TB mechanical harddrives with 150MB/sec (and multi disks RAID to saturate a SATA3 interface) isn't going to fundamentally alter this throughput equation. On the other hand, a potential 4TB SSD with 1GB/sec might be the sweet spot if the processing is not cpu-constrained. But then again, new technology enables the goal posts to move[1], and people will invent new workloads that require a 100-node cluster each having a 10TB SSD to solve bigger problems. [1]http://en.wikipedia.org/wiki/Induced_demand http://en.wikipedia.org/wiki/Induced_demand
- mark_l_watson 11y agoGreat point. I don't have applications needing Hadoop right now, but in cases where (for example) I had a ton of data in S3 and needed to make some passes through it then Elastic MapReduce was appropriate and inexpensive. Also, there are good machine learning libraries available for Hadoop and also Spark, and the overhead of using Hadoop and/or Spark might make sense to have head room for larger data sets without recoding.
- bsg75 11y ago> A lot of processing jobs can't use indexes (except for lookup tables) so a PostgreSQL 4TB db with a schema for OLTP workloads doesn't solve the issue. What are the cases where an RDBMS can't use indexes on fact tables?
- sokoloff 11y agoSearching within fields is one example. WHERE street_address LIKE "%MAIN%ST%" WHERE xml_blob LIKE "%<deprecrated_feature_node_name%" to find documents that need to be updated when you remove that feature, etc. "Just use Mongo" or "just use the latest pgsql and convert to JSON" or "just do XYZ" is often a less practical answer than "just query/analyze the database you already have".
- foobarge 11y agoContains the famous ``Too big for Excel is not "Big Data".''
- gesman 11y ago...but resumes looks better after that :)
- kf5jak 11y agoI upvoted this purely because of the title!
- jowiar 11y agoUsing "big data" technology is more than just an factor of input size. If your operations making the data bigger, you might well need a bigger tool. You can start with a data set on the middle-to-high-end of memory-sized, but if you have an operation in there which combinatorially explodes the input data, Spark is going to start looking pretty good.
- lemmsjid 11y agoAt my company, we operate on very large amounts of data. Let's pretend, though, that we woke up this morning and had a 600mb dataset. We'd still have hundreds of jobs running on that 600mb dataset. Most of them would ideally run as often as possible--we'd only schedule them periodically because of performance. Some of them would explode the 600mb dataset into several gigs. Most of them would copy the dataset several times over. Some of them would be memory intensive. Some of them would be CPU intensive. Some of them would be disk intensive. Many of those jobs emanate counters, logs, and metrics, which we want to track over time. We'd want to track the overall resource utilization of those jobs. We'd want the particular jobs to be killed automatically if they use too many resources. If the jobs are in danger of taking up too many resources in the system, we want them to be automatically queued and scheduled piecemeal. We'd want the jobs to be written in a variety of frameworks targeted to the particular task at hand. Given those hundreds of jobs competing for scarce resources, we'd want to not have to completely rewrite the system if the 600mb dataset became a 1 gig dataset, or a 1 tb dataset. Sure, we could put together a multi server ssh/bash job scheduling system, but in the end that's kind of what hadoop is, except having addressed a lot of problems that you don't even know you're going to have up front.
- lmm 11y agoWith a 600mb dataset you probably don't need batch jobs. You can afford to run ad-hoc reports when they're requested. You can process data as it arrives, streaming it into a structured representation in postgres or similar. You can afford to stop the pipeline when you get an item whose format doesn't match, and investigate, rather than having to tolerate a certain error percentage. You don't need a scheduling framework because if you ever have too many jobs running you can kill a few, you're not going to lose hours of work by doing so. You want metrics but IME hadoop isn't actually very good at that. If you're doing enough processing that you can't do it on one beefy database server then sure, the concerns that go with that are similar to when you have enough data that you can't handle it on one beefy database server. But most people with 600mb of data don't want or need to do multiple cpu-hours worth of processing on it.
- michaelochurch 11y agoI could be wrong here, but my sense of Hadoop is that it's going to prove to be a dead end in the long run. Granted, evolutionary dead ends in technology can be very lucrative (look at Cobol, or Java). They aren't exciting, but there's money in them. That said, the aesthetic compromises made to appeal to the Java community, scaled out over years or decades, tend to make things bloated and hard to use. I don't see people being excited to do Hadoop, or wanting to learn it on their own time. It's much more interesting to learn the fundamentals of distributed systems than to learn a specific set of Java classes. Personally, I'm finding myself increasingly burned out on "frameworks" that appeal to out-of-touch corporate decision-makers (who are attracted to the idea of a packaged "product" that solves everything) but tend to be overfeatured while under-delivering (especially if time to learn how to use the framework properly is included) when it counts. I'm much more interested in the modularity that you see in, say, the Haskell community where "framework" means something that would be a microframework anywhere else.
- leeleelee 11y agoA lot of people probably run into problems from poor design, poor programming, or just inability to problem solve. For example (two real examples I've faced): (1) Company xyz is having problems generating reports that pull from an sql database. They take too long, IT says it's because the data is too large and their hardware is old. Obviously this is big data! Wrong, re-writing the sql queries and indexing some fields solved the problem. Like, beginner-level stuff. Nobody had even looked at the code or structure of the database when initially trying to solve the problem. or (2) Company xyz can't load a dataset into memory to perform some computations on it, it's too large. Needs better hardware or big data solution obviously. With millions of rows and many columns, ok. Only 6 of those columns were even used by the processing job, and a fraction of the rows were relevant. There was no need to fit the entire dataset into memory in the first place, problem solved. I could go on. Lots of companies just lack the skills and knowledge to do things right, and have no idea what's going on under the hood or how things work. So naturally, they are drawn to the marketing buzz surrounding big data, cloud computing, data warehousing, machine learning, etc etc etc.
- polsoul 11y agoBeing a dba for a decade already, resolving issues similar to the ones you described, are the sweatest I've experienced in my professional career. Unfortunately, I don't believe that fixing the problem in the database by understanding and resolving it in a logical manner is the way things will be in the near future. You know, RDBMSs aren't sexy ... developers developers developers , web developers web developers web developers :) the newest API and/or language are much more important nowadays for the newcomers. I cannot believe it, it's out of my mind, but that's the reality. Noone wants to utilize the database at its fullness anymore, everything needs to be done at the gazilion of app servers by using java/.net or the trendy 3rd/4th generation language of the day. My only hope is that the data is the thing that is not going away, and this forces us to think in a logical way about it(eventually devs/web devs will start doing the data maniplation and etc the right way, I guess you're pretty familiar with the approach in mind). Thanks
- nostrademons 11y ago