3 ms·
While I understand the sentiment behind this post, I think it misses one crucial point: It costs time, effort, and very smart people to build the "Bugati"-like
by xtacy 12y ago
While I understand the sentiment behind this post, I think it misses one crucial point: It costs time, effort, and very smart people to build the "Bugati"-like system as they describe, instead of the current systems (that are more like "Toyotas", to name one).
I haven't seen the paper yet, so I can't be sure, but I think the numbers might ignore many factors: First, you need some kind of abstract, exchangeable storage (e.g., protobufs) to work with the data in many languages. Third, there's the file-system and all its intricacies. Fourth, it's unlikely that any compute environment will be dedicated only to one application (there's scheduling, resource management, and all that, which means there are hidden costs to doing network IO due to contention, protocol quirks, etc.). And finally, any realistic application is more than just "solving" the problem in the fastest way possible. Requirements change all the time, new features will be added, the code needs to be readable, understandable, maintainable, etc.
It's possible to do all the above AND be super efficient, but it requires a tremendous level of understanding of a system at all levels that it can be quite challenging, and frankly, with business requirements, it's probably not worth the time. If there's a framework that gives you abstraction but compiles to the fastest possible specific implementation AND makes a programmer productive, I would love to read up more!
- ploxiln 12y agoI've worked at a place had a good excuse for using a real "big data" processing system - a large hadoop cluster on bare metal - because the dataset was way bigger than would fit on a single machine. The cluster had over 20 machines that had 10 or so 3TB disks each, and a large computation might use half the dataset. Even if you could put all the data in a storage appliance and read it from a single machine, you needed the separate memories and data busses to just sift through the data and keep the relevant bits in non-swapped memory in less than a day. I think the important lesson from this paper, and which a few researchers also learned from our cluster, is that the amount of inefficiency that can be in software, and then removed by competent programming, is astronomical these days. You see many arguments like yours - programmer time is expensive, you need to be a super expensive expert, just pay for the systems, blah blah... it underestimates the cost of the ridiculous inefficiency, and overestimates the cost of competency. Even scientists in-experienced with serious programming can often get appreciably better at writing their data processing jobs before their first job finishes when you're dealing with the really big data. What a lot of people call big data isn't even big data. They'd rather go through the motions of setting up and using a big-data processing system and using it poorly, than learn better software engineering skills, even if that would take less of (theirs + others) time, amortized over the next few months of their work. This isn't toyota vs bugatti. This is... freight train with conductor, engineer, station staff, and one car of payload... vs getting a drivers license for a large van.