4 ms·
Unless you've got a ridiculous RAID array, there's plenty of times where processing anything over 100GB without a cluster is going to suck. Sequentially reading
by juiceandjuice 14y ago
Unless you've got a ridiculous RAID array, there's plenty of times where processing anything over 100GB without a cluster is going to suck. Sequentially reading 1TB of data is going to take a few hours even with a RAID array, or be very expensive with SSDs.
I maintain a database that's over 1TB and Oracle handles it very well. The trick there is understanding that, under absolutely no circumstance should you ever need to do a full table scan, because the table isn't designed for that.
So, I'd argue that a monolithic database is okay even up to 10TB even with slower disks as long as you never need to touch more than 10% of it. If you need to touch 100% of the data 100% of the time, I'd say anything over 100GB is too big for one machine.
The reality is that it just depends. There's times where you're going to want a hadoop cluster even for only 16GB of data, and there's going to be times where a database is going to be fine with 10TB of data.