6 ms·
RAID doesn't meet his definition of "scalable" because it has a central controller. This whole post could be summed up as "ACID doesn't scale", which has been
by TimothyFitz 17y ago
RAID doesn't meet his definition of "scalable" because it has a central controller.
This whole post could be summed up as "ACID doesn't scale", which has been proven. Consistency, Availability or Partition Tolerance; pick two (http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.20.1495 http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.20.1...).
Good introduction to non-ACID databases:
http://highscalability.com/drop-acid-and-think-about-data http://highscalability.com/drop-acid-and-think-about-data
- Confusion 17y agoRAID setups can be outfitted with redundant controllers, in which case it does meet his definition of 'scalable'.
- nettdata 17y agoOh please. The mere fact that he's talking "RAID" instead of SAN speaks volumes. (No pun intended). Any storage engineer worth their salt would be rolling their eyes right now. The doc you linked to was authored in 2002, about the same time that Oracle's RAC was introduced (late 2001). Some of his citations are from the 80's. THE 80'S! DB technology has come a LONG way in 8 years, and that paper is no longer valid, unless you're talking about some basic, "minor"/open source db technologies. If you want to say "open source DB technologies have problems scaling", then go right ahead, and I'll agree. Just don't mind those of us who continue to build large, scalable systems, using the proper DB technologies, that disprove that "sql doesn't scale" generalization.
- rjurney 17y agoSANS are too expensive. They don't scale cost-effectively compared to commodity PC hardware.
- nettdata 17y agoI'd beg to differ. Go take a look at Adam Leventhal's work with the Fishworks stuff. Specifically, go check out the Sun Storage 7310. It will scale HUGE, and is nowhere near the stupid cost of NetApp or EMC or the other major vendors. I'd love to see how a bunch of commodity PC's will scale to 100TB, and still be manageable, and have anywhere near the same feature sets. Again, if you're railing against something as ubiquitous as a SAN, I'm not sure there's anything I can say to change your mind. Not that I'm really here to change your mind. Again, the original focus was about SQL not scaling, and you seem to be fixated on the cost of that scaling.
- rjurney 17y agoYou can use Hadoop to scale to 100TB using Commodity PCs and still be manageable, much easier and cheaper than you can use Oracle to do same. The featureset could easily DWARF those available using Oracle, as you can mapreduce all the data using your own hadoop jobs, and it can hold whatever kind of data you want it to - you're not limited to a static schema and precomputed summaries. HIVE is running on commodity PCs at facebook on more than 100TB, and they prefer it to their enormous, overly expensive Oracle OLAP system. So there's your answer, for one use case. But of course, if you're updating the data often - you wouldn't use hadoop (or if you were, you would run HBase or some such on top of it). But there are many use cases where you only write once at that scale. And in those cases, from my perspective - its much nicer to scale on commodity hardware than on big iron.
- moe 17y agoI'd love to see how a bunch of commodity PC's will scale to 100TB, and still be manageable 100T is 666 spindles when you go RAID10 with 300G drives (common size in the SAS/FC area). If you go S-ATA then it'll fit on 200 spindles using 1T drives. I'll stick to the latter variant for now because I'm too lazy to lookup Sun's pricing for SAS spindles, for my comparison below. So, 200 spindles amounts to roughly 15 hosts (throwing in a few spares for good measure). The whole setup will comfortably fit into one rack, including the FibreChannel machinery and other fluff that you'll likely want. Thus from the hardware side this is trivially managable, 100T is just not a lot of data nowadays. On the software side it's up to your creativity and mostly depends on what you actually need to store. I've seen people setup commodity postgres clusters, as well as more fancy things like HDFS, GFS or homegrown storage layers that way. And it worked. And, regardless of the indeed relatively sane pricing of the Sun (formerly StorageTek) products, the bottomline is what makes the difference. In figures, for a 100T SAN on the 7310 you're looking at something like $50k for one head, plus around $75k for three trays. We're in the $125k ballpark, hardware-only. And I'm being rather optimistic here: This setup actually holds only 96T and the head is maxed out (3 trays max per controller/head). That means your next upgrade will incur another $25k markup for the next head, good thing you didn't ask for 150T... Squeezing the same amount of storage into 15 min-spec supermicro pizzaboxes I arrive at roughly $3000 per node, including spindles. A good buyer will get them cheaper. This commodity cluster sets us back only ~$45k in hardware. That's 1/3 the price of the 7310 solution, being optimistic on the Sun and pessimistic on the commodity side. That kind of difference makes up for a lot of development effort for the custom solution - most of which is a one time investment anyways and yields better flexibility in the long run. It's also the reason why most of the big boys don't use off-the-shelf SANs for their primary storage.
- derefr 17y agoWhy are there no open-source implementations of the "proper DB technologies", then?
- nostrademons 17y agoPeople can still make money charging for them? Open-sourcing is the last step in the technology lifecycle, after the technology has become widely understood and commoditized. When people can still make money off something, they will. What's the incentive for them to give it away?
- derefr 17y agoUsually, an ideological one. The first group of Linux kernel developers (past the "toy" stage) all hated Microsoft with a passion, and wished to deprive it of as much revenue as possible. They wanted to give people a "free alternative", and lower the total investment people [that is, they] would have to make in owning a computer. I could imagine an analogous situation with some developers and Oracle.
- trezor 17y agoBeing ideologically for open software and an open operating system is not the same as "hating Microsoft with a passion". Mixing these two different concepts into the same bag is misleading, as the first one represents an ideological conviction and the other merely childish spite. And if you insist on being cheap, don't be surprised when it turns out that your free database was indeed some cheap stuff which doesn't come fully featured.
- derefr 17y agoIt's a causal relationship. If you believe that all software should be free (the GNU folks—those "first kernel developers"), then you must at least dislike any company which tries to profit from the creation of artificial scarcity of software. The two groups (the idealists and the "haters"), which now have very little overlap, originally started much the same. > Don't be surprised when it turns out that your free database was indeed some cheap stuff which doesn't come fully featured. But if being "fully featured" is the goal, then it would be very surprising indeed if what you considered a "competitor" was not, in fact, fully featured. It would not, then, by definition, be a competitor. Or, at least, it would not be worth calling version 1.0 yet. To refine that: Oracle currently has no FOSS competitors, because there are no FOSS databases that are trying to compete with Oracle. They may be trying to take parts of Oracle's market share, but this is a different thing—optimizing their fit for a situation where Oracle itself is a bad fit.