4 ms·
Back in college in '97 or so, in our databases class on the first day, the professor asked, "Who's worked with databases before?" A bunch of hands went up. "Oh,
by jdk 13y ago
Back in college in '97 or so, in our databases class on the first day, the professor asked, "Who's worked with databases before?" A bunch of hands went up. "Oh, sorry, let me rephrase, who's worked with databases larger than a few dozen gigs?" Only one or two hands remained up. "If it's smaller than that, just save yourself the effort and use a flat-file instead."
15 years later and it's the same thing, plus a few orders of magnitude.
- corresation 13y agoInteresting, but it isn't really similar advice. Indeed, it sounds like absolutely terrible advice (unless there is some context that is missing).
- AndrewDucker 13y agoI suspect it should have been "Who's worked with databases too big to fit entirely in memory". Because if it all fits then you don't have to worry too much about performance overheads. If it doesn't then you need to think about how you're going to avoid table-scans and the like.
- corresation 13y agoI remember buying a pretty kick-ass server at the time and it had, if I recall correctly, 196MB of memory. However I would still strongly contest the analogy even if the database did fit in memory (the author may be off by magnitudes themselves, as in 1997 a 2GB database was a pretty substantial, unweidly thing for most people) -- doing the simple steps of putting your data in a database instantly enables enormous flexibility in the use of that data at very little cost or overhead, with better to enormously better performance than the average person is going to yield with a flat file. This situation (Hadoop for big data), in contrast, is about throwing away a lot of flexibility, and paying a large performance price, to add big-data scale out flexibility. It is, in many ways, the opposite situation.
- dalke 13y agoMany programs use SQLite in part because its ACID properties make it an excellent way to save system state. A lot of people use databases to implement persistent user state on top of stateless HTTP, even if the data itself is small enough to fit into memory. This latter use was known even in 1997. For example, the book "Database Backed Web Sites: The Thinking Person's Guide to Web Publishing" was published on Jan. 1 of that year. So even when that advice was offered, it was wrong.
- cmccabe 13y agoYes, it is terrible advice. Just like this article. Repeats all the usual myths about Hadoop... that it's just about MapReduce, that it doesn't support indexes (hint: Hive has a CREATE INDEX command, guess what it does?) and then adds some of its own. After seeing this, plus an article advocating web frontend programming in C, I'm starting to think this place is going downhill fast.
- paul_f 13y agoWhat horrible advice. Hope you dropped the class.
- VLM 13y agoWould have been hilarious to listen to the lectures about normalization and ACID topics... assuming there were any. "Cod Normal Form? Never heard of it. I like my fish sticks made of haddock anyway." Come to think of it, I've worked with guys who apparently learned everything they know about databases from that prof's database class, unfortunately.
- scott_s 13y agoAside from the ACID properties others mentioned, if you're using a relational database (and most people are), there may be non-performance related benefits. Some data is inherently relational, and it can be easier to manage it using relational abstractions such as SQL.
- pnathan 13y agoThat is perhaps the worst advice I've ever heard. I spent a summer some years ago writing what, in abstract, were hand-coded SELECT and INSERT statements for CSVs. It was a waste of time; the number of bugs was ridiculous and the data integrity was attained by brute force. Today I'd use sqlite and/or postgres.