5 ms·
Inspired by https://twitter.com/garybernhardt/status/600783770925420546 https://twitter.com/garybernhardt/status/600783770925420546
by lukegb 11y ago
Inspired by https://twitter.com/garybernhardt/status/600783770925420546 https://twitter.com/garybernhardt/status/600783770925420546
- jacquesm 11y agoHe's selling himself short.
- mosselman 11y agoCan someone explain in a bit more detail what this is about? Is the 'joke' that running data computation in RAM is faster than what? From disk?
- icebraining 11y agoI believe it's more "no, you don't need an Hadoop cluster of 20 machines, your data fits in the RAM of one machine".
- JDDunn9 11y agoRight, RAM an order of magnitude faster than disk, so calculations will be performed very quickly. Big data usually implies clusters of servers because the data won't fit on one server (even on the disk).
- jacquesm 11y agoBig Data usually implies 'big dollars', not necessarily a large amount of data. Simply use some in-efficient algorithm and a datastore with sufficient overhead and you're in Big Data territory. Re-do the same thing using an optimal algorithm operating on a compact datastructure and you make it look easy, fast and cheap. Of course you're not going to make nearly as much money.
- jordanthoms 11y agoThere is no point deploying a heavy, complex (and usually pretty slow due to the overheads involved) distributed database, when you could just buy a server with xTB ram, load any sql database on it, and run your queries in a fraction of the time. If your data is so large that it can't fit in the RAM of a single machine, then distributed databases make more sense (since loading data off disk is very slow, modulo SSD).
- jacquesm 11y agoFor some problems SQL would already be way too much overhead.
- e12e 11y agoCould you give a concrete example? If your working set is small, say 1TB -- and so fits in RAM -- for what kind of problems would using SQL be so much of an overhead that you need a different approach? And what would that approach be? I suppose you could have a massive set of linear equations that you might be able to fit into 1TB of RAM, but would be difficult to work with as tables in Postgres?
- jacquesm 11y agoTake a graph, you could use an SQL database to store it and do your graph analysis using SQL, or, alternatively, you could convert your graph to an extremely compact in-memory format and then do your analysis on that. Much better efficiency for the same size problem, bonus: you can now analyze much larger graphs with the same hardware.
- e12e 11y agoOr maybe a bit of both: http://stackoverflow.com/questions/27967093/how-to-aggregate-matching-pairs-into-connected-components-in-python http://stackoverflow.com/questions/27967093/how-to-aggregate... I appreciate you taking the time to answer -- and I get that there's a reason for why we have graph databases. But I really meant something more concrete, as in here's a real-world example that isn't feasible to do on machine X with postgresql, but easy(ish) with a proper graph structure/db -- rather than "not all data structures are easy to map to database tables in a space-efficient manner".
- jacquesm 11y agoOk, one more example: A German company holds a very large amount of profile data and wanted to search through it. On disk storage in the 100's of gigabytes. Smart encoding of the data and a clever search strategy allowed the identification of 'candidate' records for matches with fairly high accuracy, fetching the few records that matched and checking if they really were matches sped things up two orders of magnitude over their SQL based solution. It's very much dependent on how frequently you update the data and whether or not (re)loading the data or updating your structure in memory can be done efficient or not to determine whether or not such an approach is useful or not but going from 'too long to wait for' to 'near instant' for the result of a query is a nice gain. In the end 'programmer efficiency' versus 'program efficiency' is one trade-off and cost of the hardware to operate the solution on is another. Making those trade-offs and determining the optimum can be hard. But a rule of thumb is that a solution built up out of generic building blocks will usually be slower, easier to set up, will use more power and will be more expensive to operate but cheaper to build initially than a custom solution that is more optimal over the longer term. So for a one-off analysis such a custom solution would never fly, but if you need to run your queries many 100's of times per second and the power bill is something that worries you then a more optimal solution might be worth investing in.
- learnstats2 11y agoData that fits in RAM doesn't need any "Big Data" solutions.
- JonnieCache 11y agoThe subtext is that running a fancy distributed system is more exciting and beneficial for ones resume than simply buying a massive bloody server and putting postgres on it, and that people are making tech decisions on this basis.
- Anderkent 11y agoThis of course ignores that it's much easier to get your hands on a cluster of average machines than one massive bloody server, and all the non-performance-oriented benefits of running a cluster (availability etc.). Much easier to request a client provisions 20 of their standard machines, or get them from AWS. People don't like custom hardware, and for good reason.
- falcolas 11y agoAmazon offers some bloody huge servers... 32 core, 256GB RAM, and 48TB HDD space. d2.8x large
- brianwawok 11y agoThat is 4k a MONTH for 256gb of ram. If you could do the same job on a fleet of 8-16GB servers.. you can get a lot more CPU for a lot less dollars. Depends if you really need everything on 1 machine or not (as of course nothing will beat same machine in memory locality)
- jules 11y agoNot true, 8x16GB costs as much as 1x256 on Amazon. The issue here is that Amazon is hilariously expensive in general. Hetzner will rent you a 256GB server for €460 per month. Or you can buy one from Dell for $5000. These are not high numbers, in 1990 you paid more than that for a "cheap" home computer. For the price of a floppy drive back then you can now get a 32GB server.
- pquerna 11y agorackspace, onmetal-memory[1]: 512 GB, $1650/mo (3.22 $/gb/mo) softlayer, dual Xeon 2000 Series: 512GB, $1,823.00/mo (3.56 $/gb/mo) these are on-demand prices. pre-pay, or use a term discount, and its cheaper. Build it yourself: You can build a Dell or similar on a 2-Xeon-proc (E5 series), your main limit is getting good prices on 16x 32GB DIMMS. But lets say you can buy the RAM for ~$6500, then its just dependent on the rest of your kit, lets say $10,000 flat for the whole server. $277.77/mo over 36 months, but you still need network infrastructure, and you might want a new one in 12 months, but you get the general idea. [1] - http://www.rackspace.com/en-us/cloud/servers/onmetal http://www.rackspace.com/en-us/cloud/servers/onmetal
- isp 11y agoSee also: https://www.chrisstucchio.com/blog/2013/hadoop_hatred.html https://www.chrisstucchio.com/blog/2013/hadoop_hatred.html
- jacquesm 11y agoYou should post that if it hasn't been posted already, it's a much better way to make the case than the current link.
- isp 11y agoDone: https://news.ycombinator.com/item?id=9582060 https://news.ycombinator.com/item?id=9582060 I originally saw it on HN, but almost two years ago. Old comments: https://news.ycombinator.com/item?id=6398650 https://news.ycombinator.com/item?id=6398650
- feld 11y agoPeople will build gigantic compute clusters with expensive storage backends when their entire dataset fits in memory. If it fits in memory, it's going to be magnitudes faster to work with than on any other infrastructure you can build. So the trick is, you take their "big data problem" and hand them a server where everything can be hot in memory and their problem no longer exists.
- emodendroket 11y agoMost people who think they have "big data" problems actually don't have "big" data at all.