5 ms·
Are problems that fit in memory on a single machine (96GB in this article) considered "big data" now? This buzzword is becoming absolutely meaningless. But yeah
by paulasmuth 11y ago
Are problems that fit in memory on a single machine (96GB in this article) considered "big data" now? This buzzword is becoming absolutely meaningless. But yeah, "Hyper-interactive visualytics at scale.", go nvidia PR.
Also, while it's absolutely amazing that they are able to scan 240B rows/sec, I wonder what one would use this capability for if they can only keep a few hundred million records around at a time? The difference between taking 10 or 100ms to scan the dataset should hardly matter to a user that is running "interactive analytics queries".
- reitzensteinm 11y agoYou can rent a 128gb machine at Hetzner for $130 a month. I think we can all agree to draw a line in the sand and say if a kid doing a paper run could afford your servers, you don't get to call it big data. (not to take away from the linked project, which looks technically awesome)
- Smerity 11y agoI disagree with the "big data" line in the sand being whether "a kid doing a paper run could afford your servers". Using spot instances on AWS, you can have a 100 machine cluster with 1.5TB RAM for only $3 per hour. Working at Common Crawl, some of the coolest projects I've seen have been side projects or weekend projects by interested volunteers. WikiReverse[1] cost only $64USD to parse the metadata of all 3.6 billion pages (even cheaper if you avoided EMR fees) whilst Yelp extracted 748 million US phone numbers from 2 billion pages for $10.60USD. These days, big data (regardless of how you define the term) is within the reach of the kid with a paper route. [1]: https://wikireverse.org/ https://wikireverse.org/ [2]: http://engineeringblog.yelp.com/2015/03/analyzing-the-web-for-the-price-of-a-sandwich.html http://engineeringblog.yelp.com/2015/03/analyzing-the-web-fo...
- deleted 11y ago[deleted]
- riquito 11y agoI appreciate the analogy but it's not like if the server were free big data would disappear, isn't it?
- threeseed 11y agoBig Data as the industry sees it is more than just the size of the data. It is about the nasty work of data validation, consolidation, ETL, unification e.g. single customer view, analytics etc. Some of these can be technically challenging even on multi gigabyte files and require a range of multi disciplinary techniques. And Big Data increasingly means taking old EDW concepts and reimagining them for a more real time world.
- rubyfan 11y agoNo offense but what you just described is everything that is wrong with enterprise "Big Data" today. What many consultants are holding up as a big data panacea is just elbow grease and institutional fortitude... there is no free lunch.
- comboy 11y agoIf you know how to do even a simple WHERE on 750M rows in memory in 100ms without using GPU then please share some info, I'd love to learn about it.
- DannoHung 11y agoArray based columnar storage and SIMD comparisons aren't fast enough?
- rburhum 11y agoeh, any where query that hits an index where the 750M rows are not going to be scanned. The query can easily return in 100ms... fetching the records is a different story.
- paulasmuth 11y agoI actually recently left my job at google to start my own company and that's what we do. We didn't consider using GPUs yet as the queries we are running for the specific usecase we are currently working on (web/ecommerce analytics) are all IO bound [so the first step if we wanted to speed stuff up would be moving the data onto ssds or into memory as we are currently storing everything on disk. however this hasn't been necessary so far]. Our product allows you to run interactive queries (segment/aggregate/etc) against a billion record dataset (in the terrabytes; basically a giant logfile that stores an entry for every interaction we observe on the website/app with lots of metadata) with O(seconds) latency. We wrote our own code that does "the heavy lifting" but it's pretty much the same that everbody seems to be doing right now and comes down to minimizing the # of bytes read at query time; we index all data into a columnar format (so we only need to load a subset of the data and can pack it tightly for most queries) and then slice the query up into lots of individual shards that we compute in parallel (again, being bound mainly by the 200megs/sec disk read we get out of each machine). Without being completely serious now but regarding your question; "a simple WHERE" (scanning each row and applying a predicate fn with no shared state/shuffle/merge phase) is pretty straightforward to run on a huge dataset as you can split the execution up into as many pieces as you want and run them individually. Since this allows you to basically choose the problem size per shard you could scan those 750m rows on a ginormous cluster of tamagochis if you can fit a single row onto one ;) If you or somebody else would like to talk more, please shoot me an email: paul@deepcortex.io