4 ms·
I actually recently left my job at google to start my own company and that's what we do. We didn't consider using GPUs yet as the queries we are running for the
by paulasmuth 11y ago
I actually recently left my job at google to start my own company and that's what we do. We didn't consider using GPUs yet as the queries we are running for the specific usecase we are currently working on (web/ecommerce analytics) are all IO bound [so the first step if we wanted to speed stuff up would be moving the data onto ssds or into memory as we are currently storing everything on disk. however this hasn't been necessary so far]. Our product allows you to run interactive queries (segment/aggregate/etc) against a billion record dataset (in the terrabytes; basically a giant logfile that stores an entry for every interaction we observe on the website/app with lots of metadata) with O(seconds) latency.
We wrote our own code that does "the heavy lifting" but it's pretty much the same that everbody seems to be doing right now and comes down to minimizing the # of bytes read at query time; we index all data into a columnar format (so we only need to load a subset of the data and can pack it tightly for most queries) and then slice the query up into lots of individual shards that we compute in parallel (again, being bound mainly by the 200megs/sec disk read we get out of each machine).
Without being completely serious now but regarding your question; "a simple WHERE" (scanning each row and applying a predicate fn with no shared state/shuffle/merge phase) is pretty straightforward to run on a huge dataset as you can split the execution up into as many pieces as you want and run them individually. Since this allows you to basically choose the problem size per shard you could scan those 750m rows on a ginormous cluster of tamagochis if you can fit a single row onto one ;)
If you or somebody else would like to talk more, please shoot me an email: paul@deepcortex.io
- comboy 11y agoSo as far as I understand it, indexes mentioned by others are another story, if index fits in mem then it's great. But as far as I understood this is about what in classic db would be called a sequential scan (that is we don't know beforehand what we will look for). Of course IO is a bottleneck, but it's not that simple even if you have it in mem already. As you said when sharding is possible you can just throw more boxes at it. But on a single machine the problem is checking what's in your own memory fast enough using CPU. That's where GPUs come in. Even apart from the name of your company, the idea seems very interesting. I suppose I'm not the only one curious, so could just say a few more words? Are you working on some new technology or is it more like scalable RDBMS focused on latency as a service? I really like the idea of the latter btw, I'd like to use something like Amazon RDS or similar solution, but even ignoring the price they don't seem to be able to beat few boxes with fine tuned postgres when it comes to performance on large datasets. On the other hand, with big data often this the data is the money and companies cannot afford to let somebody else store it. Btw, maybe put some box with e-mails subscription on your website? I'd love to learn how it will develop but of course without it I will forget about it in a few days.
- paulasmuth 11y ago> But as far as I understood this is about what in classic db would be called a sequential scan You misunderstood my comment then. The scheme I described is used to do "full table scans" on the data. Also this is not something I came up with; there is a ton of research on this and it's how a number of fairly well known DB/analytics products work. > But on a single machine the problem is checking what's in your own memory fast enough using CPU. That's where GPUs come in. You can already scan these data volumes in millisconds on traditional hardware without doing too much fancy stuff or using GPUs. My original point was that I have a hard time coming up with an "interactive analytics" usecase where you'd want to process TB/s on one node that can keep only a few gigs around -- the improvement of doing this on GPUs instead of the old fashioned way on general purpose hardware seems to be that a query over the same dataset returns a few milliseconds faster. I reckon a user running interactive sql queries doesn't really care if their queries return in 10, 50 or 100ms -- I am not even sure I would be able to notice that difference myself. However, if you actually had a usecase for this you could still use the approach I described and make it as fast as you desire by scaling it out [tweaking the shard size] on conventional/commodity machines. > you can just throw more boxes at it. But on a single machine The linked article is discussing a setup with at least 8 distinct processing units, too. > Even apart from the name of your company, the idea seems very interesting. I suppose I'm not the only one curious, so could just say a few more words? Are you working on some new technology We are focused on delivering actionable insights as well as solving some very specific data problems for our customers right now. IMHO the hard part of doing that is making sure we understand the customer's domain, gather/track the right data from their systems and then work with them to slice and dice and visualize this data to discover "signal" from the "noise" which we can then feed back and use to optimize their website/app. The technical part of being able to handle the relatively large volumes of (non-preaggregated) source data is really just the means to that end. We do not currently offer an "off-the-shelf" version of our product, but again if you or somebody else would like to talk more, please ping me. Also I hope I not coming across too negatively here. The linked research (compiling SQL to LLVM IR) sure is exciting stuff. I just couldn't help but feel that the hyperbole PR speak was a bit too strong with this one.