4 ms·
> But as far as I understood this is about what in classic db would be called a sequential scan You misunderstood my comment then. The scheme I described is us
by paulasmuth 11y ago
> But as far as I understood this is about what in classic db would be called a sequential scan
You misunderstood my comment then. The scheme I described is used to do "full table scans" on the data. Also this is not something I came up with; there is a ton of research on this and it's how a number of fairly well known DB/analytics products work.
> But on a single machine the problem is checking what's in your own memory fast enough using CPU. That's where GPUs come in.
You can already scan these data volumes in millisconds on traditional hardware without doing too much fancy stuff or using GPUs. My original point was that I have a hard time coming up with an "interactive analytics" usecase where you'd want to process TB/s on one node that can keep only a few gigs around -- the improvement of doing this on GPUs instead of the old fashioned way on general purpose hardware seems to be that a query over the same dataset returns a few milliseconds faster. I reckon a user running interactive sql queries doesn't really care if their queries return in 10, 50 or 100ms -- I am not even sure I would be able to notice that difference myself. However, if you actually had a usecase for this you could still use the approach I described and make it as fast as you desire by scaling it out [tweaking the shard size] on conventional/commodity machines.
> you can just throw more boxes at it. But on a single machine
The linked article is discussing a setup with at least 8 distinct processing units, too.
> Even apart from the name of your company, the idea seems very interesting. I suppose I'm not the only one curious, so could just say a few more words? Are you working on some new technology
We are focused on delivering actionable insights as well as solving some very specific data problems for our customers right now. IMHO the hard part of doing that is making sure we understand the customer's domain, gather/track the right data from their systems and then work with them to slice and dice and visualize this data to discover "signal" from the "noise" which we can then feed back and use to optimize their website/app. The technical part of being able to handle the relatively large volumes of (non-preaggregated) source data is really just the means to that end. We do not currently offer an "off-the-shelf" version of our product, but again if you or somebody else would like to talk more, please ping me.
Also I hope I not coming across too negatively here. The linked research (compiling SQL to LLVM IR) sure is exciting stuff. I just couldn't help but feel that the hyperbole PR speak was a bit too strong with this one.