4 ms·
To me the problem seems more that on a GPU you can only execute fast queries on fairly small datasets. I would assume that the space of problems that still fit
by paulasmuth 10y ago
To me the problem seems more that on a GPU you can only execute fast queries on fairly small datasets. I would assume that the space of problems that still fit into a few gigs of GPU RAM but need to be answered rapidly with a lot of parallelism is fairly small.
In my anecdotal sample, most of the users of hadoop/bigquery/etc seem to use it less because of raw query speed but more because the datasets are simply to large for any classical solution (i.e. much larger than fits into RAM on a single machine)
- tmostak 10y agoYou can fit up to 256GB of GPU RAM in a standard server these days, backed by terabytes of CPU RAM as cache. In a rack we can easily deploy on terabytes of GPU RAM. This is big enough to handle datasets of over a hundred billion rows, which is not big data by everyone's standard but big enough that its almost intractable to handle in real-time with other solutions, particularly if the use cases requires a lot of scans (witness Mark's other benchmarks).
- paulasmuth 10y ago256GB is 256 billion bytes. So you can only fit over a hundred billion rows into 256GB if you have zero overhead and every row is smaller than roughly two bytes... ;) [so you couldn't even store a single proper integer in each row in your example] Also, querying a static dataset that you can fit completely into RAM is not exactly "intractable to handle in real-time" without a GPU. You can do that just fine using the CPU right now. Case in point: Even MySQL can handle 256gb of rows in memory with ease if you give it that much ram... Essentially, any single-machine database (gpu or not) can only process the same datasets that traditional single-machine databases (like MySQL) can already handle, albeit maybe a bit faster. I think a good example for a "newsql" usecase that does _not_ fit into classical solutions like MySQL or Postgres would be web tracking (clickstream) data: Even a medium sized web property (say alexa top 100 in germany or france) will generate dozens of gigabytes of tracking data per hour. over the period of months, this adds up to hundreds of terrabytes of data. There is no way to load that much data in one piece into any of the classical databases. EDIT: Maybe I should add a disclaimer: I'm the founder of another open-source database product that could be considered a competitor to MapD.
- tmostak 10y agoEven MySQL can handle 256gb of rows in memory with ease if you give it enough ram The scan performance of MySQL on that much data is going to be dramatically slower. MySQL may be good at indexed accesses but is simply not build for analytics workloads like those tested in the blog post. Not sure exactly how MySQL benches against Postgres, but the latter takes minutes over the same queries. http://tech.marksblogg.com/billion-nyc-taxi-rides-postgresql.html http://tech.marksblogg.com/billion-nyc-taxi-rides-postgresql...
- paulasmuth 10y agoThat's not a fair comparison though. I only quickly read the postgres benchmark you linked, but it looks like you stored the data on disk (!) on a machine with only 16GB of ram for the postgres benchmark, while storing the data completely in 96GB of very fast RAM for the mapd benchmark. Unless I misread the post it's just comparing apples and oranges. If anything, the two benchmarks show that one piece of hardware (RAM) is orders of magnitude faster than the other one (SSDs). My previous comment was specifically about MySQL running on a machine where the whole dataset fits into memory. Of course, I'm sure mapd is faster than mysql/postgres for some usecases. But the benchmark doesn't prove that in a fair comparison. EDIT: Maybe I should add a disclaimer: I'm the founder of another open-source database product that could be considered a competitor to MapD.
- tmostak 10y agoThat's not a fair comparison though. I only quickly read the postgres benchmark you linked, but it looks like you stored the data on disk (!) on a machine with only 16GB of ram for the postgres benchmark, while storing the data completely in 96GB of very fast RAM for the mapd benchmark. Fair enough, but even if the data is in RAM a CPU solution will still be much slower. See this benchmark running the same queries on a 7-node Redshift cluster. http://tech.marksblogg.com/billion-nyc-taxi-rides-redshift-large-cluster.html http://tech.marksblogg.com/billion-nyc-taxi-rides-redshift-l... And MySQL for all its strengths is not in an analytics database and for these types of queries will be much slower than Redshift.
- 10y ago