4 ms·
Yeah, I agree this is surprising. If all the data in their CPU implementation were stored in a well-packed columnar format, I'd expect something close to theore
by etrain 11y ago
Yeah, I agree this is surprising. If all the data in their CPU implementation were stored in a well-packed columnar format, I'd expect something close to theoretical memory throughput on the query they describe (intermediate counts table fits easily into L1/L2 cache and results are commutative and associative). Cache-coherence issues in the aggregation code are a possible cause.
If I were building a speed-freak analytics database I'd be focusing on making my CPU implementations as fast as possible, since that's what 99% of potential customers are already running. Assuming you get to 40GB/s on this type of query, that's only a factor of 5 slower than the GPU implementation. I'd imagine that for most workloads, a factor of 5 speedup that requires new hardware and lots of energy is kind of a non-starter.
- AceJohnny2 11y ago> If I were building a speed-freak analytics database I'd be focusing on making my CPU implementations as fast as possible, since that's what 99% of potential customers are already running. Well, this is by Nvidia, so of course their incentive is otherwise...
- bsprings 11y agoThis work was not done by NVIDIA, it was done by MapD, a startup. NVIDIA is promoting the work of a partner; it's a guest post on the NVIDIA developer blog.
- tmostak 11y agoHi, one of the original authors here. You're confusing rows per second with bytes per second. We're measuring here in rows per second, with random data that takes four bytes per record. So the group-by is showing around 24 GB/sec for CPU and roughly 1 TB/sec for the GPUs. Admittedly we are a small startup and there are some optimizations still to be made on CPU (we're working on being NUMA-aware, for example), but the CPU performance is not bad and still much higher for group by than you see in other databases. You have to remember that we're building a full database and visualization system and not just optimizing for a single benchmark. In addition we're trying to make the point that you can hit this 1-2TB/sec (depending on the query) on a single server with GPUs, which means assuming the compressed data fits in 192GB of GPU RAM you can get much higher performance that you would see out of a whole rack of beefy CPU servers, particularly when you take into account the network overhead that distributed databases suffer. Furthermore, the bandwidth of Nvidia's Pascal architecture (to be released next year) should have at least 2X the bandwidth by using High Bandwidth Memory and likely significantly higher memory sizes, so our speedups will only increase.
- vardump 11y agoYou seem to have two CPU sockets, so I'm assuming you have totally 8 memory channels. You should have 100 GB/s CPU memory bandwidth available. Without NUMA optimization, bandwidth across QPI is obviously much, much less. In that benchmark you have 8x Nvidia K80s, $4,595.99 each [1], which you compare against 2x unspecified 8 core Intel Xeon CPUs. How would it look like if you had $36k worth of servers? For every 2 GPUs, you can buy 4 servers with 2 sockets of Xeons each, 64 GB of RAM per server [2]. Total bandwidth for 8 GPUs $ worth of two socket Intel servers is 1600 GB/s. Totally 1 TB of RAM. Granted, it takes more space and more power and communication over ethernet, FC, etc. is slower. On the other hand, instead of just 64 GB per node, you could also expand each 2 CPU node to 1 TB. That's totally 1 * 4 * 4 = 16 TB. [1]: http://www.amazon.com/Nvidia-Accelerator-passive-cooling-900-22080-0000-000/dp/B00Q7O7PQA http://www.amazon.com/Nvidia-Accelerator-passive-cooling-900... [2]: Maybe not exactly this model, but they should have something within the parameters: http://www.supermicro.nl/products/system/2U/2027/SYS-2027TR-HTRF_.cfm http://www.supermicro.nl/products/system/2U/2027/SYS-2027TR-...
- pdeva1 11y agohave you considered the cost of your system? the aws g2.8xlarge comes with 16gb vram and will cost almost $1,900 per month. the 192gb of vram you are talking about will cost $22,000 per month!! A system with 192gb of normal ram will cost almost an order of magnitude lower than that. How does it make sense to run MapD in that case considering the enormous cost.
- Bedon292 11y agoIt doesn't make sense to run it on AWS, and not just because of cost. Which is why it is run on bare metal, which costs approximately the same as two months of AWS.