7 ms·
Postgres on the GPU
- hippich 13y agoNVidia's CUDA only (for now?) Anyone can explain why opensource projects embrace CUDA over OpenCL? As I understand OpenCL is more generic API which could be potentially used with CPUs and GPUs.
- apaprocki 13y agoOne reason why lots of things typically use CUDA is because NVidia makes datacenter rackmounted GPU gear. So there is no need to run generically if that is the only available "production" hardware.
- tlarkworthy 13y agothere are matrix multiplication routines developed for CUDA. OpenCL you would have to do everything from the ground up. So Nvidia gave everyone a head start for numerical computation, and that edge has ever since snowballed.
- sanxiyn 13y agoYou can use http://viennacl.sourceforge.net/ http://viennacl.sourceforge.net/ for OpenCL.
- eelsen 13y agoTry programming in both CUDA and OpenCL and see which one you would choose.
- dljsjr 13y agoThere are probably a lot of factors. I worked on CUDA code for around a year, and used to understand the landscape pretty well, but if I were to start a high-performance computing project today I'd probably take my lumps and go with OpenCL. There would be a lot of lumps. Firstly, CUDA is just more mature; there is a very large and well-established set of libraries for a lot of common operations, there is a decent sized community, and Nvidia even produces specialized hardware (Tesla cards) designed just for CUDA. Second, all that generic-ness of OpenCL doesn't come for free. With Nvidia, you're just working with one architecture; CUDA cards. Optimizing your kernels is much easier. OpenCL is just generically parallel, so you could have any sort of crazy heterogeneous high-performance computing environment you have to fiddle with (any number of CPU's with different chipsets and any number of GPU's with different chipsets). I haven't used OpenCL myself, but almost purely anecdotally I have heard many people say that CUDA is often slightly faster[1] and the code is easier to write. TL;DR: CUDA sacrifices flexibility for ease of development and performance gains. OpenCL wants to be everything for everyone, and comes with the typical burdens. [1]: Maybe this is a result of OpenCL being more generic and so harder to optimize.
- smosher 13y agoSo why would you go with OpenCL? Is it portability? For a while I thought that was worth it but I am having a hard time remaining convinced of that.
- cronin101 13y agoI've been working on a rather large computation library using OpenCL. OpenCL is useful for providing an abstraction over multiple device types. If you are only interested in producing highly-tuned parallel code to execute on NVidia hardware, I suggest sticking to CUDA for the above reasons. I utilised the OpenCL programming interface to write code that would run the same kernel functions on CPU and/or GPU devices (using heuristics to trade-off latency/throughput) which is something that is not possible afaik using the CUDA toolchain. TL;DR YMMV and horses for courses.
- apaprocki 13y agoFYI regarding highly-tuned code -- An ex ATI/AMD GPU core designer told me that the price you pay for writing optimized code in OpenCL versus the device specific assembler is roughly 3x. Something to keep in mind if you're targeting a large enough system to OpenCL and you find spots that can't be pushed any faster.
- cronin101 13y agoUnlike previous versions, OpenCL 2.0 been shown to only be about 30%[1] slower than CUDA and can approach comparable performance given enough optimisation. Since I am working on code generation of Kernels to perform dynamic tasks, I can't afford to write at the lowest level available. (I'm accelerating Python/Ruby routines though so OpenCL gives a significant bonus without much pain at all.) [1] http://dl.acm.org/citation.cfm?id=2066955 http://dl.acm.org/citation.cfm?id=2066955 (Sorry about the paywall, I access through University VPN)
- deleted 13y ago[deleted]
- fixxer 13y agoYou're preaching to the choir. Looks like NEC funded this. NVidia seems to be the preferred hardware for institutions/big companies. I'm not sure if this is because NVidia's architecture is better for supercomputers or if they're simply better at marketing to those types of customers NVidia funds a lot of academics in my space, and I've found academia to be very anti-open source for those reasons, which amuses me greatly. Case in point, Matlab. Why is this taught in a world with Python/Numpy/MatplotLib?
- mjn 13y agoMatlab seems like inertia/culture to me: it's the longtime de-facto standard in engineering. Since it's what everyone uses, it's got packages for everything, and papers will often come with prototype Matlab implementations. Roughly like the cultural position R holds in statistics. Matlab's hold on engineering is also bolstered by its widespread use in industry: students want to learn it, because it's what their future employers use, and professors / research scientists like to use it because it's what their industrial collaborators use. In my area of CS (artificial intelligence) it seems considerably less popular. I don't really remember how to use it, since the last time I used it seriously was in some engineering (but not CS) courses in undergrad.
- eru 13y agoIn addition, the MathWorks have so far managed not too screw up too badly and are keeping Matlab up to date. (They are definitely quite nice as an employer.)
- fixxer 13y agoI've heard this argument a lot (since 1999), but I'm not confident it holds true anymore. The free alternatives are so good. You may be right regarding Professors, but that is also changing as they age out.
- mjn 13y agoCould be; we don't use Matlab much in my own research area, so recent change could've happened under my radar. When I've occasionally had contact with engineers in industry, though, Matlab still seemed to be everywhere. The most recent two examples were someone doing DSP, and someone doing mechanical engineering, and both had all their stuff built on top of Matlab+Simulink.
- tkahn6 13y agoFrom my experience porting CUDA code to OpenCL code, CUDA is much cleaner and more succinct since it is able to assume a lot about the underlying hardware.
- guard-of-terra 13y agoAnother question, is CUDA actually fast enough? Because in bitcoin mining, it's always ATI/AMD and OpenCL because ATI cards are like ten times faster than Nvidia cards. This is because of architecture differencies. Does it not affect this postgres table-scanning task? I wonder if they did any benchmarks.
- duskwuff 13y agoAMD's advantage in Bitcoin mining was purely due to an architectural quirk: their shader cores supported bitwise rotation, but Nvidia's didn't. Bitwise rotation is a rare instruction outside of certain crypto algorithms (like SHA256!), so this really means very little for general-purpose performance. http://www.extremetech.com/computing/153467-amd-destroys-nvidia-bitcoin-mining http://www.extremetech.com/computing/153467-amd-destroys-nvi...
- guard-of-terra 13y agoScrypt too? Because Nvidia performance on scrypt is bad either.
- mrb 13y ago"AMD's advantage in Bitcoin mining was purely due to an architectural quirk" False. I authored a Bitcoin miner utilizing this quirk (bit_align). I was also the first to leverage another instruction exclusive to AMD (bfi_int): https://bitcointalk.org/?topic=2949 https://bitcointalk.org/?topic=2949 bit_align "only" gave AMD a 1.7x advantage over Nvidia. The biggest perf gains (2x-3x!) came from the fact AMD has more execution units: https://en.bitcoin.it/wiki/Why_a_GPU_mines_faster_than_a_CPU#Why_are_AMD_GPUs_faster_than_Nvidia_GPUs.3F https://en.bitcoin.it/wiki/Why_a_GPU_mines_faster_than_a_CPU... (I also authored this section of the wiki).
- jtc331 13y agoThis is because the hashing algorithm is highly dependent on integer rotate right instruction. AMD's implements it in 1 clock cycle, Nvidia in 3. So it's a special case.
- 13y ago
- jacquesm 13y agoCUDA is a lot more mature and easier to program for. OpenCL likely has the future, but it takes more work to set it up and if you have Nvidia cards it is harder to optimize.
- potkor 13y agoNvidia software has always been much better than AMD. AMD on Linux is a complet disaster. So if you work on GPU it's natural to go Nvidia.
- papsosouid 13y ago>AMD on Linux is a complet disaster. So is nvidia. Both companies produce absolutely horrible drivers. Not just for linux either, tons of bluescreens, crashes and other windows instability issues are video driver bugs. That is what happens when the sole concern is speed, and stability is totally ignored.
- cmccabe 13y agoI have to agree with the grandparent. AMD's graphics drivers for Linux are a complete disaster. NVidia's are only a partial disaster. And sometimes that's the best you get.
- nwhitehead 13y agoThe fun part of parallel programming is getting things running on your GPU, parallelizing the algorithm, then tuning and optimizing the code. This is easier, faster, and more pleasant in CUDA with it's mature tools and ecosystem. That's why open source projects often use CUDA. The advantage of OpenCL is that it runs on more platforms (not just NVIDIA). The problem is that it's more complicated and more of a headache. My advice to programmers is to start with CUDA and play around with your problem for a while. Time spent learning how GPUs work, what kinds of operations are efficient, and how to parallelize algorithms is not wasted if you switch to OpenCL later. Once you've made some progress then make an informed decision about whether you want to go to production with CUDA or OpenCL.
- homerowilson 13y agoNvidia uses ECC RAM on their GPU compute cards, an important consideration for serious HPC computing.
- wtallis 13y agoNo, they don't. They implement ECC by re-purposing some of the existing RAM to hold the parity data. Enabling ECC reduces the usable amount of memory, and also can hurt performance. It offers some improved reliability, but it's nothing like a real server-grade memory system.
- DigitalTurk 13y agoReally, really cool! I've toyed with the idea of doing pattern matching (and graph rewriting) on the GPU before but this looks like it's much more advanced than I thought was feasible. I'm surprised they went with CUDA instead of OpenCL though. CUDA is proprietary NVidea technology and does not work for non-NVidea devices.
- cs702 13y agoIf I understood correctly, this module allows you to export data to, and then access (read-only), foreign tables architected specifically for faster copying of data to/from GPU, allowing you to speed up queries that benefit from GPU computation by 10-20x. Nice. On the surface, this seems very similar to https://news.ycombinator.com/item?id=5592886 https://news.ycombinator.com/item?id=5592886 except it's nicely integrated with Postgres. Or am I missing something? tmostak?
- masklinn 13y ago> this module allows you to export data to, and then access (read-only), foreign tables architected specifically for faster copying of data to/from GPU That seems to be pretty much it, although the read-only part is a limitation of Postgres rather than a feature of the module. Furthermore, it seems to be dispatching queries intelligently so you can perform all queries against the FDW table with (I hope) minimal overhead, if the qualifier can't be compiled to a kernel the fdw will run it as a normal on-CPU qualifier. That's a thoughtful touch.
- rarrrrrr 13y agoBrilliant! Perhaps it's just my scars showing, but I'm concerned about database system stability with active GPU hardware added to the box. I probably wouldn't add the GPU hardware to a master but rather do the queries that would benefit from it on a streaming replica, where an occasional kernel panic won't be so severe. Regardless, looking forward to trying it.
- jonknee 13y agoIt's read only...
- oakwhiz 13y agoI think that he is referring to the tendency of GPU drivers to crash. Even if the DB is read-only from the GPUs' standpoint, if the DB goes down because of faulty drivers it's still a problem.
- NuZZ 13y agoOn the flip side, if this is sufficiently adopted, it could present motivation to driver developers, thus improved drivers. Perhaps this and the linux gaming movement could mean some symbiosis for driver development. Ignoring Windows as I guess I don't really take Windows servers too seriously.
- spitfire 13y agoI notice the GPU load time is about 53ms. This is a discrete graphics card so I do wonder how an integrated APU will affect this. I can imagine the overhead there being virtually nil. The long term trend is for the GPU to merge with the CPU, so I think we'll see more of this in the future.
- perlgeek 13y agoAll the examples seem to use numbers (integers and floats). It would be interesting to see if it can work efficiently with variable-width strings, which is the main workload that I encounter. But even if not, I see the value working with lots of data.
- hippich 13y agoGPU is good at calculating stuff (nVidia in particular at calculating floats.) I do not see reasons to do text stuff with GPU
- qb45 13y agoGPUs had been used for some kind of fuzzy string matching [1] and worked well thanks to their huge memory bandwidth. [1] http://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm http://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorith...
- metrix 13y agoPostgres on the GPU just screams HSA http://hsafoundation.com http://hsafoundation.com and more importantly AMD and their upcoming Kaveri processor. which will significantly reduce the latency of performing GPU calculations since the CPU and GPU will share the same cache.
- agnsaft 13y agoThis is cool, however, Postgres could probably achieve higher performance without the GPU as well if they added concurrency on the CPU for certain type of operations (e.g. aggregation, sorting, etc). That would be a killer feature.
- jaakl 13y agoThe query sample is interesting: it seems to be find objects (locations) nearby given point. But there is CPU-only solution to index the data properly and have 1000x faster response times: PostGIS. So if before you threw more CPU/RAM if you did not know how to make things faster in smart way, then now you throw more GPU. Anyway, for sure there can be cases where GPU -based solution could be better alternative than using traditional existing solutions like specialized indexes.
- rektide 13y agoPgOpenCL takes a different approach- instead of a Foreign Data Wrapper (FDW) that exposes tables, it's a language for writing postgres functions in, where the language is just opencl. http://www.slideshare.net/3dmashup/pgopencl http://www.slideshare.net/3dmashup/pgopencl Alas, Tim's been talking about this for two years now, and as far as we know he's the only one whose ever seen the code.
- Rickasaurus 13y agoVery cool. This should be upvoted more.
- DiabloD3 13y agoI don't understand why people keep writing new code in legacy hardware that most people don't own. Its written in CUDA, not OpenCL. Stop that.
- ketralnis 13y agoYeah, why do people keep solving the problems they have rather than the problems you wish they had? Bastards.
- kyhoolee 13y agointeresting, I'm also currently working with GPU