12 ms·
Libcu++: Nvidia C++ Standard Library
- lars 6y agoIt really is a tiny subset of the C++ standard library, but I'm happy to see they're continuing to expand it: https://nvidia.github.io/libcudacxx/api.html https://nvidia.github.io/libcudacxx/api.html
- roel_v 6y agoYeah, really tiny... At first I thought 'wow this is a game changer', but then I looked at your link and thought 'what's the point?'. Can someone explain what real problems you can solve with just the headers in the link above?
- happyweasel 6y agoIt runs on the GPU?
- jpz 6y agoI guess that the point is that when writing CUDA code (which looks like C++), you can use these libraries which are homogenous with CPU code. Looking at the functions, chrono/barrier etc require CPU level abstractions, so using the STL versions (which are for the CPU) aren't going to work really.
- TillE 6y agoI would have expected the <algorithm> header, but instead...synchronization primitives? std::chrono? I'm completely baffled about how that would be useful, but that's probably because I know very little about CUDA.
- blelbach 6y agoGPUs are parallel processors. So, yes, synchronization primitives are the highest priority. We focused on things that require /different/ implementations in host and device code. The way you implement std::binary_search is the same in host and device code. Sure, we can stick `__host__ __device__` on it for you, but it's not really high value. Synchronization primitives? Clocks? They are completely different. In fact, the machinery that we use to implement both the synchronization primitives and clocks has not previously been exposed in CUDA C++.
- blelbach 6y agohttps://youtu.be/75LcDvlEIYw https://youtu.be/75LcDvlEIYw https://youtu.be/VogqOscJYvk https://youtu.be/VogqOscJYvk
- shaklee3 6y agoNvidia has had many members on the c++ standards committee for a while.
- blelbach 6y agoToday, you can use the library with NVCC, and the subset is small. We'll be focusing on expanding that subset over time. Our end goal is to enable the full C++ Standard Library. The current feature set is just a pit stop on the way there.
- scott31 6y agoA pathetic attempt to lock developers into their hardware.
- daniel-thompson 6y agoI think CUDA itself is the locking attempt; this is just a tiny cherry on top.
- jpz 6y agoThey seem to be pushing the barrier on innovation on GPU compute. It seems a little unfair to call that pathetic, whatever strategic reasons they have to find OpenCL unappetising (which simply enables their sole competitor in truth.) Their decision making seems rational, of course it's not ideal if you're consumer. We would like the ability to bid off NVidia with AMD Radeon. Convergence to a standard has to be driven by the market, but it's impossible to drive NVidia there because they are the dominant player and it is 100% not in their interests. It doesn't mean they're a bad company. They are rational actors.
- my123 6y agoWith nvc++, they are converging towards a standardised source code standard: https://developer.nvidia.com/blog/accelerating-standard-c-with-gpus-using-stdpar/ https://developer.nvidia.com/blog/accelerating-standard-c-wi... However, this notably doesn't cover binaries, which are GPU vendor specific in that case, so AMD for example would have to provide a C++ compiler implementing stdpar for GPUs targeted to their hardware.
- deleted 6y ago[deleted]
- gj_78 6y agoAgree++. They are good at hardware and should stay that way.
- my123 6y agoThe thing is: that hardware isn't very usable without good software, and an easy to use software stack at that. That's what NVIDIA understood and made them what they are today.
- RcouF1uZ4gsC 6y agoFor everyone wondering where are all the data structures and algorithms, vector and several algorithms are implemented by Thrust. https://docs.nvidia.com/cuda/thrust/index.html https://docs.nvidia.com/cuda/thrust/index.html Seems the big addition of the Libcu++ to Thrust would be synchronization.
- blelbach 6y agoYep, that's correct. My team develops Thrust, CUB, and libcu++.
- BoppreH 6y agoUnfortunate name, "cu" it's the most well known slang for "anus" in Brazil (population: 200+ million). "Libcu++" is sure to cause snickering.
- NullPrefix 6y agoThis only affects developers. Limited scope. Wasn't there something related about Microsoft Lumia phones?
- kitd 6y agocf. the Vauxhall Nova car "No va" means "doesn't go" in Spanish.
- andrepd 6y agoAlso Hyundai Kona, "cona" means "cunt" or "pussy" in Portuguese.
- FridgeSeal 6y agoWow, Kona Bikes [0] must have a fun time in Portugal then.. [0] https://konaworld.com/ https://konaworld.com/
- fullstop 6y agoI wonder how kona coffee sells over there.
- sterwill 6y agoI think it's unlikely that Spanish speakers would have been confused about the word "nova" when used as a car name. In Spanish "nova" describes the same astronomical event we call a "nova" in English: a new light in the sky. Additionally Spanish "nuevo" and English "new" seem to share the same root. My point is these words all mean similar things to English- and Spanish-speaking car buyers.
- Mr_lavos 6y agoDoes this mean you can do operations on struct's that live on the GPU hardware?
- shaklee3 6y agoYou have been able to do that for a long time with UVA.
- blelbach 6y agoSince Unified Memory. UVA, or Unified Virtual Addressing, just ensured that a GPU-private object wouldn't have the same address as a CPU-private object.
- shaklee3 6y agoYou're right, sorry. Mixing up terms.
- blelbach 6y agoNot your fault, we don't make it easy. The acroynms are terrible! That's why I typically spell out the full term. My first week at NVIDIA: Me, to very senior engineer: something something UVM. Very senior engineer: What's UVM? Me: Unified Virtual Memory. Very senior engineer: Don't call it that, call it Unified Memory, no abbreviation. TLAs are evil. Me: What's TLA? Very senior engineer: Three letter acronym.
- fanf2 6y ago“Whenever a new major CUDA Compute Capability is released, the ABI is broken. A new NVIDIA C++ Standard Library ABI version is introduced and becomes the default and support for all older ABI versions is dropped.” https://github.com/NVIDIA/libcudacxx/blob/main/docs/releases/versioning.md https://github.com/NVIDIA/libcudacxx/blob/main/docs/releases...
- MichaelZuo 6y agoIt’s interesting that they use the word to broken to describe incompatible machine code. Well if the code is recompiled for each new version then it’s different from the old machine code, that’s by definition. Does any major software vendor support older versions of the ABI or machine code?
- londons_explore 6y agoFamously Microsoft does with Windows. That's how an exe file from 25 years ago can still run today.
- formerly_proven 6y agoRunning 32 bit x86 code on a AMD64 machine is possible on most operating systems which supported both of these, and has probably more to do with AMD64 supporting that execution model.
- londons_explore 6y agoTry that on Linux and you'll find most libraries no longer have the same entry points and that various data structures have changed leading to fun fun crashes... The kernel itself has maintained (mostly) ABI compatibility though.
- formerly_proven 6y agoThat's a "you're holding it wrong" problem, though. Projects like GTK or Qt never claimed they'd be backwards-compatible 26 years (Qt has specific backwards-compatibility API and ABI guarantees and are in my experience pretty diligent about it), so if you want a binary to work for a long time, you have to ship your own versions of these. Libraries like Xlib on the other hand are very stable and much more similar to the Win32 API in that respect. In theory Linux has versioning for libraries, in practice it is never used correctly and useless anyway, since distros generally only keep around one version of everything, so even if you'd link against a specific version (e.g. libfoobar.so.2.21 instead of libfoobar.so.2, which will break if you don't recompile and/or patch the source), it wouldn't exist _anyway_ after a few updates. And that's mostly because distros never promised you'd be able to run binaries built outside their packaging infrastructure anyway; it being common practice and sometimes working doesn't imply it's guaranteed to work. Hence why C applications only linking these "basic" libraries (libc, Xlib, zlib, ...) are regarded as so stable and portable, because they're built and linked against system components which rarely change. (Keep in mind to build this kind of binary on ancient systems, otherwise glibc will make sure it won't work everywhere).
- gj_78 6y agoI really do not understand why a (very good) hardware provider is willing to create/direct/hint custom software for the users. Isn't this exactly what a GPU firmware is expected to do ? Why do they need to run software in the same memory space as my mail reader ?
- dahart 6y agoWhat do you mean about running in the same memory space? Your operating system doesn’t allow that. Is your concern about using host memory? This open source library doesn’t automatically use host memory, users of the library can write code that uses host memory, if they choose to. How would a firmware help me write heterogeneous bits of c++ code that can run on either cpu or gpu?
- gj_78 6y agoIMHO, the question is not that we need code to run on CPUs and GPUs , we do need that, The question is whether the GPU seller has to control both sides. Until I buy a CPU from nvidia I want to keep some kind of independence. When will we be able to use a future riscv-64 CPU with an nvidia GPU ? we will let the answer to nvidia ?
- dahart 6y agoYou can use this library to write code that runs on both risc-v and a GPU! You seem to be pretty confused about what this library is. It’s not exerting any control. It’s open source! It’s strictly optional, and it only allows developers to do something they actually want, to write code that will compile for any type of processor that a modern c++ compiler can target.
- gj_78 6y agoAgain, I see what you mean. I am even against nvidia advising the developers to use such or such C++ library (be it GNU). It is not their role to do that. We need smarter and more shining GPUs from nvidia, not software. I would say .... The hardware must be sold independently of the software ... but it is a bit too complex, I know.
- lionkor 6y ago> Promising long-term ABI stability would prevent us from fixing mistakes and providing best in class performance. So, we make no such promises. Wait NVidia actually get it? Neat!
- matheusmoreira 6y agoThis is an awesome quote... Same argument used by the Linux kernel developers.
- jlebar 6y agoThis is super-cool. For those of us who can't adopt it right away, note that you can compile your cuda code with `--expt-relaxed-constexpr` and call any constexpr function from device code. That includes all the constexpr functions in the standard library! This gets you quite a bit, but not e.g. std::atomic, which is one of the big things in here.
- davvid 6y agoHere's a somewhat related talk from CppCon '19: "The One-Decade Task: Putting std::atomic in CUDA" https://www.youtube.com/watch?v=VogqOscJYvk https://www.youtube.com/watch?v=VogqOscJYvk
- einpoklum 6y ago1. How do we know what parts of the library are usable on CUDA devices, and which are only usable in host-side code? 2. How compatible is this with libstdc++ and/or libcu++, when used independently? I'm somewhat suspicious of the presumption of us using NVIDIA's version of the standard library for our host-side work. Finally, I'm not sure that, for device-side work, libc++ is a better base to start off of than, say, EASTL (which I used for my tuple class: https://github.com/eyalroz/cuda-kat/blob/master/src/kat/tuple.hpp https://github.com/eyalroz/cuda-kat/blob/master/src/kat/tupl... ). ... partial self-answer to (1.): https://nvidia.github.io/libcudacxx/api.html https://nvidia.github.io/libcudacxx/api.html apparently only a small bit of the library is actually implemented.
- blelbach 6y ago> apparently only a small bit of the library is actually implemented. Yep. It's an incremental project. But stay tuned. > I'm somewhat suspicious of the presumption of us using NVIDIA's version of the standard library for our host-side work. Today, when using libcu++ with NVCC, it's opt-in and doesn't interfere with your host standard library. I get your concern, but a lot of the restrictions of today's GPU toolchains comes from the desire to continue using your host toolchain of choice. Our other compiler, NVC++, is a unified stack; there is no host compiler. Yes, that takes away some user control, but it lets us build things we couldn't build otherwise. The same logic applies for the standard library. https://developer.nvidia.com/blog/accelerating-standard-c-with-gpus-using-stdpar https://developer.nvidia.com/blog/accelerating-standard-c-wi... > Finally, I'm not sure that, for device-side work, libc++ is a better base to start off of than, say, EASTL (which I used for my tuple class: https://github.com/eyalroz/cuda-kat/blob/master/src/kat/tupl https://github.com/eyalroz/cuda-kat/blob/master/src/kat/tupl... ). We wanted an implementation that intended to conform to the standard and had deployment experience with a major C++ implementation. EASTL doesn't have that, so it never entered our consideration; perhaps we should have looked at it, though. At the time we started this project, Microsoft's Standard Library wasn't open source. Our choices were libstdc++ or libc++. We immediately ruled libstdc++ out; GPL licensing wouldn't work for us, especially as we knew this project had to exchange code with some of our other existing libraries that are under Apache- or MIT-style licenses (Thrust, CUB, RAPIDS). So, our options were pretty clear; build it from scratch, or use libc++. I have a strict policy of strategic laziness, so we went with libc++.