6 ms·
Genuinely, why? - For new code, all of the functions [here](https://libc.llvm.org/gpu/support.html#libc-gpu-support https://libc.llvm.org/gpu/support.html#libc
by C-programmer 2y ago
Genuinely, why?
- For new code, all of the functions [here](https://libc.llvm.org/gpu/support.html#libc-gpu-support https://libc.llvm.org/gpu/support.html#libc-gpu-support) you can do without just fine.
- For old code:
* Your project is large enough that you are likely use using an unsupported libc function somewhere.
* Your project is small enough that you would benefit from just implementing a new kernel yourself.
I am biased because I avoid the C standard library even on the CPU, but this seems like a technology that raises the floor not the ceiling of what is possible.
- tredre3 2y ago> this seems like a technology that raises the floor not the ceiling of what is possible. In your view, how is making GPU programming easier a bad thing?
- convolvatron 2y agothat's clearly not a bad thing. however encouraging people to run mutating, procedural code with explicit loops and aliasing maybe isn't the right path to get there. particularly if you just drag forward all the weird old baggage with libc and its horrible string conventions. I think any programming environment that treats a gpu as a really slow serial cpu isn't really what you want(?)
- quotemstr 2y agoWhat if it encourages people to write parallel and functional code on CPUs? That'd be a good thing. Influence works both ways. The bigger problem is that GPUs have various platform features (shared memory, explicit cache residency and invalidation management) that CPUs sadly don't yet. Sure, you could expose these facilities via compiler intrinsics, but then you end up code that might be syntactically valid C but is alien both to CPUs and human minds
- fc417fc802 2y ago> is alien both to CPUs and human minds On the contrary I would love that. The best case scenario in my mind is being able to express the native paradigms of all relevant platforms while writing a single piece of code that can then be compiled for any number of backends and dynamically retargeted between them at runtime. It would make debugging and just about everything else SO MUCH EASIER. The equivalent of being able to compile some subset of functions for both ARM and x86 and then being able to dynamically dispatch to either version at runtime, except replace ARM with a list of all the GPU ISAs that you care about.
- JonChesterfield 2y agoOne thing this gives you is syscall on the gpu. Functions like sprintf are just blobs of userspace code, but others like fopen require support from the operating system (or whatever else the hardware needs you to do). That plumbing was decently annoying to write for the gpu. These aren't gpu kernels. They're functions to call from kernels.
- almostgotcaught 2y ago> One thing this gives you is syscall on the gpu i wish people in our industry would stop (forever, completely, absolutely) using metaphors/allusions. it's a complete disservice to anyone that isn't in on the trick. it doesn't give you syscalls. that's impossible because there's no sys/os on a gpu and your actual os does not (necessarily) have any way to peer into the address space/schedular/etc of a gpu core. what it gives you is something that's working really really hard to pretend be a syscall: > Traditionally, the C library abstracts over several functions that interface with the platform’s operating system through system calls. The GPU however does not provide an operating system that can handle target dependent operations. Instead, we implemented remote procedure calls to interface with the host’s operating system while executing on a GPU. https://libc.llvm.org/gpu/rpc.html https://libc.llvm.org/gpu/rpc.html.
- quotemstr 2y agoIt's a matter of perspective. If you think of the GPU as a separate computer, you're right. If you think of it as a coprocessor, then the use of RPC is just an implementation detail of the system call mechanism, not a semantically different thing. When an old school 486SX delegates a floating point instruction to a physically separate 487DX coprocessor, is it executing an instruction or doing an RPC? If RPC, does the same instruction start being a real instruction when you replace your 486SX with a 486DX, with an integrated GPU? The program can't tell the difference!
- almostgotcaught 2y ago> It's a matter of perspective. If you think of the GPU as a separate computer, you're right. this perspective is a function of exactly one thing: do you care about the performance of your program? if not then sure indulge in whatever abstract perspective you want ("it's magic, i just press buttons and the lights blink"). but if you don't care about perf then why are you using a GPU at all...? so for people that aren't just randomly running code on a GPU (for shits and giggles), the distinction is very significant between "syscall" and syscall. people who say these things don't program GPUs for a living. there are no abstractions unless you don't care about your program's performance (in which case why are you using a GPU at all).
- nickysielicki 2y agohttps://developer.nvidia.com/blog/simplifying-gpu-application-development-with-heterogeneous-memory-management/#unified_memory_after_hmm https://developer.nvidia.com/blog/simplifying-gpu-applicatio...
- deleted 2y ago[deleted]
- JonChesterfield 2y ago> Genuinely, why? > ... this seems like a technology that raises the floor not the ceiling of what is possible. The root cause reason for this project existing is to show that GPU programming is not synonymous with CUDA (or the other offloading languages). It's nominally to help people run existing code on GPUs. Disregarding that use case, it shows that GPUs can actually do things like fprintf or open sockets. This is obvious to the implementation but seems largely missed by application developers. Lots of people think GPUs can only do floating point math. Especially on an APU, where the GPU units and the CPU cores can hammer on the same memory, it is a travesty to persist with the "offloading to accelerator" model. Raw C++ isn't an especially sensible language to program GPUs in but it's workable and I think it's better than CUDA.
- rbanffy 2y ago> Lots of people think GPUs can only do floating point math. IIRC, every Raspberry Pi is brought up by the GPU setting up the system before the CPU is brought out of reset and the bootloader looks for the OS. > it is a travesty to persist with the "offloading to accelerator" model. Operating systems would need to support heterogeneous processors running programs with different ISAs accessing the same pools of memory. I'd LOVE to see that. It'd be extremely convenient to have first-class processes running on the GPU MIMD cores. I'm not sure there is much research done in that space. I believe IBM mainframe OSs have something like that because programmers are exposed to the various hardware assists that run as coprocessors sharing the main memory with the OS and applications.
- als0 2y ago> I'm not sure there is much research done in that space. There is. And the finest example I can think of is Barrelfish https://barrelfish.org https://barrelfish.org
- rbanffy 2y agoInteresting - it resembles a network of heterogeneous systems that can share a memory space used primarily for explicit data exchange. Not quite what I was imagining, but probably much simpler to implement than a Unix where the kernel can see processes running on different ISAs on a shared memory space. I guess hardware availability is an issue, as there aren't many computers with, say, an ARM, a RISC-V, an x86, and an AMD iGPU sharing a common memory pool. OTOH, there are many where a 32-bit ARM shares the memory pool with 64-bit cores. Usually the big cores run applications while the small ARM does housekeeping or other low-latency task.