6 ms·
GPU Offload in Rust: Portable, Safe, and Fast
- bicepjai 2mo agoI write all my code in Rust because I am a Rustacean. In many of my custom LLM inference engine projects, the biggest fight has always been bindings. I don’t want to maintain and write bindings; also, if I use an existing project that provides bindings, then I have to wait for the owner to update or fork it and then maintain it on top. It has been a big headache. Running Rust core on GPU sounds like something I will try from day one. Kudos to the team and will watch it closely.
- deleted 2mo ago[deleted]
- MisterTea 2mo ago> I write all my code in Rust because I am a Rustacean. Why is Rust the only language to have literal acolytes? It's as if people believe they are part of some collective computer religion ushering in the messiah. Weird, man.
- bicepjai 2mo agoHahaha, I have heard that argument. I generally love different programming languages. If you have written code a lot in C/C++ and have battle scars with security and null pointer handling, then Rust is definitely a breath of fresh air. It makes you feel why did I have to be even in that fight? It’s like driving stick your whole life in heavy traffic, and someone just handed you an automatic :) you didn’t realize how much mental overhead that was until it was gone. Respect for addressing something that fundamental at the design level. If something deserves praise, it’s okay to praise it. Give it a try before writing it off as religion.
- MisterTea 2mo ago> If you have written code a lot in C/C++ and have battle scars with security and null pointer handling, then Rust is definitely a breath of fresh air. Whats new is old. Ada was that breath of fresh air for me. Safety through a type system with built in concurrency since 1983. Spark gives you provable safe code. > It’s like driving stick your whole life in heavy traffic, and someone just handed you an automatic :) Eh, depends on the vehicle (weight, gearing, tires). My 2006 Civic was a dog in traffic but my 2002 Pathfinder was a breeze.
- surajrmal 2mo agoThe longer you starve the more delicious the food tastes. Programming language advances were ignored by a significant number of folks for decades. The sudden relief of getting something like rust can leave a strong impact. People also like community and rust has a decent one. It's not wrong for folks to want to feel part of a community they have shared interests or values with.
- stouset 2mo agoIt's cute you think that Rust is the only language that has people who are whole converts.
- MisterTea 2mo agoShow me the proof.
- explodes 2mo agoThe onus isn't on them to correct your incorrect generalizations. It is on you as a person to guide yourself. You can do it, I believe in you.
- pezezin 2mo agoHave you never heard of Pythonistas, or Ruby fans, or Perl fans back in the day, or the how Java (plus a healthy dose of XML) were going to be the One True Language?
- iberator 2mo ago[flagged]
- aabhay 2mo agoGenuinely curious what these constant changes are. In fact I don’t feel they’re moving fast enough to give us improvements to core.
- frollogaston 2mo agoIt was changing quickly in the earliest releases, about 10 years ago. Like one day I pulled our repo and there was new "?" syntax, but that was a feature I'd been wanting anyway. Edit: Oh, async/await was a bigger and more recent one, 2019. I've heard that this wasn't an easy decision for them but was kinda needed.
- aabhay 2mo agoA few years in between releases with a sane versioning system seems fine to me. If you don’t want new features don’t build with that new toolchain. It’s not javascript where you need to support all possible browsers.
- frollogaston 2mo agoTeams will disagree over what toolchain to use, and you will read others' code, so this doesn't dodge the issue. Otherwise there'd be no complaint about C++. I don't think Rust is bloated though, every feature has a very good reason.
- speedstyle 2mo agoWould they disagree, new toolchains work great with old code? And even add 95% of the features to older editions, just not new keywords or inference defaults. I think in C++ people complain about the effort to migrate --std, rather than having to learn variants or concepts. It does add to the complexity of course, and it's useful to agree on a consistent style, I just mean Rust doesn't have the particular issues with everyone having to migrate in lockstep, or anyone needing to for new tools
- Thomashuet 2mo agoThat's promising but did they publish any code? I can't find anything in the abstract.
- supermatt 2mo agoIt is a part of the rust codebase: https://rustc-dev-guide.rust-lang.org/offload/internals.html https://rustc-dev-guide.rust-lang.org/offload/internals.html https://github.com/rust-lang/rust/issues/131513 https://github.com/rust-lang/rust/issues/131513
- deleted 2mo ago[deleted]
- jasonjmcghee 2mo ago> the rust-gpu project has to emulate pointers[8], which we consider a blocking issue for most HPC benchmarks. Why is it a blocking issue? I feel like this is very aligned with the goals of rust-gpu.
- minraws 2mo agoPointers are sort of needed for high performance memory management for HPC targets for existing design patterns, maybe we can think of better solutions down the line but it's hard for me to say anything I just use/abuse CUDA pointers as well.
- adgjlsfhk1 2mo agoJulia has pretty good design heritage for how to deal with this sort of thing. you build the right abstractions and everything works (the main key is making sure the compiler elides bounds checks)
- rfgplk 2mo agoFascinating how many people still overcomplicate offloading to GPUs.
- MBCook 2mo agoHow so? I don’t know anything about this area.
- konradha 2mo agoWhat's the easy way here? Linking CUDA into your Rust binary?
- binsquare 2mo ago:) This might be an relevant read: https://smolmachines.com/engineering/gpu-over-vsock https://smolmachines.com/engineering/gpu-over-vsock
- dc443 2mo agoI just read this tool up and down but still can't conceive of a use case.
- binsquare 2mo agothat's fair the connection I'm making is that both sit on the same interface: the CUDA driver api (module load, launch, memcpy). the paper is about compiling kernels using safe rust instead of CUDA C++, what I did is reimplement the other side of it using rust for the calls from inside a VM get forwarded over vsock to the host driver via api-remoting.
- foltik 2mo agoNo, it’s not relevant to simple GPU offload, stop spam linking your slop blog.
- binsquare 2mo ago
- jheriko 2mo ago[dead]
- maxchisto 2mo agodoes anyone know Mojo well enough to comment how Rust + gpu-offload compares to it?
- giancarlostoro 2mo agoMojo is not fully open sourced yet, but it will eventually be, would be an interesting comparison though.
- maxchisto 2mo agoMojo's OSS status doesn't prevent us from evaluating its memory model, writing and benchmarking kernels in it, etc
- giancarlostoro 2mo agoSure, and I realized after I posted the std lib is opened up, not sure how much of it will reveal the underlying Mojo specifics though.
- Alephinitesimal 2mo agoThe NVIDIA+AMD support is the part I find really interesting. I know OpenMP and SYCL can already target multiple GPU vendors, but doing this while keeping Rust's safety model seems pretty compelling. I'm curious how portable the performance is in practice.
- boywitharupee 2mo agois this mainly about making host binaries self-contained for heterogenous workloads? also, seems like this is mostly targeted towards HPC audience?
- whateverboat 2mo ago> This module is under active development. Once upstream, it should allow Rust developers to run Rust code on GPUs. We aim to develop a rusty GPU programming interface, which is safe, convenient and sufficiently fast by default. This includes automatic data movement to and from the GPU, in a efficient way. We will (later) also offer more advanced, possibly unsafe, interfaces which allow a higher degree of control. I really appreciate the work and the effort that went into this. However, such an approach has previously not really worked for C++ with LLVM offload. Why would it work for Rust?
- winningChild 2mo ago[dead]
- winningChild 2mo ago[dead]
- aw1621107 2mo ago> However, such an approach has previously not really worked for C++ with LLVM offload. Why would it work for Rust? I think that will depend on the exact reason(s) C++ with LLVM offload didn't work out? If Rust differs from C++ in a way that addresses pain points/failure modes/etc. from the C++ attempt, for instance, then perhaps it isn't unreasonable to think Rust could succeed where C++ didn't (c.f., Mozilla's pre-Rust attempts to parallelize Firefox's CSS styling engine). Inversely, if Rust doesn't do things differently in the right way perhaps one might expect the effort to also not work out. Or maybe the problems are entirely non-technical and things could work out in either language.
- erupti 2mo ago[flagged]
- ux266478 2mo ago> However, such an approach has previously not really worked for C++ with LLVM offload. Why would it work for Rust? They're very different languages, with different semantics. Without reading more than the synopsis of the paper, they're 100% leveraging the substructural type system and will have a really tight requirement for you to use a certain kind of Rust code at the CPU/GPU boundary.
- YuechenLi 2mo agoSo... why go through LLVM at all instead of having the MIR target PTX/HIP C directly then? If they really wanted a vendor neutral solution for Rust GPU, that already exists: you write the CPU side code, including buffering, allocation, concurrency, etc through Vulkan binding and consume the compute kernel in SPIR-V from HLSL/GLSL/WGSL etc. As it stands, the way they use Rust here feels more like using it like TypeScript types/interfaces than anything else. Again, the size of most operations that should be done on the GPU is known ahead of time before compilation, so it's very much possible to statically allocate memory at compile time instead of going through all this trouble to write what's essentially a Rust shaped DSL for GPU compute.
- sanxiyn 2mo agoPeople go through trouble to write Python-shaped DSL for GPU compute. We will go "why o why?", but apparently such things are necessary to succeed in the market.
- YuechenLi 2mo agoWell, I suppose fake Python is better than fake C++ at least. But yeah, I think Python's dominance in science/ML will eventually pass, just as FORTRAN and Matlab did before.
- minraws 2mo agoBecause it's convenient? Shader and vulkan semantics can be quite limiting and annoying to write. Maybe it doesn't matter in a post AI world but perhaps it will allow better abstractions. No need to yuck someone else's yum.
- gmueckl 2mo agoThe compute shader side of Vulkan is actually fairly smooth compared to graphics. It's not really that different from low level CUDA or OpenCL. The Vulkan complexities overwhelmingly concern rasterization and raytracing.
- Driftbench 2mo agoOwnership tracking should map well to GPU memory lifetimes. That's one place Rust has a real edge over C++.
- deleted 2mo ago[deleted]
- Arsen-V 2mo ago[flagged]
- andreypk 2mo ago[dead]
- contrahax 2mo agoI’ve been having success doing this with AdaptiveCpp (formerly OpenSYCL) bindings in Rust to get PG OLAP and geospatial workloads to run on a metal/cuda gpus - pure rust the whole way through sounds really nice!
- Archit3ch 2mo ago> multi-vendor GPU compilation framework Technically true, since it supports NVIDIA and AMD. But we have a different definition of portability, if I cannot bring a Metal device and expect it to work.