12 ms·
LibreCUDA – Launch CUDA code on Nvidia GPUs without the proprietary runtime
- JonChesterfield 2y agoVery nice! That's essentially all I want from a cuda runtime. It should be possible to run llvm libc unit tests against this, which might then justify a corresponding amd library that does the same direct to syscall approach.
- snihalani 2y agoFor a non cuda n00b, what problem does this solve?
- heyoni 2y agoLike anything open source it allows you to know and see exactly what your machine is doing. I don’t want to speculate too much but I remember there being discussions around whether or not nvidia could embed licensing checks and such at the firmware level.
- samstave 2y ago> licensing checks and such at the firmware level. Could you imaging an age where the NVIDIA firmware does LLM/AI/GPU license checking before it does operations on your vectors? (Hello Oracle on SUN e650, My old Friend) ((Worse would be a DRM check against deep-faking or other Globalist WEF Guardrails)) ((oracle had(has) an age olde function where if you bought a license for a single proc and threw it inot a dual proc sun enterprise server with an extra proc or so - it knew you have several hundred K to spend on an additional e650 so why not have an extra ~$80K for an additional oracle proc license. Rather than make the app actually USE the additional proc - as there were no changes to oracles garbage FU Maxwell))
- IntelMiner 2y ago"Globalist WEF Guardrails" Tell us what you really feel
- samstave 2y agoIf you use all the GPTs enough - you'll see them clear as day... And by saying "Tell us how you really feel" reveals, you may not have thought of The Implications of the current state of AI. (I can give you a concrete example of the WEF guardrails: I have a LBB of some high profile names that are all related around a specific person, then I wanted to see how they were related to one another from a publicly available data-set "that which is searchable on the open internet" And several GPTs stated "I do not feel comfortable revelaing the connections between these people without their consent" I was trying to get a list of public record data for whom the owners and affiliates of shared companies were... If you go down political/financial/professional rabbit holes using various data-mining techniques with augmenting searches and connections via public GPTs (paid even) -- You see the guardrails real fast (hint - they invlove power, money, and particular names and organizations that you hit guardrails against)
- kelnos 2y agoI don't necessarily disagree with your overall point (I don't know much about it either way), but I'm not sure your example does a great job illustrating it. If you tried the same thing, but with non-high-profile names, would it give you the same response? If so, the charitable (and probably correct) interpretation is that this is a general privacy guardrail, not one that's there to protect the powerful/rich.
- dmnmnm 2y ago> If so, the charitable (and probably correct) interpretation is that this is a general privacy guardrail, not one that's there to protect the powerful/rich Considering that some of the champions behind machine learning, like Google, are companies that made a living out of violating your privacy just to serve more ads to your eyeballs.. I wouldn't be so charitable. Tech bros have an inherent disregard for the privacy of others or for author rights for that matter. Was anyone asked if their art could be used to train their replacement? Power for me, not for thee.
- Q6T46nT668w6i3m 2y agoQuadro?
- cyberpunk 2y agoIt was even worse than that. Even if you created a resource pool with only the one CPU on a dual system they wanted licenses for both as you could “potentially” use both CPUs. On VMware they extended this to every CPU in the cluster. A gigantic shower of absolute grifters.
- queuebert 2y agoTwo obvious problems that come to mind are 1. Replacing the extremely bloated official packages with lightweight distribution that provides only the common functionality. 2. Paving the way for GPU support on *BSD.
- einpoklum 2y agoIt doesn't solve problem (1.) ; even when complete, this will replace the CUDA driver and its associated library - which is a very small part of CUDA. As for (2.) - this is just CUDA, not GPU use in general. I wonder whether nouveau is relevant for BSDs (I have no idea...)
- jstanley 2y agoDo you still need to be running the proprietary nvidia graphics driver, or is that completely unrelated?
- tptacek 2y agoPresumably yes, if it functions through an ioctl interface.
- kaladin-jasnah 2y agoYou will need an NVIDIA driver (the README says as much), be it the proprietary or open source modules. Looks like this is performing RM (Resource Manager, which is the low-level API that is used to communicate with the NVIDIA proprietary driver using ioctl calls) API calls. If you look in the src/nvidia directory, many of the header files are RM API call header files from the NVIDIA open source kernel module, containing various structures and RM API command numbers (not sure if this is the official term). Fun thing, the open source modules takes some proprietary things and moves them to the GSP firmware. Incidentally, the open source modules actually communicate with the GSP firmware using the RM API as well. This understanding may be correct, but now instead of some RM calls being handled in kernel space they are forwarded to the firmware and handled there.
- KeplerBoy 2y agoWhat's a CUDA elf file? Is it binary SASS code, so one would still need a open source ptxas alternative?
- mike64_t 2y agoYes, the Nvidia SASS ISAs are not documented and emitting them is non trivial due to Nvidia GPUs not handling pipeline hazards in Hardware and requires the compiler to correctly schedule instructions to avoid race conditions. The only available code that does this can be found in MESA, but even they say "//this is bs and we know it" in a comment above their instruction latencies, which you also can't easily figure out. Replacing ptxas is highly non trivial. I will attempt to do so, but it increasingly looks like ptxas is here to stay. I started working on a nvcc + cuda SDK replacement which already works surprisingly well for a day of work. However, ptxas is in my sight. But I know this is something that to my knowledge nobody that wasn't fed Nvidia documentation under license has ever successfully accomplished.
- shmerl 2y agoSince ZLUDA was taken down (by request from AMD of all parties), it would be better to have some ZLUDA replacement as a general purpose way of breaking CUDA lock-in. I.e. something not tied to Nvidia hardware.
- KeplerBoy 2y agoThat's a problem on a different level of the CUDA stack. Having a compiler that takes a special C++ or python dialect and compiles it to GPU suitable llvm-ir and then to a GPU binary is one thing (and there's progress on that side: triton, numba, soonish mojo), being able to launch that binary without going through the nvidia driver is another problem.
- shmerl 2y agoYeah, the latter one is more useful for effective lock-in breaking.
- codedokode 2y agoCannot Vulcan compute be used to execute code on GPU without relying on proprietary libs? Why not?
- Conscat 2y agoYou still require a Vulkan driver to do anything with it. Until last year, Nvidia hardware required a proprietary Vulkan driver (prior to Nvvk), and anything pre-Pascal still requires that.
- codedokode 2y agoYes but you can use any GPU with Vulkan, not only NVIDIA.
- cyber_kinetist 2y agoVulkan Compute's semantics are limited by SPIR-V and thus cannot implement all of the features CUDA provides (ex. there is no proper notion of a "pointer") Also it's much more convenient to use plain C++ rather than a custom shading language, especially if you're writing complex numerical code or need some heavy templated abstractions to do powerful stuff. And the CUDA tooling itself is just much easier to use compared to Vulkan, with its seamless integration of host / device code.
- wackycat 2y agoI have limited experience with CUDA but will this help solve the CUDA/CUDNN dependency version nightmare that comes with running various ML libraries like tensorflow or onnx?
- bstockton 2y agoMy experience, over 10 years building models with libraries using CUDA under the hood, this problem has nearly gone away in the past few years. Setting up CUDA on new machines and even getting multi GPU/nodes configuration working with NCCL and pytorch DDP, for example, is pretty slick. Have you experienced this recently?
- jokethrowaway 2y agoyes, especially if you are trying to run various different projects you don't control some will need specific versions of cuda right now I masked cuda from upgrades in my system and I'm stuck on an old version to support some projects I also had plenty of problems with gpu-operator to deploy on k8s: that helm chart is so buggy (or maybe just not great at handling some corner cases? no clue) I ended up swapping kubernetes distribution a few times (no chance to make it work on microk8s, on k3s it almost works) and eventually ended up installing drivers + runtime locally and then just exposing through containerd config
- trueismywork 2y agoThat's torches bad software distribution problem. No one can solve it apart from torch distributors
- amelius 2y agoBy the way, can anyone explain why libcudnn takes on the order of gigabytes on my harddrive?
- lldb 2y agoPrimarily because it has specialized functions for various matrix sizes which are selected at runtime.
- greenavocado 2y agoThe authors better start thinking about the trademark infringement notice coming their way
- allan_s 2y agosomething like kudo ?
- greenavocado 2y agohttps://news.ycombinator.com/item?id=41195332 https://news.ycombinator.com/item?id=41195332
- gavindean90 2y agokudo seems to be in keeping with your comment. I am not sure what you are getting at.
- mango3355 2y ago[flagged]
- greenavocado 2y agoIt's far too similar.
- kkielhofner 2y agoNaming anything is hard and I don’t have better suggestions but when you’re doing something that’s already poking at something a big corp holds dearly hitting on trademark while you’re at it makes it really easy for them.
- greenavocado 2y agoYou can't have the CUDA substring in the name or anything a court would deem potentially confusing. Even if "CUDA" wasn't registered, using a similar name could be seen as an attempt to pass off the product as affiliated with or endorsed by NVIDIA. The similarity in names could be construed as an attempt to unfairly benefit from NVIDIA's reputation and market position. If the open-source project implements techniques or methods patented by NVIDIA for CUDA, it could face patent infringement claims. If CUDA is considered a famous mark, using a similar name could be seen as diluting its distinctiveness, even if the products aren't directly competing. If domain names similar to CUDA-related domains are registered for the project, this could potentially lead to domain dispute issues. It's a huge can of worms.
- Onavo 2y agoWhat about the extra bits like CuDNN?
- deleted 2y ago[deleted]
- m3kw9 2y agoWhy is there a need to do this?
- curious_cat_163 2y ago1. To learn, how. 2. Nvidia needs to be challenged with OSS. They are far too big to be left alone. 3. To have some fun.
- nasretdinov 2y agoSuch a missed opportunity to call it CUDA Libre...
- phoronixrly 2y agoUnfortunately both would seem to be infringing on nvidia's trademark... We just can't have nice things..
- mrbungie 2y agoThen call it Cuba Libre.
- phoronixrly 2y agoSounds too similar to cuda.. :(
- chii 2y agocuba libre is a drink name, which won't infringe on trademark, as it's not really possible to confuse it with CUDA.
- diggan 2y agoOr go even further, call it "Culo Libre", only two letters away anyways.
- pezezin 2y agoIf nVidia can release a library called cuLitho, I don't see why not.
- deleted 2y ago[deleted]
- londons_explore 2y agoA trademark isn't a total prohibition on using someone else's name. You can still use their name where there is no likelihood of consumer confusion. Obviously many companies choose not to to avoid a lawsuit over the issue - but it's unlikely NVidia would win over this method name.
- daft_pink 2y agoI think the point of open cuda is to run it on non NVIDIA gpus. Once you have to buy NVIDIA gpus what’s the point. If we had true you competition I think it would be far easier to buy devices with more vram and thus we might be able to run llama 405b someday locally. Once you already bought the NVIDIA cards what’s the point
- londons_explore 2y agoThe NVidia software stack has the "no use in datacenters" clause. Is this a workaround for that?
- why_only_15 2y agoSpecifically the clause is that you cannot use their consumer cards (e.g. RTX 4090) in datacenters.
- candiddevmike 2y agoThat's why we run all of our ML workloads in a distributed GPU cluster located in every employee's house
- pplante 2y agoThe bonus is free heating for every employees household!
- rurban 2y agoFree cooling also. You cannot really run a big GPU with external cooling. I needed rather big 15cm isolated cooling tubes to get the heat out of the building.
- seniorThrowaway 2y agoyou joke but I've thought about doing this
- snvzz 2y agoMoving to HiP on LibreCUDA should probably be the first step for projects that are dependent on CUDA to gain platform freedom.
- fishcrackers 2y ago[dead]
- koimaster 2y ago[dead]
- georgehotz 2y agoCool to see one of these in C, particularly if it can be binary compatible. Why not s/libreCuInit/cuInit? If you are interested in open source runtimes, tinygrad has them in Python for both AMD and NVIDIA, speaking directly to the kernel through ioctls and poking the command queues. https://github.com/tinygrad/tinygrad/blob/master/tinygrad/runtime/ops_nv.py https://github.com/tinygrad/tinygrad/blob/master/tinygrad/ru... https://github.com/tinygrad/tinygrad/blob/master/tinygrad/runtime/ops_amd.py https://github.com/tinygrad/tinygrad/blob/master/tinygrad/ru...
- ZoomerCretin 2y agoIncredible! Any plans to support SASS instructions for Nvidia GPUs, or only PTX?
- georgehotz 2y agoWe'll get there as we push deeper into assemblies. RDNA3 probably first, since it's documented and a bit simpler.
- mike64_t 2y agoHow do you plan on finding instruction latencies for eg. sm_89?
- JonChesterfield 2y agoThat's interesting. This looks like you've bypassed the rocm userspace stack entirely. I've been looking for justification to burn libhsa.so out of the dependency graph for running llvm compiled kernels on amdgpu for ages now. I didn't expect roct to be similarly easy to drop but that's a clear sketch of how to build a statically linked freestanding x64 / gcn blob. Excellent. (I want a reference implementation of run-simple-stuff which doesn't fall over because of bugs in libhsa so that I know whatever bug I'm looking at is in my compiler / the hardware / the firmware)
- georgehotz 2y ago
- mattiasfestin 2y agoCould this cause legal issues with tech export bans and restrictions in Nvidia’s driver, specifically if LibreCUDA circumvents those restrictions?
- mango3355 2y ago[flagged]
- mango3355 2y ago[flagged]
- gnulinux 2y agoDoes it make sense to buy Nvidia GPUs as a linux user in 2024 anyway? I thought Nvidia has abysmal linux support, if you don't have Nvidia GPU what's the point of LibreCUDA?
- jfarina 2y agoIt does if you work in data science/machine learning.
- einpoklum 2y agoThis is a nice initiative, but for now it covers almost none of the API. It needs to be about 100x bigger or so, in terms of coverage, before we can consider using it except as a proof-of-concept.
- mango3355 2y ago[flagged]