6 ms·
I noticed CUDA 11.0 was almost ready for release last week when I went to install CUDA and the default download page linked to the 11.0 Release Candidate. The 1
by usmannk 6y ago
I noticed CUDA 11.0 was almost ready for release last week when I went to install CUDA and the default download page linked to the 11.0 Release Candidate. The 10.1 and 10.2 links were buried behind a link off to the side labeled "legacy". The thing is, no library you use is going to be supporting the CUDA 11.0 RC, that's ridiculous. For example, Pytorch stable is on 10.2 and Tensorflow only goes up to 10.1.
This is generally indicative of how poorly organized the CUDA documentation and installation instructions are. The Conda dependency manager has made this a lot easier recently. Especially by, e.g., providing pytorch binaries. Though if you want to use packages like NVIDIA Apex for mixed precision DL[0] you're going to be in for a huge headache trying to compile torch from source while also managing your cuda and nvcc version, which sometimes must be the same but sometimes can not be![1]
[0] Yes, I'm aware that Apex was very recently brought into torch but it seems that the performance issues haven't been ironed out yet.
[1] https://stackoverflow.com/questions/53422407/different-cuda-versions-shown-by-nvcc-and-nvidia-smi https://stackoverflow.com/questions/53422407/different-cuda-...
- liuliu 6y agoYeah. To make the matter worse, they have updated libnccl-dev from their apt repo to be CUDA-11 based a few weeks ago. That breaks my CUDA app (because it is still on 10.2 and not compatible) in interesting ways. apt-hold libnccl-dev for a while and waiting for this release.
- jjoonathan 6y agoYeah, and the CUDA 10.0 official Visual Studio demo project build was broken for... looks like a year, at least, because they didn't want to populate the toolkit path. NVidia, you're better than this. https://forums.developer.nvidia.com/t/the-cuda-toolkit-v10-0-directory-does-not-exist/65821/6 https://forums.developer.nvidia.com/t/the-cuda-toolkit-v10-0... > The Conda dependency manager has made this a lot easier Yeah but conda is "Let's do dependency management with a SAT solver, it'll be great!" On a good day, it's just slow. On a bad day, the SAT solver spins for hours before failing to converge. On a really bad day, the SAT solver does something "clever." I've had a couple of really bad days this year. I'm really starting to not like conda very much.
- ajtulloch 6y agoYou might find https://github.com/TheSnakePit/mamba https://github.com/TheSnakePit/mamba useful, especially if you are slowed down by package resolution.
- jjoonathan 6y agoThat looks worth a look for sure!
- techwizrd 6y agoConda's SAT solver for dependency management is the bane of my existence. For pip-installable packages, I'll almost always turn to pip rather than conda even when in a conda environment.
- jabl 6y agoMy biggest gripe with the conda depencency manager is that it doesn't keep track of which packages own which files, and if multiple packages own the same file the last one to be installed will happily scribble over whatever was there before. With hilarious results, of course. This means that keeping a conda installation up to date is often very tricky, when upgrading you frequently have to uninstall and reinstall some packages. It works better if you start from scratch with a requirements.yml file.
- pjc50 6y ago> "Let's do dependency management with a SAT solver, it'll be great!" Debian managed something like this over 20 years ago in dpkg. But somehow people must keep reinventing the wheel.
- jjoonathan 6y agoI thought the debian SAT solver was a maintainer tool rather than something that ran every time? In any case, conda's implementation is really quite awful by comparison and they would have been well served by copying something that works instead of building something that doesn't.
- 6y ago
- bbatsell 6y agoI have to use containers with nvidia-docker because NVIDIA so consistently and relentlessly breaks things without so much as a glance at backward compatibility.
- dijksterhuis 6y agoI moved our Deep Learning servers over to Docker images + JupyterHub DockerSpawners recently because maintaining all the various version dependencies between frameworks was an absolute PITA. Images are publicly available here in case anyone else needs something similar: https://hub.docker.com/u/uodcvip https://hub.docker.com/u/uodcvip
- etaioinshrdlu 6y agoThe annoying thing is that nvidia-docker is still not great. You still have to deal with the driver installed outside the container, and it makes a big difference. Furthermore it seems like even the CUDA runtime is typically not installed in the container, but rather injected in by the nvidia-docker container runtime. It is not fun to deal with.
- dmm 6y agoYou don't have to use nvidia-docker to use cuda with docker. I made my own cuda containers based on Debian and pass the devices to the docker run command. I mount the libcuda and libnvidia libraries as volumes. I think that's what you mean by injecting the runtime. Here's an example Dockerfile: https://github.com/dmm/docker-debian-cuda/blob/master/Dockerfile https://github.com/dmm/docker-debian-cuda/blob/master/Docker... And here's an example docker run command: docker run -it --rm $(ls /dev/nvidia* | xargs -I{} echo '--device={}') $(ls /usr/lib/x86_64-linux-gnu/{libcuda,libnvidia}* | xargs -I{} echo '-v {}:{}:ro') dmattli/debian-cuda:10.0-buster-debug /bin/bash Verbose but it works fine. You still have to have the nvidia driver installed on the host system.
- TheGuyWhoCodes 6y agoI'm never sure of the relation between the driver, nvidia-docker and the container with a specific cuda version. Last time I tried it the cuda inside the container tough it was using some old driver version while a much newer version was installed on the host. So I had to manual install the older version, not sure where the issue was but maybe it was because I was using the deprecated nvidia-docker version 2 which is still needed to pass gpu resources to containers run inside kubernetes.
- fock 6y agospeaking of poorly organized: did they fix their embedded dependencies yet for glibc 2.30 in actual tagged releases of tensorflow?