4 ms·
I have to use containers with nvidia-docker because NVIDIA so consistently and relentlessly breaks things without so much as a glance at backward compatibility.
by bbatsell 6y ago
I have to use containers with nvidia-docker because NVIDIA so consistently and relentlessly breaks things without so much as a glance at backward compatibility.
- dijksterhuis 6y agoI moved our Deep Learning servers over to Docker images + JupyterHub DockerSpawners recently because maintaining all the various version dependencies between frameworks was an absolute PITA. Images are publicly available here in case anyone else needs something similar: https://hub.docker.com/u/uodcvip https://hub.docker.com/u/uodcvip
- etaioinshrdlu 6y agoThe annoying thing is that nvidia-docker is still not great. You still have to deal with the driver installed outside the container, and it makes a big difference. Furthermore it seems like even the CUDA runtime is typically not installed in the container, but rather injected in by the nvidia-docker container runtime. It is not fun to deal with.
- dmm 6y agoYou don't have to use nvidia-docker to use cuda with docker. I made my own cuda containers based on Debian and pass the devices to the docker run command. I mount the libcuda and libnvidia libraries as volumes. I think that's what you mean by injecting the runtime. Here's an example Dockerfile: https://github.com/dmm/docker-debian-cuda/blob/master/Dockerfile https://github.com/dmm/docker-debian-cuda/blob/master/Docker... And here's an example docker run command: docker run -it --rm $(ls /dev/nvidia* | xargs -I{} echo '--device={}') $(ls /usr/lib/x86_64-linux-gnu/{libcuda,libnvidia}* | xargs -I{} echo '-v {}:{}:ro') dmattli/debian-cuda:10.0-buster-debug /bin/bash Verbose but it works fine. You still have to have the nvidia driver installed on the host system.
- TheGuyWhoCodes 6y agoI'm never sure of the relation between the driver, nvidia-docker and the container with a specific cuda version. Last time I tried it the cuda inside the container tough it was using some old driver version while a much newer version was installed on the host. So I had to manual install the older version, not sure where the issue was but maybe it was because I was using the deprecated nvidia-docker version 2 which is still needed to pass gpu resources to containers run inside kubernetes.
- markus92 6y agoWe use Singularity as our container provider for the exact same reason. For now it has worked great and driver/CUDA updates haven't broken any containers yet over the past 2 years. Something good about Singularity (which I bet you could also do with Docker) is that it automatically binds the right NVIDIA stuff into the container. It also works fine unprivileged :)