Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
samnco
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Opportunistically mining cryptocurrencies in Kubernetes
(medium.com)
2 points
by
samnco
9y ago
|
0 comments
2.
▲
Simple Perf Test for Tensorflow Serving on K8s
(medium.com)
3 points
by
samnco
9y ago
|
0 comments
3.
▲
by
samnco
9y ago
Hello, OK I gave it a try and you are absolutely right. For the nvidia-smi, I could run it the /dev/nvidia0, which is cool. I was also able to run it unprivileged. I guess my mistake was to believe the example from the docs and no
4.
▲
by
samnco
9y ago
Aaah that is interesting. Let me dive into this later today and test my charts without that. It would actually make my life way easier for charting. I got that from a very early stage work and never questioned it again (the /dev stuf
5.
▲
by
samnco
9y ago
privileged containers are required for the GPU to be shared with the containers. By default, the bundle come with a "auto" tag, which will activate privileged containers just when GPUs are detected. You can enforce "false&quo
6.
▲
by
samnco
9y ago
I have not, but it is technically possible. the PSU is the double 1100W with the GPU enablement kit. Up to 4x PCI-x 16x full speed. Also up to 1.5TB RAM, and 8x 3.5" HDD or 16x 2.5". I didn't go this far though ($$$...)
7.
▲
by
samnco
9y ago
Here you go :) https://drive.google.com/file/d/0B1CCk51NQ4koSmkxSmxWb1E5Y0E... Replicating is not very hard. You need a lightweight x86 machine for MAAS, which takes ~20min to install, one VLAN for the iDRAC (IPMI
8.
▲
by
samnco
9y ago
Yes, each GPUs has a 4x -> 16x and a 4x-4x extender, in addition to the m.2 -> PCI-e 4x adapter. So many potential failure points in there. The sole use case is CUDA. Essentially I wanted a portable cluster with GPUs and that did the
9.
▲
by
samnco
9y ago
I don't know. Maybe the make of the extenders isn't very good, I saw other people with similar issues. The PSU is the Corsair AX1500i (1500W), with 10x lines for GPUs. It's robust on paper, didn't have any problem with j
10.
▲
by
samnco
9y ago
Actually, it was a fun DIY project I did a while ago. You can read about it here: https://hackernoon.com/installing-a-diy-bare-metal-gpu-clust... It works, but the GPUs aren't very stable at 4x vs. a normal 16x.
11.
▲
by
samnco
9y ago
Typically, you would have a set of "helm charts" (packages) for your application(s). So deploying, without data, would be something like a sequence of "helm install app-appId --values /path/to/config/for&
12.
▲
by
samnco
9y ago
Ah you are right, I forgot about glusterfs. My bad. Canonical at this stage only supports Ceph commercially, but it doesn't mean GlusterFS is not a good option. I haven't tried it myself, so can't tell. Anyone?
13.
▲
by
samnco
9y ago
You have several options for this. If it is non HA, then you can pin a RC to a specific node, and use hostpath storage. if the container fails, it will always respawn on the same node, maximizing uptime and also having max capacity from you
14.
▲
by
samnco
9y ago
You have several options: * Run Ceph in separate nodes and connect it to the cluster. With Juju, you can do that from the bundle, as Ceph is also a supported workloads. This gives you scale for storage * Run Ceph within the cluster with a H
15.
▲
by
samnco
9y ago
Ahah, good point. Really the ETH stuff was "because I can". But in the same charts repository you will find a Tensorflow chart. My previous series of blogs [0] was about exactly that. A nice addition as well for compute intensive
16.
▲
by
samnco
9y ago
There are a few killer features that you would benefit at any size and that I really love * self healing: when you create a deployment/replica set. it will be maintained at all cost, so if the app has a memory leak or anything goes wro
17.
▲
by
samnco
9y ago
Canonical will officially support GPUs when they lands GA upstream. The flag is beta as of now in the Canonical Distribution of Kubernetes. Paying customers either for the managed or supported solutions get a best effort for GPU, and this