Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
g413n
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
g413n
1y ago
yeah we're very interested in trying toploaders, we'll do a test rack next time we expand and switch to that if it goes well. w.r.t. testing the main thing we did was try to buy a bit from each supplier a month or two ahead of tim
32.
▲
by
g413n
1y ago
yeah colo help has been great, we had a power blip and without any hassle they covered the cost and installation of UPSes for every rack, without us needing to think abt it outside of some email coordination.
33.
▲
by
g413n
1y ago
not caring about redundancy/reliability is really nice, each healthy HDD is just the same +20TB of pretraining data and every drive lost is the same marginal cost.
34.
▲
by
g413n
1y ago
fwiw our first test rack has been up for about a year now and the full cluster has been operational for training for the past ~6 months. having it right down the block from our office has been incredibly helpful, I am a bit worried abt what
35.
▲
by
g413n
1y ago
yeah we just have the 100gig link, atm that's about all the gpu clusters can pull but we'll prob expand bandwidth and storage as we scale. I guess worth noting that we do have a bunch of 4090s in the colo and it's been super
36.
▲
by
g413n
1y ago
someone has to go and power-cycle the machines every couple months it's chill, that's the point of not using ceph
37.
▲
by
g413n
1y ago
the doodles are great
38.
▲
by
g413n
1y ago
in a datacenter context failure rates are just a remote-hands recurring cost so it's not too bad with front-loaders e.g. have someone show up to the datacenter with a grocery list of slot indices and a cart of fresh drives every few mo
39.
▲
by
g413n
1y ago
7.5k for zayo 100gig so that's like half of the MRC
40.
▲
by
g413n
1y ago
No mention of disk failure rates? curious how it's holding up after a few months