15 ms·
Building the heap: racking 30 petabytes of hard drives for pretraining
- g413n 1y agoNo mention of disk failure rates? curious how it's holding up after a few months
- ClaireBookworm 1y agogood point
- bayindirh 1y agoThe disk failure rates are very low when compared to decade ago. I used to change more than a dozen disks every week a decade ago. Now it's an eyebrow raising event which I seldom see. I think following Backblaze's hard disk stats is enough at this point.
- gordonhart 1y agoBackblaze reports an annual failure rate of 1.36% [0]. Since their cluster uses 2,400 drives, they would likely see ~32 failures a year (extra ~$4,000 annual capex, almost negligible). [0] https://www.backblaze.com/cloud-storage/resources/hard-drive-test-data https://www.backblaze.com/cloud-storage/resources/hard-drive...
- joering2 1y agoTheir rate will probably be higher since they are utilizing used drives. From the spec: 2,400 drives. Mostly 12TB used enterprise drives (3/4 SATA, 1/4 SAS). The JBOD DS4246s work for either.
- antisthenes 1y agoNot necessarily, since disk failures are typically U-shaped. Buying used drives eliminates the high rate of early failure (but does get you a bit closer to the 2nd part of the U-curve). Typically most drives would become more obsolete before hitting the high failure rate of the right side of the U-curve from longevity (7+ years)
- toast0 1y agoI bet you still have a higher early failure rate because of the stress from transportation, even if there's no funny business. And I expect some funny business because used enterprise drives often come with wiped SMART data, some drives may have been retired by sophisticated clients who decided they were near failure.
- dist-epoch 1y agoPhysically moving the drive tends to reset the U-shape. Some will be damaged.
- deleted 1y ago[deleted]
- cjaackie 1y agoThey mentioned the cluster being used enterprise drives, I can see the desire to save money but agree, that is going to be one expensive mistake down the road. I should also note personally for home cluster use, I learned quickly that used drives didn’t seem to make sense. Too much performance variability.
- g413n 1y agoin a datacenter context failure rates are just a remote-hands recurring cost so it's not too bad with front-loaders e.g. have someone show up to the datacenter with a grocery list of slot indices and a cart of fresh drives every few months.
- guywithahat 1y agoUsed drives make sense if maintaining your home server is a hobby. It's fun to diagnose and solve problem in home servers, and failing drives give me a reason to work on the server. (I'm only half-joking, it's kind of fun)
- jms55 1y agoIf I remember correctly, most drives either: 1. Fail in the first X amount of time 2. Fail towards the end of their rated lifespan So buying used drives doesn't seem like the worst idea to me. You've already filtered out the drivers that would fail early. Disclaimer: I have no idea what I'm talking about
- dboreham 1y agoOver in hardware-land we call this "the bathtub curve".
- g413n 1y agowe don't have perfect metrics here but this seems to match our experience; a lot of failures happened shortly after install before the bulk of the data download onto the heap, so actual data loss is lower than hardware failure rates
- dylan604 1y agoI've mentioned this story before, but we had massive drive failures when bringing up multiple disk arrays. We get them racked on a friday afternoon, and then I wrote a quick and dirty shell script to read/write data back and forth between them over the weekend that was to kick in after they finished striping the raid arrays. By quick and dirty I mean there was no logging, and just a bunch of commands saved as .sh. Came in on Monday to find massive failures in all of the arrays, but no insight into when they failed during the stripe or during stressing them. It was close to 50% failure rate. Turned out to be a bad batch from the factory. Multiple customers of our vendor were complaining. All the drives were replaced by the manufacturer. It just delayed the storage being available to production. After that, not one of them failed in the next 12 months before I left for another job.
- jeffrallen 1y ago> next 12 months before I left for another job Heh, that's a clever solution to the problem of managing storage through the full 10 year disk lifecycle.
- ClaireBookworm 1y agogreat write up, really appreciate the explanations / showing the process
- nharada 1y agoSo how do they get this data to the GPUs now...? Just run it over the public internet to the datacenter?
- bayindirh 1y agoThey can rent a dark fiber for themselves for that distance, and it'll be cheap. However, as they noted they use 100gbps capacity from their ISP.
- nee1r 1y agoWe want to get darkfiber from the datacenter to the office. I love 100Gbps
- dylan604 1y agoI'm now envisioning a poster with a strand of fiber wearing aviators with large font size Impact font reading Dark Fiber with literal laser beams coming out of the eyes.
- geor9e 1y agoDoes San Francisco really still have dark fiber? That 90s bubble sure did overshoot demand.
- madsushi 1y agoDWDM tech improvements have outpaced nearly every other form of technology growth, so the same single pair of fiber that used to carry 10 Mbps can now carry 20 Tbps, which is a 2,000,000x multiplier. The same somewhat-fixed supply of fiber can go a very long way today, so the price pressure for access is less than you might expect.
- dpe82 1y agoI think these days folks say "dark fiber" for any kind of connection you buy. It bothers me too.
- miniman1337 1y agoUsed Disks, No DR, not exactly a real shoot out.
- nee1r 1y agoTrue, though this is specifically for pretraining data (S3 wouldn't sell us used disk + no DR storage).
- p_ing 1y agoYou're in a seismically active part of the world. Will the venture last in a total loss scenario?
- nee1r 1y agoWe're currently 1/1 for the recent 4.3 magnitude earthquake (though if SF crumbles we might lose data)
- p_ing 1y ago4.3 is a baby quake. I'd hope that you'd be 1/1!
- antonkochubey 1y agoThey spent $300,000 on drives, with AWS they would have spent 4x that PER MONTH. They're already ahead of the cloud.
- p_ing 1y agoAWS/cloud doesn't factor into my question what so ever. Loss of equipment is one thing. Loss of all data is quite a different story.
- Sanzig 1y agoI do appreciate the scrappiness of your solution. Used drives for a storage cluster is like /r/homelab on steroids. And since it's pretraining data, I suppose data integrity isn't critical. Most venture-backed startups would have just paid the AWS or Cloudflare tax. I certainly hope your VCs appreciate how efficient you are being with their capital :)
- leejaeho 1y agohow long do you think it'll be before you fill all of it and have to build another cluster LOL
- nee1r 1y agoAlready filled up and looking to possibly copy and paste :)
- giancarlostoro 1y agoSo, others have asked, and I'm curious myself are you sourcing the videos yourselves or third parties?
- tomas789 1y agoMy guess would be they are running some dummy app like quote of the day or something and it records the screen at 1fps or so.
- not--felix 1y agoBut where do you get 90 million hours worth of video data?
- deleted 1y ago[deleted]
- _1tem 1y agoAnd not just any video data, they specifically mentioned screen recordings for agentic computer uses. A very specific kind of video. My guess is they have a partnership with someone like Rewind.ai
- Barbing 1y ago“For your privacy, your screen and audio recordings are stored locally and NEVER leave your Mac.” Tell me it’s only someone _like_ Rewind and not actually them! Quoting from the Privacy page they link in their header.
- hengheng 1y ago> we prob don't want to have it in Europe Best indicator as to what that is.
- conception 1y agoArrr matey
- bobbob1921 1y agoAssuming my calculation is accurate, 90,000,000 hours of video using around 30 PB comes to an average bit rate of about 760k. (Hard to guess though bc I doubt they’re using up all the space they provision day1) So my guess is either CCTV type of footage where there’s large gaps of motion / high GOP / big codec gains - or something like desktop recordings which are generally very low bit rate even though they can be high res. At that bitrate I can’t imagine it’s something like YouTube video. (Unrelated to the bitrate maybe it’s something like all older public domain videos). I would love to have an idea of what type of videos they are using (just out of curiosity)
- mschuster91 1y agoShows how crazy cheap on prem can be. tips hat
- nee1r 1y agotips hat back
- stackskipton 1y agoNot included is overhead of dealing with maintenance. S3/R2 generally don’t require OPS type dedicated to care and feeding. This type of setup will likely require someone to spend 5 hours a week dealing with it.
- mschuster91 1y agoI once had about three racks full of servers under my control, admittedly they weren't a ton of disks, but still the hardware maintenance effort was pretty much negligible over a few years (until it all went to the cloud). The majority of server wrangling work I spent dealing with OS updates and, most annoyingly, OpenStack. But that's something you can't escape even if you run your stuff in the cloud...
- stackskipton 1y agoWith S3/R2 whatever, you do get away from it. You dump a bunch of files on them and then retrieve them. OS Updates, Disk Failures, OpenStack, additional hardware? Pssh, that's S3 company problem, not yours. $LastJob we ran a ton of Azure Web App Containers, alot of OS work no longer existed so it's possible with Cloud to remove alot of OS toil.
- nee1r 1y agoTrue, this is a large reason why we chose to have the datacenter a couple blocks away from the office.
- hanikesn 1y agoWhy 5h a week? Just for hardware?
- g413n 1y agothe doodles are great
- nee1r 1y agoThanks! Lots of hard work went into them.
- zparky 1y ago$125/disk, 12k/mo depreciation cost which i assume means disk failures, so ~100 disks/mo or 1200/yr, which is half of their disks a year - seems like a lot.
- devanshp 1y agono, we wanted to be conservative by depreciating somewhat more aggressively than that. we have much closer to 5% yearly disk failure rates.
- AnotherGoodName 1y agoIt's an accounting term. You need to report the value of assets of your company each reporting cycle. This allows you to report company profit more accurately since the 2400 drives aren't likely not worth what the company originally paid. It's stated as a tax write-off but people get confused with that term (they think X written off == X less tax paid). It's better to correctly state it as a way to more accurately report profit (which may end up with less company tax paid but obviously not 1:1 since company tax is not 100%). So anyway you basically pretend you resold the drives today. Here they are assuming in 3 years time no one will pay anything for the drives. Somewhat reasonable to be honest since the setup's bespoke and you'll only get a fraction of the value of 3 year old drives if you resold them.
- zparky 1y agooh i see, thanks! i might be too used to reading backblaze reports :p
- ttfvjktesd 1y agoThe biggest part that is always missing in such comparisons is the employee salaries. In the calculation they give $354k/year of total cost per year. But now add the cost of staff in SF to operate that thing.
- g413n 1y agosomeone has to go and power-cycle the machines every couple months it's chill, that's the point of not using ceph
- paxys 1y agoSo the drives are never going to fail? PSUs are never going to burn out? You are never going to need to procure new parts? Negotiate with vendors?
- buckle8017 1y agoThey mention data loss is acceptable, so im guessing they're only fixing big outages. Ignoring failed hdds week likely mean very little maintenance.
- theideaofcoffee 1y agoThis concern troll that everyone trots out when anyone brings up running their own gear is just exhausting. The hyperscalers have melted people’s brains to a point where they can’t even fathom running shit for themselves. Yes, drives are going to fail. Yes, power supplies are going to burn out. Yes, god, you’re going to get new parts. Yes, you will have to actually talk to vendors. Big. Deal. This shit is -not- hard. For the amount of money you save by doing it like that, you should be clamoring to do it yourself. The concern trolling doesn’t make any sort of argument against it, it just makes you look lazy.
- immibis 1y agoVery good point. There was something on the HN front page like this about self-hosted email, too. I point out to people that AWS is between ten to one hundred times more expensive than a normal server. The response is "but what if I only need it to handle peak load three hours a day?" Then you still come out ahead with your own server. We have multiple colo cages. We handle enough traffic - terabytes per second - that we'll never move those to cloud. Yet management always wants more cloud. While simultaneously complaining about how we're not making enough money.
- OutOfHere 1y agoIs it correct that you have zero data redundancy? This may work for you if you're just hoarding videos from YouTube, but not for most people who require an assurance that their data is safe. Even for you, it may hurt proper benchmarking, reproducibility, and multi-iteration training if the parent source disappears.
- nee1r 1y agoDefinitely much less redundancy, this was definitely a tradeoff we made for pretraining data and cost.
- Sanzig 1y agoDid you do any kind of redundancy at least (eg: putting every 10 disks in RAID 5 or RAID Z1)? Or I suppose your training application doesn't mind if you shed a few terabytes of data every so often?
- g413n 1y agoatm we don't and we're a bit unsure whether it's a free lunch wrt adding complexity. there's a really nice property of having isolated hard drives where you can take any individual one and `sudo mount` it and you have a nice chunk of training data, and that's something anyone can feel comfortable touching without any onboarding to some software stack
- randomfinn 1y agoI wonder if snapraid would work for this. Especially if your data is mostly written once and then just read, it could be an easy way to add redundancy while keeping isolated individual drives.
- RagnarD 1y agoI love this story. This is true hacking and startup cost awareness.
- nee1r 1y agoThanks!! :)
- boulos 1y agoIt's quite cheap to just store data at rest, but I'm pretty confused by the training and networking set up here. It sounds like from other comments that you're not going to put the GPUs in the same location, so you'll be doing all training over X 100 Gbps lines between sites? Aren't you going to end up totally bottlenecked during pretraining here?
- g413n 1y agoyeah we just have the 100gig link, atm that's about all the gpu clusters can pull but we'll prob expand bandwidth and storage as we scale. I guess worth noting that we do have a bunch of 4090s in the colo and it's been super helpful for e.g. calculating embeddings and such for data splits.
- mwambua 1y agoHow did you arrive at the decision of not putting the GPU machines in the colo? Were the power costs going to be too high? Or do you just expect to need more physical access to the GPU machines vs the storage ones?
- g413n 1y agoWhen I was working at sfcompute prior to this we saw multiple datacenters literally catch on fire bc the industry was not experienced with the power density of h100s. Our training chips just aren't a standard package in the way JBODs are.
- huxley_marvit 1y agodamn this is cool as hell. estimate on the maintenance cost in person-hours/month?
- jonas21 1y agoNice writeup. All of the technical detail is great! I'm curious about the process of getting colo space. Did you use a broker? Did you negotiate, and if so, how large was the difference in price between what you initially were quoted and what you ended up paying?
- nee1r 1y agoWe reached out to almost every colocation space in SF/some in Fremont to get quotes. There wasn't a difference between the quote price and what we ended up paying, though we did negotiate terms + one-time costs.
- toomuchtodo 1y agoPlease consider posting the quotes, even if you have to redact colo names.
- archmaster 1y agoHad the pleasure of helping rack drives! Nothing more fun than an insane amount of data :P
- nee1r 1y agoThanks for helping!!!
- miltonlost 1y agoAnd how much did the training data cost?
- jimmytucson 1y agoJust wanted to say, thanks for doing this! Now the old rant... I started my career when on-prem was the norm and remember so much trouble. When you have long-lived hardware, eventually, no matter how hard you try, you just start to treat it as a pet and state naturally accumulates. Then, as the hardware starts to be not good enough, you need to upgrade. There's an internal team that presents the "commodity" interface, so you have to pick out your new hardware from their list and get the cost approved (it's a lot harder to just spend a little more and get a little more). Then your projects are delayed by them racking the new hardware and you properly "un-petting" your pets so they can respawn on the new devices, etc. Anyways, when cloud came along, I was like, yeah we're switching and never going back. Buuut, come to find out that's part of the master plan: it's a no-brainer good deal until you and everyone in your org/company/industry forgets HTF to rack their own hardware, and then it starts to go from no-brainer to brainer. And basically unless you start to pull back and rebuild that muscle, it will go from brainer to no-brainer bad deal. So thanks for building this muscle!
- theideaofcoffee 1y agoI'm not op, but thanks for this. Like I mentioned in another comment, the wholesale move to the cloud has caused so many skills to become atrophied. And it's good that someone is starting to exercise that skill again, like you said. The hyperscalers are mostly to blame for this, the marketing FUD being that you can't possibly do it yourself, there are too many things to keep track of, let us do it (while conveniently leaving out how eye-wateringly expensive they are in comparison).
- tempest_ 1y agoThe other thing the cloud does not let you do is make trade offs. Sometimes you can afford not to have triple redundant 1000GB network or a simple single machine with raid may have acceptable down time.
- g413n 1y agoyeah this it means that even after negotiating much better terms than baseline we run into the fact that cloud providers just have a higher cost basis for the more premium/general product.
- pronoiac 1y agoI wonder if they'll go with "toploaders" - like Backblaze Storage Pods - later. They have better density and faster setup, as they don't have to screw in every drive. They got used drives. I wonder if they did any testing? I've gotten used drives that were DOA, which showed up in tests - SMART tests, short and long, then writing pseudorandom data to verify capacity.
- g413n 1y agoyeah we're very interested in trying toploaders, we'll do a test rack next time we expand and switch to that if it goes well. w.r.t. testing the main thing we did was try to buy a bit from each supplier a month or two ahead of time, so by the time we were doing the full build that rack was a known variable. We did find one drive lot which was super sketchy and just didn't include it in the bulk orders later. diversity in suppliers helps a lot with tail risk
- joshvm 1y ago"don't have to screw in every drive" is relative, but at least tool-less drive carriers are a thing now. A lot of older toploaders from vendors like Dell are not tool-free. If you bought vendor drives and one fails, you RMA it and move on. However if you want to replace failed drives in the field, or want to go it alone from the start with refurbished drives... you'll be doing a lot of screwing. They're quite fragile and the plastic snaps easily. It's pretty tedious work.
- tempest_ 1y agoUsed Supermicro machines of this generation and very cheap (all things considered) https://www.theserverstore.com/supermicro-superstorage-ssg-6049p-e1cr60l-4u-60-bay-lff-rackmount-server https://www.theserverstore.com/supermicro-superstorage-ssg-6...
- synack 1y agoIPMI is great and all, but I still prefer serial ports and remote PDUs. Never met a BMC I could trust.
- jeffrallen 1y agoTry Lenovo. Their BMCs Don't Suck (tm).
- toast0 1y agoSerial over IPMI, plus ipmi power control is pretty good when it works. Supermicro X10 and newer was pretty nice. X9 and X8 not as nice; it's not helpful when the serial over ipmi drops during reboot and doesn't come back in a reasonable amount of time, and then the graphical mode needs ancient java webstart with os and platform specific jni, oof.
- fragmede 1y agoMy question isn't why do it yourself. A quick back of the envelope math shows AWS being much more expensive. My question is why San Francisco? It's one of the most expensive real estate markets in the US (#2 residential, #1 commercial), and electricity is expensive. $0.71/KwH peak residential rate! A jaunt down 280 to San Jose's gonna be cheaper, at the expense of. having to take that drive to get hands on. But I'm sure you can find someone who's capable of running a DC that lives in San Jose and needs a job so the SF team doesn't have to commute down to South Bay. Now obviously there's something to be said for having the rack in the office, I know of at least two (three, now) in San Francisco, it just seems like a weird decision if you're already worrying about money to the point of not using AWS.
- hnav 1y agoArticle says their recurring cost is $17.5k, they'll spend at least that amount in terms of human time tending to their cluster if they have to drive to it. It's also a question of magnitudes, going from $0.5m/mo to $0.05m/mo (hard costs plus the extra headaches of dealing with cluster) is an order of magnitude, even if you could cut another order of magnitude it wouldn't be as impactful.
- renewiltord 1y agoProblem when you self-roll this is that you inevitably make mistakes and the cycle time of going down and up ruins everything. Access trumps everything. You can get a DC guy but then he doesn't have much to do post setup and if you contract that you're paying mondo dollars anyway to get it right and it's a market for lemons (lots of bullshitters out there who don't know anything). Learned this lesson painfully.
- g413n 1y agoit's not just in sf it's across the street from our office this has been incredibly nice for our first hardware project, if we ever expand substantially then we'd def care more about the colo costs.
- tarasglek 1y agoi am still confused what their software stack is, they dont use ceph but bought netapp, so they use nfs?
- OliverGuy 1y agoThe NetApps are just disk shelves, can plug it into a SAS controller and use whatever software stack you please.
- tarasglek 1y agobut they have multiple head nodes, so its some distributed setup or just active/passive type thing?
- hnav 1y agoI'm guessing the client software (outside the dc) is responsible for enumerating all the nodes which all get their own IP.
- nee1r 1y agoWe have a custom barebones solution that uses a hashring to route the files!
- toast0 1y agoI think each rack is one head node and several disk shelves (10?). No dual headed shelves.
- trebligdivad 1y agoThe networking stuff seems....odd. 'Networking was a substantial cost and required experimentation. We did not use DHCP as most enterprise switches don’t support it and we wanted public IPs for the nodes for convenient and performant access from our servers. While this is an area where we would have saved time with a cloud solution, we had our networking up within days and kinks ironed out within ~3 weeks.' Where does the switch choice come into whether you DHCP? Wth would you want public IPs.
- giancarlostoro 1y ago> Wth would you want public IPs. So anyone can download 30 PB of data with ease of course.
- buzer 1y ago> Wth would you want public IPs. Possibly to avoid needing NAT (or VPN) gateway that can handle 100Gbps.
- bombcar 1y agoI don't know what they're doing, but Mikrotik can perhaps route that → https://mikrotik.com/product/ccr2216_1g_12xs_2xq#fndtn-testresults https://mikrotik.com/product/ccr2216_1g_12xs_2xq#fndtn-testr... and is about the cost of their used thing. And I think this would be a banger for IPv6 if they really "need" public IPs.
- dustywusty 1y agoExactly what I came in to say, CCR2216 can do this for < $2k, and does it well.
- xp84 1y agoNo DHCP doesn't mean public IPs nor impact the need for NAT, it just means the hosts have to be explicitly configured with IP addresses, default gateways if they need egress, and DNS. Those IPs you end up assigning manually could be private ones or routable ones. If private, authorized traffic could be bridged onto the network by anything, such as a random computer with 2 NICs, one of which is connected eventually to the Internet and one of which is on the local network. If public, a firewall can control access just as well as using NAT can.
- OliverGuy 1y agoAren't those netapp shelves pretty old at this point? See a lot of people recommending against them even for homelab type uses. You can get those 60 drive SuperMicro JBODs for pretty cheap now, and those aren't too old, would have been my choice. Plus, the TCO is already way under the cloud equiv. so might as well spend a little more to get something much newer and more reliable
- g413n 1y agoyeah it's on the wishlist to try
- bobbob1921 1y agoThanks to op for actually replying to the various comments here - really appreciate that (and for the initial of course!)
- drnick1 1y agoEveryone should give AWS the middle finger and start doing this. Beyond cost, it's a matter of sovereignty over one's computing and data.
- twoodfin 1y agoIf this is a real market, I’d expect AWS to introduce S3 Junkyard with a similar durability and cost structure. They probably still won’t budge on the egress fees.
- Barbing 1y ago>S3 Junkyard There it is, the answer to how to mitigate brand damage when risking distance between themselves and some of those 9s.
- g413n 1y agowe would be so down to buy s3 junkyard tbh we were going around begging various storage clouds to offer us this before giving up and building it ourselves
- alchemist1e9 1y agoWould have been much easier and probably cheaper to buy gear from 45drives.
- renewiltord 1y agoThe cost difference is huge. Modern compute is just so much bigger than one would think. Hurricane Electric is incredibly cheap too. And Digital Realty in the city are pretty good. The funny thing is that the Monkeybrains guys will make room for you at $75/amp but that isn't competitive when a 9654 based system pulls 2+ amps at peak. Still fun for someone wanting to stick a computer in a DC though. Networking is surprisingly hard but we also settled for the cheapo life QSFP instead of the new Cisco switches that do 800 Gbps that are coming. Great writeup. One that would be fun is about the mechanics of layout and cabling and that sort of thing. Learning all that manually was a pain in the ass. It's not just written down somewhere and I should have done it when I was doing it but now I no longer am doing it and so can't provide good photos.
- ThinkBeat 1y agoSo now you have all - your storage in one place - you own all backup, -- off site backup (hot or cold) - uptime worries - maintenance drives -- how many can fail. before it is a problem - maintenance machines -- how many can fail. before it is a problem - maintenance misc/datacenter - What to do the electricity is cut off suddenly -- do you have a backup provider? -- disel generators? -- giant batteries? -- Will the backup power also run cooling? -natural disaster -- earthquake -- flooding -- heatwave - physical security - employee training / (esp. if many quit) - backup for networking (and power for it) - employees on call 24/7 - protection against hacking +++++ I agree that a lot of cloud providers overcharge by a lot, but doing it all yourself gives you a lot of headaches. co-hosting would seem like a valuable partial mitigator.
- pclmulqdq 1y agoMost of these come from your colo provider (including a good backup power and networking story), and you can pay remote hands for a lot of the rest. Things like "protection from hacking" also don't come from AWS.
- yread 1y agoYou could get pretty close to the cost 1$/TB/month using Hetzner's sx135 with 8x22TB so 140TB in raidz1 for 240 eur. Maybe you get a better rate if you rent 200 of them. Someone else takes care of a lot of risks and you can sleep well at night
- nodja 1y agoI don't think Hetzner provides locations in SF. Those 100GBit connections don't do much if they need to connect outside the city the rest of the equipment is in, but maybe peering has gotten better and my views are outdated.
- fuzzylightbulb 1y agoYou're good. The speed of light through a glass fiber is still just as slow as it ever was.
- g413n 1y agoyeah it's totally plausible that we go with something like this in the future. We have similar offers where we could separate out either the financing, the build-out, or both and just do the software. (for Hetzner in particular it was a massive pain when we were trying to get CPU quotas with them for other data operations, and we prob don't want to have it in Europe, but it's been pretty easy to negotiate good quotes on similar deals locally now that we've shown we can do it ourselves)
- mx7zysuj4xew 1y agoYou cannot use hetzner for anything serious. They'd most likely claim abuse and delete your data wholesale without notice
- fapjacks 1y ago100% this. Hetzner has no problems completely blowing away whatever you've got running for arbitrary reasons. And their support is incresibly bad.
- 1y ago
- deleted 1y ago[deleted]
- intalentive 1y ago“Solve computer use” and previous work is audio conversation model. How do these go together? Is the idea to replace keyboard and mouse with spoken commands? a la Star Trek
- g413n 1y agojust general research work. Once the recipes are efficient enough the modality is a smaller detail. On the product side we're trying to orient more towards 'productive work assistant' rather than the default pull of audio models towards being an 'ai friend'.
- nerpderp82 1y agoMake me transparent aluminum!
- coleca 1y agoFor a workload of that size you would be able to negotiate private pricing with AWS or any cloud provider, not just CloudFlare. You can get a private pricing deal on S3 with as little as half a PB. Not saying that your overall expenses would be cheaper w/a CSP than DIY, but its not exactly an apples to apples comparison of taking full retail prices for the CSPs against eBayed equipment and free labor (minus the cost of the pizza).
- g413n 1y agoegress costs are the crux for AWS and they didn't budge when we tried to negotiate that we them, it's just entirely unusable for AI training otherwise. I think the cloudflare private quote is pretty representative of the cheaper end of managed object-bucket storage. obv as we took on this project the delta between our cluster and the next-best option got smaller, in part bc the ability to host it ourselves gives us negotiating leverage, but managed bucket products are fundamentally overspecced for simple pretraining dumps. glacier does a nice job fitting the needs of archival storage for a good cost, but there's nothing similar for ML needs atm.
- epistasis 1y agoWhat sort of deal are you taking about? Would it be 50% or more?
- master_crab 1y agoYou can get way higher than 50% discounts with AWS (or any cloud) depending upon the scale of the buy.
- oasisbob 1y agoNot for that minimum 0.5PB volume. Even at 10PB, the storage commit discounts won't be anywhere near 50%. Probably more like 10-20%, if that.
- landryraccoon 1y agoTheir electricity costs are $10K per month or about $120K per year. At an interest rate of 7% that's $1.7M of capital tied up in power bills. At that rate I wonder if it makes sense to do a massive solar panel and battery installation. They're already hosting all of their compute and storage on prem, so why not bring electricity generation on prem as well?
- moffkalast 1y agoLet's just say we're not seeing all of these sudden private nuclear reactor investments for no reason.
- datadrivenangel 1y agoAt 120K per year over the three year accounting life of the hardware, that's 360k... how do you get to 1.7M?
- landryraccoon 1y agoIt seems unlikely to me that they'll never have to retrain their model to account for new data. Is the assumption that their power usage drastically drops after 3 years? Unless they go out of business in 3 years that seems unlikely to me. Is this a one-off model where they train once and it never needs to be updated?
- Onavo 1y ago> We kept this obsessively simple instead of using MinIO or Ceph because we didn’t need any of the features they provided; it’s much, much simpler to debug a 200-line program than to debug Ceph, and we weren’t worried about redundancy or sharding. All our drives were formatted with XFS. What do you plan to do if you start getting corruption and bitrot? The complexity of S3 comes with a lot of hard guarantees for data integrity.
- g413n 1y agoour training stack doesn't make strong assumptions about data integrity, it's chill
- htrp 1y ago>We threw a hard drive stacking party in downtown SF and got our friends to come, offering food and custom-engraved hard drives to all who helped. The hard drive stacking started at 6am and continued for 36 hours (with a break to sleep), and by the end of that time we had 30 PB of functioning hardware racked and wired up. So how many actual man hours for 2400 drives?
- g413n 1y agoaround 250
- Havoc 1y agoCool write-up. I do feel sorry for the friends that go suckered into doing a bunch of grunt work for free though
- g413n 1y agoyeah that's why we started paying people near the second half- not super clearly stated in the blogpost, but the novelty definitely wore off with plenty of drives left to stack, so we switched strategies to get it done in time. I think everyone who showed up for a couple hours as part of the party had a good time tho, and the engraved hard drives we were giving out weren't cheap :p
- pighive 1y agoHDDs - are never one time costs. Do datacenters also offer ordering and replacing HDDs?
- epistasis 1y agoWith 30PB it's likely they will simply let capacity fall as drives fail. They apparently have zero need for redundancy in their use case, and the failure rate won't be high enough to take out a significant percentage of their capacity.
- Symbiote 1y agoThey offer replacing, yes, but normally expect you to order the new one. (Usually covered by a warranty, sent next business day.)
- supermatt 1y agoWhere does one get “90 million hours of video data”?
- hmcamp 1y agoI’m also curious about this. I don’t recall seeing that mentioned in the article
- supermatt 1y agoIts in the first sentence: "We built a storage cluster in downtown SF to store 90 million hours worth of video data."
- amarcheschi 1y agoThey were asking for the source of those data
- supermatt 1y ago/facepalm
- neilv 1y agoAs a fan of eBay for homelab gear, I appreciate the can-do scrappiness of doing it for a startup. To adapt the old enterprise information infrastructure saying for startups: "Nobody Ever Got Fired for Buying eBay"
- Scramblejams 1y agoFun piece, thanks to the author. But for vicarious thrills like this, more pictures are always appreciated!
- echelon 1y agoIf the authors chime in, I'd like to ask what "Standard Intelligence PBC" does. Is it a public benefit corp? What are y'all building?
- nee1r 1y agoWe did want more pictures!! Recently bought a Sony A7III to capture more fun moments like this. We're working on pretraining computer action models from the ground up—hence the pretraining data cluster. We're a public benefit corp because we think its important for AGI to built in the public's interest + are planning on automating a lot of the work done on computers!
- Scramblejams 1y ago"The best camera is the one you have with you." Looking forward to the next buildout post!
- kid64 1y agoMany colos disallow photography.
- ThrowawayTestr 1y agoDIY is always cheaper than paying someone else. Great write-up.
- akreal 1y agoHow is/was the data written to disks? Something like rsync/netcat?
- nee1r 1y agoWe use the same nginx rust server to do file writes, it's done via web requests
- lucb1e 1y agoThe linked Discord post is also interesting and fun to read. Most of the post is more serious but this is one of the small gems: > One thing we discovered very quickly was that [world cup] goals scored showed up in our monitoring graphs. This was very cool because not only is it neat to see real-world events show up in your systems, but this gave our team an excuse to watch soccer during meetings. We weren’t “watching soccer during meetings”, we were “proactively monitoring our systems’ performance.” https://discord.com/blog/how-discord-stores-trillions-of-messages https://discord.com/blog/how-discord-stores-trillions-of-mes... It is linked as evidence for Discord using "less than a petabyte" of storage for messages. My best guess is that they multiplied node size and count from this post, which comes out to 708 TB for the old cluster and 648 in the new setup (presumably it also has some space to grow)
- g413n 1y agoyeah we weren't sure about putting that number esp whether it includes all the image attachments, but in any case it's at least around the right reference class for the largest text data operations.
- lisbbb 1y agoGarbage in, garbage out, 30 petabyte edition
- speransky 1y agoWhy re-invent the wheel instead of using Lustre filesystem? It's easy to deploy on such a small filesystem; it is easy enough. POSIX interface, multiple clients, supports high-speed networking...
- yodon 1y agoAny startup that has enough money to casually buy a two-letter domain name has too much money, period. Kind of like counting the number of Aeron chairs at startups of old. Not a good sign.
- lmm 1y agoIt's a .inc domain name, are those worth anything?
- 63 1y agoLooks like unregistered two letter .inc domains are going for $2300/yr. Certainly <5% of the cost of a single developer.
- dangoodmanUT 1y agoit's .inc... those aren't expensive
- 0xbadcafebee 1y agoFor massive amounts of high-performance storage, the cloud is absolutely the most expensive option, by far. Even just 100+TB is ridiculously expensive on any cloud provider. If your company revolves around large amounts of data, it can make sense to keep it on-prem... ...but only if you compute the TCO. The bandwidth, peering, service contracts, available power, cooling, networking, rack capacity, half-decent smart hands, spare gear, etc, etc. The disks won't be the majority of your bill, and the logistics are difficult. It can still be cheaper than $CLOUD, but you have to deal with all the cost and complexity that comes with DIY, so do your homework first.
- jillesvangurp 1y agoAWS obviously does the same but better. That's why they are so rich. Their cost is a tiny percentage of their revenue. They buy cheap servers, and then run lots of vms on them. Each of which delivers 10s/100s of $ per month. That server pays for itself in revenue within weeks/months. And it will be in service until it stops working which could be over five years. Same with storage, networking, gpus, etc. They've spent years optimizing everything so their monthly costs are going to be much lower than what these guys managed on their first attempt. They'll be using less energy. They run their own internet backbones and infrastructure, they design their own hardware and source components directly from the best suppliers, they have exclusive deals with energy providers, etc. Every thing these guys did, AWS does way better. And yet they charge 40x more. AWS at cost price would probably be 60-80x less than what they charge; if not more. Cloudflare undercuts them a bit because they are smaller but they can do the same things. So do MS, Google, and everybody else. This market is ripe for disruption. There should not be a need to shovel hundreds of billions per year into AWS revenue for the industry. The same business operating at 20% margins would be a game changer. And most of this stuff is commodity stuff. Why is there not more competition in this space driving pricing down aggressively? What's keeping competitors off the market?
- silisili 1y agoInertia. It's the new 'nobody gets fired for buying..' A previous project I worked on had relatively little traffic, and AWS costs were rather insane for that. I proposed exploring OVH or DO and probably get costs down to 2 digits per month. Upper management would hear nothing of it - AWS was what they wanted, costs be damned. They were more protecting their own jobs than making a technical decision, I think.
- Maxion 1y agoUnless the savings would be more than 100k Eur / 300k++ USD (I.e. total cost of one employee, it's not really worth it. Even then, moving to new infrastructure carries high risk for business disruptions which can cause an even bigger dip in revenue. The cloud providers have definitely optimized their pricing for maximum profit extraction. Costs are high, and in many cases it's not high enough to actually warrant changing infrstructure to cheaper alternatives. Sticking with AWS / Azure / GCP carries other benefits, too. You're more likely to find engineers who are experienced with those cloud platforms over, say, OVH.
- azinman2 1y agoWhere does one acquire 90M hours of video without being YouTube?
- Hobadee 1y agopr0n
- NitpickLawyer 1y ago"I swear I'm seeding those just so I get my ratio up" :)
- bilekas 1y agoOr the counter FB argument "It's not illegal because I'm NOT seeding."
- Barbing 1y agoAnywhere as long as you can avoid “legal/practice/business slog”. Success is defined by $1.5b settlements. :) just kidding but also curious where besides torrents
- hengheng 1y agoMy guess is automated surveillance, which is also where this whole play has to be headed.
- fuzzfactor 1y agoSeems like that would be a good niche, not only for avoiding massive copyright considerations. Also, it's some of the most boring footage where there's overwhelming amounts that's about the least desirable thing for humans to sit and watch every minute of. Why send a human to do a machine's job?
- Hobadee 1y agoWhile I don't completely disagree with all the downsides of Ceph, it also sounds like they haven't heard of Croit. I set it up at my last company and it's amazing. It takes like 99% of the Ceph headaches away, plus you get Ceph experts to talk to.
- richwisdomwise 1y ago[flagged]
- GoatInGrey 1y agoBut is it Web3 compatible?
- urbandw311er 1y agoWell done! I love the honest write up and the “can do” attitude. Must have been a lot of fun too. Out of interest why do you think you made the mistake of buying 20x more drives than you needed instead of the denser storage that you mention? Was there a reason you opted for this?
- Tepix 1y agoHe did mention that it would have been a higher up-front cost.
- g413n 1y agoI think <2x more drives than needed, not 20x (24 vs 14TB), but the racks holding the drives could've been denser. Around the same cost in any case and our colo doesn't charge for space, so it's not a big deal and we were just going with what we were familiar with, but something to try.
- urbandw311er 1y agoOops sorry, my bad! Great to read all about it - good luck with the project.
- jmakov 1y agoWonder why everybody's first pick is CEPH which is known for being hard to optimize vs e.g. SeaweedFS
- jnsaff2 1y agoIf I'd have to guess then I would think that Ceph is the only one who is truly open source and does not feature gate important parts to paid enterprise users. I did go through this couple of years ago and we ended up with Ceph as well. Combine this with reusing existing hardware that was very suboptimal for Ceph in several ways, it was a pretty bad experience and in the end for our use case AWS was able to offer a good enough pricing that the performance and reliability of S3 was a better deal than managing it ourselves. If I would do it again then I would make sure that I have the hardware setup that is ideal (plenty of SSD's for metadata, every spinning disk directly addressed as a single OSD, sound network topology and fast enough NIC's) and probably use Rook instead of cephadm. The monitoring, configuration and documentation side of Ceph is however still quite sad, it was really hard to figure out why something is slow and how to tune things faster. That said, if the Enterprise options are performing better or you at least get good support for tuning and optimizing then the alternatives could be well worth consideration.
- winterrx 1y agoGreat piece, thanks for the write up.
- Zvez 1y agoThat's basically scaled up story of 'I store my files on my computer and it is 10x cheaper than using dropbox' While disks fail rate is already explored in another threads here, there is one related thing that catch my interest. Disk failure in such setup is not just cost of new disk + replacement cost (someone has to go there and change it!). It also inconvenience with dealing with failing requests. Ok, you are willing to lose 5% of your dataset. But are your '200-lines of code' robust enough to handle such cases. What if disk didn't fail, but start to be veeeeery slow. Does your training process can efficiently skip such bad objects. Do you have enough transparency to understand how much data you already lost? Is it still below 5%? And so on and so forth. I feel like this article was written right after they built this construction and before let say 6 months of usage. Because I'm pretty sure their costs will go much higher than they calculated here. Especially if they start including hidden costs, like the work needed to be done on training side. Yes, cost for self-hosting most probably still be less than aws (aws is not cheap). But it might start to be comparable with storage solutions of small ('neo') cloud providers if you buy gpu there.