6 ms·
Maintaining large-scale AI capacity at Meta
- joaquincabezas 2y agoEven if this scale is massive and second-to-none, it’s funny how some issues are the same for all of us. In particular “Bad hosts are very bad” aka “a chain is only as strong as its weakest link” can happen with as little as a few (4) machines, and then ruin your day.
- loeg 2y agoThis is really unique to the AI training clusters for reasons I'm not super clear on. Most other types of horizontally scaled workloads can sort of tolerate a slightly underperforming host, or hosts going bad every so often, with little P99/P99.9 impact. For some reason, AI training workloads really cannot.
- eugenhotaj 2y agoThis is because everyone is training with synchronous sgd. all gpus need to synchronize on each gradient step so tail latency will kill you.
- Mehdi2277 2y agoI’ve worked at companies with async training. Async training does help on fault tolerance and also can assist with training thoroughput by being less reliant on slowest machine. It does add meaningful training noise and when we did experiments against sync training we got much more stable results with sync training and some of our less stable models would even sometimes have loss explosions/divergence issues with async training but be fine with sync training. Although even for async training generally I see dataset just sharded and if worker goes down then shard of data may be loss/skipped not some kind of smarter dynamic file assignment factoring when workers go down. Even basic things like job fails continue from last checkpoint with same dataset state for large epoch is messy when major libraries like tensorflow lack a good dataset checkpointing mechanism.
- chaos_emergent 2y agoWhen training a large model you're doing a single forward pass + backprop over multiple Infiniband-connected nodes for a single model instance, so if one node goes down it takes a logical unit of nodes down with it. For reference, GPT-4 was rumored to be around 1.7T, and doing some back-of-the-hand math[1], that's like 500-700 H100 GPUs per model instance, which means you need a multiple of that for any training parallelism whatsoever. [1] back-of-the-hand-math: 1.7T * 4 bytes = 6.8 TB; 3-4x that for activation + gradients = 27.2 TB; 27.2TB / (80GB / H100) = 349 H100s; 1.5-2x conservative multiplier accounting for not fully using node resources + memory overhead in the machine = ~500-700 H100s. truly insane numbers.
- furyofantares 2y agoThat trillion+ parameter count is the sum of each of the "experts", right?
- oersted 2y agoThe ever-circulating rumour is 1.7T - 1.8T for the whole thing. But it is not very substantiated, mostly started by SemiAnalysis and geohot based on rather loose speculation (such as API latency and price), and not much solid evidence to confirm it after that. And of course, it must have changed substantially with GPT-4-Turbo and GPT-4o. It would make sense if the cost reduction was larger than the price reduction, they probably have a higher profit margin now, and the price reduction has been very significant since GPT-4 release.
- realreality 2y agoI’m old enough to remember when companies were eager to claim that their data centers (or some aspect) were finally “carbon neutral”. Now, with the enormous data center growth for AI purposes, companies don’t even bother pretending that any of this is sustainable. At best, they might delude themselves into believing that a glorified text autocomplete program will magically solve the world’s problems, including the unsustainability of the machines running the program.
- 123yawaworht456 2y agowe could exist without any of the modern conveniences. let's tear down the electric grid and return to monke.
- realreality 2y agoThat’s right (even though you’re probably being facetious).
- AYBABTME 2y agoWe're way past that. Global warming requires (or will, soon enough) heat pumps for survival in many regions of the world. Plenty of regions require large amount of electricity for life critical functions. Degrowth isn't the answer.
- elcomet 2y agoOr alternatively, people will need to move to colder regions.
- killingtime74 2y agoOr there are just less people (as we see in developed countries birthrates)
- FuckButtons 2y agoYes, they will, which will drive conflict, xenophobia and economic destabilization in the countries those people move to, which will exacerbate global political tensions and probably wind up with us all getting wiped out in a nuclear configuration sometime before the century is over, so we might as well have really nice autocomplete before we get there.
- lumost 2y agoSo I was pondering, NVidia quarterly datacenter revenue is around 18.4 billion. Meaning that the raw cost input to the AI industry is somewhere around 14-24 billion dollars per year post depreciation. This is against known revenues of ~3.8 Billion at OpenAI and ~800 Million at Anthropic. Based on reported revenue at cohere of 20 MM - I think it's a fair assumption that the only other material revenue in the industry is in the applications side either at megacaps or smaller startups targeting various back office tasks. One could make a bearish claim on NVidia, that their revenue/valuation is unsustainable unless the AI industry grows 100x over the next few years.
- zitterbewegung 2y agoRandom commenter on HN just tried to do a bearish analysis that Morningstar or any other analysis group you can obtain. News at 11. How do we know that AI is going to be the only thing that GPGPUs are the end game?
- bamboozled 2y agoand that it grows using current day approaches and technologies...
- m3kw9 2y agoRegardless NVDA gonna pop Monday on this news
- kortilla 2y agoWhat news?
- candiddevmike 2y agoIt is unsustainable, but no one wants to see the music stop. AI hype is floating a lot of tech stocks right now, and NVIDIA is at the center of it.
- mlinhares 2y agoIf they can keep the narrative, like Elon has kept the Tesla narrative, they'll have made so much money that it doesn't matter if it is sustainable or not.
- notarealllama 2y agoMaintenance trains, interesting takeaway from this. There's a trolley problem joke in there somewhere.
- lilharddad 2y ago[flagged]
- nomilk 2y agoMeta is going hard into AI (both hardware and software), which is great to see. Something that's not super obvious is what specific features of existing apps require AI, that is, how will Meta get return on investment? Two uses I can think of are i) text and image content moderation on fb and instagram (won't need as many human reviewers if bots are as/more effective), and ii) chatbots for businesses (businesses could provide their business documentation to a meta LLM which could handle customer inquiries via messenger and whatsapp). Anything else?
- xyzzy4747 2y agoThey could make customer facing support bots for every business with a Facebook page
- ldjkfkdsjnv 2y agoTheir whole advertising business model gets better with LLM understanding of text. They can target ads better.
- Mehdi2277 2y agoThis is fair guess on intuition but working in recommender space on both content/ad recommendation, content understanding signals have pretty consistently across two companies and many projects tended to underwhelm and key signals are generally engagement signals (including event sequences) and many embeddings (user embedding, creator embedding, ad embedding, etc). The main place I’ve seen content understanding help is coldstart specially for new items by new creators.
- xwolfi 2y agoAnd you wonder sometimes if the products being advertised, themselves, couldn't matter more than the targets reading them. We've learned to recognize crappy offering and AI can try to make me read more and more ads relevant to what I'm saying, if it's a crap product, I won't pay anyway :(
- candiddevmike 2y ago
- Havoc 2y agoSounds like this opsplanner software has a fair bit of autonomy
- Narhem 2y agoThey make people into slaves. I’d be extremely happy when Facebook gets shutdown.
- deleted 2y ago[deleted]
- tgma 2y agoWe have gone full circle back to building the HPC supercomputer after quarter century of clusters.