27 ms·
Building Meta's GenAI infrastructure
- alexsereno 3y agoHonestly Meta is consistently one of the better companies at releasing tech stack info or just open sourcing, these kinds of articles are super fun
- adamnemecek 3y agoDo you find this informative?
- alexsereno 3y agoYes of course - it depends on what lens though. If you mean "I'm learning to build better from this" then no, but its very informative on Meta's own goals and mindset as well as real numbers that allow comparison to investment in other areas, etc. Also the point was mostly that Meta does publish a lot in the open - including actual open source tech stacks etc. They're reasonably good actors in this specific domain.
- rshm 3y agoI think some elements of this stack might flow into the open compute.
- CuriouslyC 3y agoYann wants to be open and Mark seems happy to salt the earth.
- bananabrick 3y agoWhat do you mean?
- CuriouslyC 3y agoIn pretty much every interview, Yann has talked about how important that AI infrastructure is open and distributed for the good of humanity, and how he wouldn't work for a company that wasn't open. Since Mark doesn't have an AI product to cannibalize, it's in his interest to devalue the AI products of others ("salting the earth").
- Legend2440 3y agoI don't see how they're devaluing other people's AI products.
- CuriouslyC 3y agoThe Llama models have played a large part in fostering the development of the open source LLM ecosystem, and I expect Llama3 to put in performance > mistral medium and anthropic haiku while being fully open and able to be run on consumer hardware.
- crakenzak 3y agoThe angle is that by releasing cutting edge AI research to the public openly, the relative difference between open source models/tech and closed source tech shrinks. Whether or not you think the "value" of AI products is proportional to their performance gap vs the next closest thing or not is up to you. Very interesting PG essay I read recently talks about the opposite of this (Superlinear returns) where if you're half as good as the next competitor, you don't get half the customers, you get 0. Essay: https://paulgraham.com/superlinear.html https://paulgraham.com/superlinear.html
- nova22033 3y agoNew Linux versions don't "salt the earth" for Windows.
- nemothekid 3y ago
- torginus 3y agoI genuinely think one of the most plausible short-term dangers of AI is the creation of lifelike bots which will be absolutely indistinguishable from real humans in short-form online interaction. Since people don't want to talk to algorithms, this would result in them shunning all social media, which is a huge danger to companies in the space.
- marmaduke 3y agoJust for comparison, Swiss CSCS new Alps system will get 5k GH200 nodes (each with a H100).
- mjburgess 3y agoI'd be great if they could invest in an alternative to nvidia -- then, in one fell swoop, destroy the moats of everyone in the industry.
- math_dandy 3y agoA company moving away from Nvidia/CUDA while the field is developing so rapidly would result in that company falling behind. When (if) the rate of progress in the AI space slows, then perhaps the big players will have the breathing room to consider rethinking foundational components of their infrastructure. But even at that point, their massive investment in Nvidia will likely render this impractical. Nvidia decisively won the AI hardware lottery, and that's why it's worth trillions.
- mjburgess 3y agoI'm more concerned to avoid nvidia (et al.) market domination, than chasing the top-edge of the genAI benefits sigmoid. This will prevent much broad-based innovation.
- hx8 3y agoThis space is so compeitive, even if Nvidia is asleep at the wheel a competitor will come and push them before too long. AMD has a history of noticing when their competitors are going soft and rapidly being compeitive.
- whiplash451 3y agoPeople said the same thing when tensorflow was all the rage and pytorch was a side project. Granted, HW is much harder than SW, but I would not discount Meta's ability to displace NVIDIA entirely.
- Cthulhu_ 3y agoI don't think they could; nvidia has tons of talent, Meta would have to steal that. Meta doesn't do anything in either consumer or datacenter hardware that isn't for themselves either. Meta is a services company, their hardware is secondary and for their own usage.
- islewis 3y agoI know we won't get it this from FB, but I'd be really interested to see how the relationship of compute power to engineering hours scales. They mention custom building as much as they can. If FB magically has the option to 10x the compute power, would they need to re-engineer the whole stack? What about 100x? Is each of these re-writes just a re-write, or is it a whole order of magnitude more complex? My technical understanding of what's under the hood of these clusters is pretty surface level- super curious if anyone with relevant experience has thoughts?
- bilekas 3y agoI'm not 100% sure but I would.make an educated guess that that cluster in the first image for example is a sample of scalable clusters, so throwing more hardware at it could bring improvements but sooner or later the cost to improvements will call for an optimization or rewrite as you call it, so a bit of both usually. It seems a bit of a balancing act really!
- tintor 3y ago"just a re-write"
- mirekrusin 3y ago...the idea is that at some point it "just re-writes" itself.
- ametrau 3y agoThe day after that, we have true AGI.
- jvalencia 3y agoThe cost of training quickly outpaces the cost of development as context length increases. So hardware is cheap until it isn't anymore, by orders of magnitude.
- 3y ago
- lvl102 3y agoThis reads more like a flex for the investment community.
- choppaface 3y agoTotal cluster they say will reach 350k H100, which at $30k street price is about $10b. In contrast, Microsoft is spending over $10b per quarter capex on cloud. That makes Zuck look conservative after his big loss on metaverse. https://www.datacenterdynamics.com/en/news/q3-2023-cloud-results-ai-investments-drive-up-results-and-capex/ https://www.datacenterdynamics.com/en/news/q3-2023-cloud-res...
- baby 3y agoWhat loss lol. Stop the fud
- Legend2440 3y agoHas literally anyone spent money on the metaverse? Maybe it'll still take off in the future, but it's a $40b loss so far.
- artninja1988 3y ago>Has literally anyone spent money on the metaverse? I guess people buy their vr headsets, if that counts. I'm not too familiar with what the "metaverse" entails though...
- baby 3y agoHow is research and development a loss. If anything it positioned themselves as the top VR/AR company right when Apple entered the market. SMH
- yuliyp 3y agoThat's a weird comparison. The GPU is only a part of the capex: there's the rest of the servers and racks, the networking, as well as the buildings/cooling systems to support that.
- KaiserPro 3y agothe biggest cost at meta is infra. > In contrast, Microsoft is spending over $10b per quarter capex on cloud. to service other people's work load. Its a different business.
- DEDLINE 3y agoI wonder if Meta would ever try to compete with AWS / MSFT / GOOG for AI workloads
- lifeisstillgood 3y agoFB does not have the flywheel of running data centres - all three of those mentioned run hyper scale datacentres that they can then juice by “investing” billions in AI companies who then turn around and put those billions as revenue in the investors OpenAI takes money from MSFT and buys Azure services Anthropic takes Amazon money and buys AWS services (as do many robotics etc) I am fairly sure it’s not illegal but it’s definitely low quality revenue
- woah 3y agoSounds like it's free equity at the very least
- lotsofpulp 3y agoHow is it free equity? Spending money to invest it somewhere involves risks. You might recover some of it if the investment is valued by others, but there is no guarantee.
- miohtama 3y agoYou do not need cash in hands to invest. Instead, you print your own money (AWS credit) and use that to drive up the valuation, because this money costs you nothing today. It might cost tomorrow though, when the company starts to use your services. However depending the deal structure they might not use all the credit, go belly up before credit is used or bought up by someone with real cash.
- deleted 3y ago[deleted]
- deleted 3y ago
- hendersoon 3y ago350k H100 cards, around ten billion dollars just for the GPUs. Less if Nvidia gives a volume discount, which I imagine they do not.
- renegade-otter 3y agoIt will be ironic if Meta sinks all this money into the new trend and finds out later that it has been a huge boondoggle, just as publishers followed Facebook's "guidance" on video being the future, subsequently gutting the talent pool and investing into video production and staff - only to find out it was all a total waste.
- deleted 3y ago[deleted]
- motoxpro 3y agoIt already paid off. When the world moved from determinisic to probablistic ad modeling. That's why their numbers are so good right now compared to every other advertiser
- blitzar 3y agoIt already paid off. FB stonk price is up lots.
- althea_tx 3y agoCan you explain more about this?
- sangnoir 3y agoApple turned off a lot of "signals" used by advertisers for precisely targeted ads via persistent user-beacons. Facebook ad placement quality (and revenue) cratered in the immediate aftermath. Meta has since gotten better at it- likely with lots of AI-assistance and their revenue numbers reflect this. The targeting is now likely probabilistic in that the advertiser now makes educated guesses on the best ads to serve based on limited or non-existent identity information. So the AI efforts would have paid back by way of higher revenues.
- gingergoat 3y agoThe article doesn't mention MTIA, meta's custom ASIC for training & inference acceleration. https://ai.meta.com/blog/meta-training-inference-accelerator-AI-MTIA/ https://ai.meta.com/blog/meta-training-inference-accelerator... I wonder if they will use it in RSC.
- dazhbog 3y agoSearched H100 and an Amazon link popped up. Good reviews. https://www.amazon.com/Tesla-NVIDIA-Learning-Compute-Graphics/dp/B0C3XH4QSJ#customerReviews https://www.amazon.com/Tesla-NVIDIA-Learning-Compute-Graphic...
- mejutoco 3y agoThose reviews are hilarious
- zerop 3y ago> At Meta, we handle hundreds of trillions of AI model executions per day Such a large number, makes sense?
- pants2 3y agoPerhaps there's some combinatorics where every time an ad or post is displayed to the user, it runs through some hundreds/thousands of candidates and computes their relevance.
- GeneralMayhem 3y agoSure. 100T/day * 1day/86400sec ~= 1B/sec. They're probably considering at least a few hundred candidates per impression, and every impression is going to go through _at least_ two models (relevance and pCTR/revenue), so you could get there just with online serving at 5Mqps, which is plausible. But they're also going to be doing a lot of stuff in batch - spam predictions, ad budget forecasts, etc - so that every candidate actually runs through four or five different models, and every actual impression could do more than that.
- dakiol 3y agoWhat's an "AI model execution"? When I ask something to ChatGPT and it answers to me, does that count as 1 "AI model execution" for OpenAI?
- sangnoir 3y agoHow many ads does Meta serve a day, and how many AI model executions are done for each one? Repeat the same for stories, post and comment recommendations on Facebook and Instagram, and you have very big numbers. To that, Add VR, internal modeling and other backoffice/ offline analyses over billions of users and you'll easily get into the trillions.
- danielhanchen 3y agofloat8 got a mention! x2 more FLOPs! Also xformers has 2:4 sparsity support now so another x2? Is Llama3 gonna use like float8 + 2:4 sparsity for the MLP, so 4x H100 float16 FLOPs? Pytorch has fp8 experimental support, whilst attention is still complex to do in float8 due to precision issues, so maybe attention is in float16, and RoPE / layernorms in float16 / float32, whilst everything else is float8?
- GamerAlias 3y agoI was thinking why is this one guy on HN so deeply interested and discussing technical details from a minor remark. Then I clocked the name. Great work on Gemma bugs
- danielhanchen 3y agoOh thanks :) I always like small details :)
- andy99 3y agoIs there float8 support in any common CPU intrinsics? It sounds interesting but curious what will be the impact if any on CPU inference.
- deleted 3y ago[deleted]
- ashvardanian 3y agoNope. Moreover, simulating it even with AVX-512 is quite an experience. Been postponing it for 2 years now... But first of all, you need to choose the version of float8 you want to implement, as the standards differ between GPU vendors.
- janwas 3y agoWe use it in gemma.cpp [1]. This hybrid of E5M2 and E4M3 decodes to bf16 in ~14 instructions, so we can do that on the fly during dot products. [1]: github.com/google/gemma.cpp
- elwell 3y ago> Meta’s long-term vision is to build artificial general intelligence (AGI)
- valzam 3y agoDon't worry, this goal will change with the next hype cycle
- latchkey 3y agoI pity the fools that think AI is just another internet hype cycle.
- brookst 3y agoI’m old enough to remember the proud, defiant declarations that the internet was just a hype cycle.
- latchkey 3y agoI got my first email in 1991 and started my first internet business in 1995 (a web dev shop). My entire life has been an endless hype cycle.
- bennyelv 3y agoWell it was wasn’t it? There was a massive boom where loads of companies over promised what they would achieve, followed by a crash when everyone realised lots of them couldn’t, followed by stability for the smaller number that could. It was the very definition of a hype cycle as far as I can see. Hype cycle doesn’t mean “useless and will go away”, you have the second upward curve and then productivity. https://en.m.wikipedia.org/wiki/Gartner_hype_cycle https://en.m.wikipedia.org/wiki/Gartner_hype_cycle
- brookst 3y agoI don’t disagree, but a lot of “analysis” was not that nuanced. At one time I worked for a company where 90% of revenue was from printed periodicals. Smart, capable executives assured the whole company that the internet was not a threat, just something college kids used for fun. Colloquial, dismissive use of “hype cycle” does not usually mean “this will change the world but foolish things, soon forgotten, will also be done in the short term”. Though I agree a deeper understanding of the term can suggest that.
- latchkey 3y ago> we have successfully used both RoCE and InfiniBand clusters for large, GenAI workloads (including our ongoing training of Llama 3 on our RoCE cluster) without any network bottlenecks. Interesting dig on IB. RoCE is the right solution since it is open standards and more importantly, available without a 52+ week lead time.
- loeg 3y agoYeah, and RoCE isn't single vendor. I'm not sure IB scales to the relevant cluster sizes, either.
- anonymousDan 3y agoIs NVLink just not scalable enough here?
- loeg 3y agoI don't know. I haven't actually worked with IB in this specific space (or since before Nvidia acquired MLNX). My experience with RoCE/IB was for storage cluster backend in the late 2010s.
- fuddle 3y agoHow much are they paying for H100's? If they are paying $10k: 350,000 NVIDIA H100 x $10k = $3.5b
- ZiiS 3y agoThey may have to pay a premium to secure ~¼ of the output; certainly unlikely to be that steep a discount.
- theptip 3y agoSemi analysis posted recently noting that Meta locked in these purchases a while ago; something like a year or more. So they probably didn’t pay today’s spot rate.
- YetAnotherNick 3y ago> $3.5b Which is a fourth of what they spent in VR/AR in a year. And Gen AI is something they could easily get more revenue as it has now become proven technology, and Meta could possibly leapfrog others because of the data moat.
- NBJack 3y agoWhat moat exactly? Much of the user data they have access to is drying up due to new regulations, some of which prohibit IIRC direct use on models as well. I'm not even sure they can use historical data. Meta certainly has an edge in engineer count, undoubtedly. But I'd say they really, really want the metaverse to succeed more to have their on walled garden (i.e. equivalent power of Apple and Google stores, etc.). There's a reason they gave a hard pass to a Google partnership.
- Dr_Birdbrain 3y agoI think the raw text inside Facebook groups is at least as valuable as Reddit data. Even if demographics data is restricted under European law, the raw text of people interacting is quite valuable.
- froonly 3y agolmfao at the Meta folks not giving any credit whatsoever to the company that actually came up with and implemented the infrastructure work.
- zone411 3y agoMeta is still playing catch-up. Might be hard to believe but according to Reuters they've been trying to run AI workloads mostly on CPUs until 2022 and they had to pull the plug on the first iteration of their AI chip. https://www.reuters.com/technology/inside-metas-scramble-catch-up-ai-2023-04-25/ https://www.reuters.com/technology/inside-metas-scramble-cat...
- axpy906 3y agoDefinitely has some pr buzz and flex in the article. Now I see why.
- pwb25 3y agoso tired of this, not everyone need to work with AI stuff. work on facebook that is a disaster page instead
- delegate 3y agoSubtitled 'Here's what you'll never be able to do'.
- ilaksh 3y ago"Everything You Wanted to Know About GenAI at Meta, Except the One Thing You Honestly Care About" (Llama 3).
- wseqyrku 3y ago> Commitment to open AI innovation I see what you did there, Meta.
- owenpalmer 3y agoHaha, I noticed that too xD
- delanyoyoko 3y agoYou've got to read "open" roughly 3x in a paragraph.
- papichulo2023 3y agoIf they release models I dont care honestly, they can brag about that as much as they want.
- dekhn 3y agoit's really interesting just how similar these systems are to the designs adopted for HPC over the past few decades. I'm salty because it took a while for the ML community to converge on this (20+K GPUs connected by a real fabric with low latency and high bandwidth).
- mrkramer 3y ago"Share this: Hacker News" Noice
- BonoboIO 3y agoI thought at first "what are you talking about", when i check my uBlock filters. Was blocking the whole "Share this" content section. Sharing on Hacker News ... they now their audience.
- mrkramer 3y agoI also use uBlock but my filters are the default ones and I saw it without any problem but tbh this is the first time that I saw some post on the Web have HN as a share option or the first time that I was surprised seeing it. Maybe it has something to do with Google ranking "trusted human information and knowledge" higher than "non-human" information and knowledge[0] or simply some Meta software engineer loves and uses HN so s/he decided to include HN as well, idk. [0] https://news.ycombinator.com/item?id=39423949 https://news.ycombinator.com/item?id=39423949
- jvanderbot 3y agoSo, I'd love to work on optimizing pipelines like this. How does one "get into" it? It seems a ML scientist with some C/C++ and infra knowledge just dips down into the system when required? Or is it CUDA/SIMD experts who move "up" into ML?
- KaiserPro 3y agoA lot of the optimisation at this level is getting data into the right place at the right time, without killing the network. Its also a group effort to provide simple to use primitives that "normal" ML people can use, even if they've never used hyper scale clusters before. So you need a good scheduler, that understand dependencies (no, the k8s scheduler(s) are shit for this, plus it wont scale past 1k nodes without eating all of your network bandwidth), then you need a dataloader that can provide the dataset access, then you need the IPC that allows sharing/joining of GPUs together. all of that needs to be wrapped up into a python interface that fairly simple to use. Oh and it needs to be secure, pass an FCC audit (ie you need to prove that no user data is being used) have a high utilisation efficiency and uptime. the model stuff is the cherry on the top
- jvanderbot 3y agoOk, but back to my main question, how do I get into this?
- willsmith72 3y agoIt looks more like an infra problem than ML. "Software architect"s mixed with devops/infra/sre people
- jvanderbot 3y agoWell since I'm not a ML engineer of any kind - that's good!
- zooq_ai 3y ago
- seydor 3y agoThis is great news for Nvidia and their stock, but are they sure the LLMs and image models will scale indefinitely? nature and biology has a preference for sigmoids. What if we find out that AGI requries different kinds of cpu capabilities
- jiggawatts 3y agoIf anything, NVIDIA H100 GPUs are too general purpose! The optimal compute for AI training would be more specialised, but then would be efficient at only one NN architecture. Until we know what the best architecture is, the general purpose clusters remain a good strategy.
- pinko 3y agoThe link mentions "our internal job scheduler" and how they had to optimize it for this work -- does anyone know what this job scheduler is called, or how it works?
- KaiserPro 3y agoit might be twine: https://www.usenix.org/system/files/osdi20-tang.pdf https://www.usenix.org/system/files/osdi20-tang.pdf but I suspect its not that, because Twine is optimised for services rather than batch processing, and doesn't really have the concept of priorities.
- radicality 3y agoI would think it’s probably that. Also, has this been renamed to Twine from Tupperware?
- benreesman 3y agoI think it’s always useful to pay attention to the history on stuff like this and it’s a rare pleasure to be able to give some pointers in the literature along with some color to those interested from first-hand experience. I’d point the interested at the DLRM paper [1]: that was just after I left and I’m sad I missed it. FB got into disagg racks and SDN and stuff fairly early, and we already had half-U dual-socket SKUs with the SSD and (increasingly) even DRAM elsewhere in the rack in 2018, but we were doing huge NNs for recommenders and rankers even for then. I don’t know if this is considered proprietary so I’ll play it safe and just say that a click-prediction model on IG Stories in 2018 was on the order of a modest but real LLM today (at FP32!). The crazy part is they were HOGWILD trained on Intel AVX-2, which is just wild to think about. When I was screwing around with CUDA kernels we were time sharing NVIDIA dev boxes, typically 2-4 people doing CUDA were splitting up a single card as late as maybe 2016. I was managing what was called “IGML Infra” when I left and was on a first-name basis with the next-gen hardware people and any NVIDIA deal was still so closely guarded I didn’t hear more than rumors about GPUs for training let alone inference. 350k Hopper this year, Jesus. Say what you want about Meta but don’t say they can’t pour concrete and design SKUs on a dime: best damned infrastructure folks in the game pound-for-pound to this day. The talk by Thomas “tnb” Bredillet in particular I’d recommend: one of the finest hackers, mathematicians, and humans I’ve ever had the pleasure to know. [1] https://arxiv.org/pdf/1906.00091.pdf https://arxiv.org/pdf/1906.00091.pdf [2] https://arxiv.org/pdf/2108.09373.pdf https://arxiv.org/pdf/2108.09373.pdf [3] https://engineering.fb.com/2022/10/18/open-source/ocp-summit-2022-grand-teton/ https://engineering.fb.com/2022/10/18/open-source/ocp-summit... [4] https://youtu.be/lQlIwWVlPGo?si=rRbRUAXX7aM0UcVO https://youtu.be/lQlIwWVlPGo?si=rRbRUAXX7aM0UcVO
- junim 3y ago[flagged]
- dougdonohoe 3y agoHaving lived through the dot-com era, I find the AI-era slightly dispiriting because of the sheer capital cost of training models. At the start of the dot-com era, anyone could spin up an e-commerce site with relatively little infrastructure costs. Now, it seems, only the hyper-scale companies can build these AI models. Meta, Google, Microsoft, Open-AI, etc.
- danielmarkbruce 3y agoIt's not quite the same thing. A model is just one part of a product. You can spin up a product with zero infra and calling APIs hosting models.
- andy99 3y agoSo far it's been pretty "democratic" - I feel in no way disadvantaged because I can't train a foundation model myself. Actually the ecosystem is a lot better than 25 years ago - there are open source (or source available) versions of basically everything you'd want to participate in modern AI/ML.
- renegade-otter 3y agoNot everything has to be AI. You can run a small business infra for MUCH less than you did back then, especially if you adjust for inflation (!). Training AI models costs a fortune, but so far it's been just front-loading costs in hopes of a windfall. We'll see what actually happens.
- codingjaguar 3y ago"By the end of 2024, we’re aiming to continue to grow our infrastructure build-out that will include 350,000 NVIDIA H100 GPUs as part of a portfolio that will feature compute power equivalent to nearly 600,000 H100s." This AI game is getting into a GPU war. Heard that Meta is pushing a lot of CPU wordloads to GPU to co-locate with model inference for infra simplicity.
- sidcool 3y agoThose are some seriously great engineering numbers. Mera, with all the negative pressure it receives (rightfully so) is an engineering powerhouse. But I do wonder how they foresee monetising this.
- pedrovhb 3y agoMeta seems to actually be taking all the right steps in how they're contributing to open source AI research. Is this a "commodotize your complement" kind of situation?
- hansonpeter 3y ago[dead]
- spencerchubb 3y agoAll this compute and my Instagram Reels feed still isn't as good as my TikTok feed
- zeroonetwothree 3y agoWhat does that have to do with Gen AI
- spencerchubb 3y agoGenAI infra is the same as regular AI infra. They used GenAI in the title because it's a buzzword.
- ipsum2 3y agoNot really. Ranking and recommendation models require different infrastructure than LLMs. The models are generally smaller and require more data processing before training.
- refulgentis 3y agoYeah, no.
- lmm 3y agoIf Gen AI doesn't have anything to do with "Meta"'s actual business then WTF are they setting all this money on fire for?
- sashank_1509 3y agoMetas backing itself into a corner with its admirable commitment to open source. Unfortunately, at some point when they decide to monetize their billions spent and try to release a closed source model, the level of vitriol they will deal with will be an order of magnitude above what even OpenAI is experiencing. I don’t think they realize that!
- Horffupolde 3y agoThe general public doesn’t care. Only developers.
- bigcat12345678 3y agoNo Meta's commitment to Open Source is well under calculation. OCP is a way to rally lower-tier vendors to form a semi-alliance to keep up with super-gorilla like AWS & Google. LLaMA has already gained much more than its cost (look at the stock price, and the open source ecosystem built surrounding LLaMA, and Google's open source Gemma models which is a proof of Meta's success). IMHO, Meta's Open Source strategy already covered at least 5 years in prospect. That's enough to finesse a 180 degree turn around if necessary (i.e., from open source to close source)
- vernallen 3y ago[dead]
- yusyyuuu 3y ago[flagged]