8 ms·
Previously posted about here: https://news.ycombinator.com/item?id=36671588 https://news.ycombinator.com/item?id=36671588 and here: https://news.ycombinator.com
by CSMastermind 3y ago
Previously posted about here: https://news.ycombinator.com/item?id=36671588 https://news.ycombinator.com/item?id=36671588 and here: https://news.ycombinator.com/item?id=36674905 https://news.ycombinator.com/item?id=36674905
With the original source being: https://www.semianalysis.com/p/gpt-4-architecture-infrastructure https://www.semianalysis.com/p/gpt-4-architecture-infrastruc...
The twitter guy seems to just be paraphrasing the actual blog post? That's presumably why the tweets are now deleted.
---
The fact that they're using MoE was news to me and very interesting. I'd love to know more details about how they got that to work. Variations in that implementation would explain the fluctuations in the quality of output that people have observed.
I'm still waiting for the release of their vision model which is mentioned here but we still know little about, sans a few demos a few months ago.
- jph00 3y agoThe previous posts are to a twitter thread that's been taken down, and the preview of a post that requires a $1000 subscription. This post however is freely available (for now at least).
- londons_explore 3y agoAnd the tweeter of the twitter thread paid the $1000, copied the useful info to twitter, and then did a credit card chargeback.
- renlo 3y agoSeems he summarized it and didn't copy it
- londons_explore 3y agoA summary isn't allowed under US copyright law. The copyright office calls them "condensations", and they are considered derivative works. His use was likely not within US copyright law. "Effect of the use upon the potential market for or value of the copyrighted work" is one of four factors a judge should use to decide if fair use applies, and it is clear that publishing the main information from an article, information which is not available elsewhere, freely, severely degrades the market for the original.
- doctor_eval 3y agoI had to ask GPT what MoE means: "MoE" in the context of artificial intelligence typically stands for "Mixture of Experts". This is a machine learning technique that is based on the idea of dividing a problem into sub-problems, solving each sub-problem with a specialized "expert" (or model), and then combining their outputs.
- ShamelessC 3y agoYep they (would) basically have 8-16 "experts" that are each about the size of GPT-3. Since they each see different batches of the dataset, they learn to model those distributions independently rather than the distribution of the whole dataset. Some of the attention is shared between them however. Then another "routing model" decides which model is most suitable for the given user prompt. Given they use relatively few experts, each one is likely similarly capable to the others on many tasks. I assume this make deployment easier and is a "more conservative" less risky approach. Even if the wrong model is chosen by the router, answers should still tend to be somewhat acceptable, for instance.
- ta988 3y agoThat's interesting because that's more or less on more level above the multi-head attention.
- refulgentis 3y agoSource?
- ShamelessC 3y agoJust theorizing from the top-level post here. No clue if it's legitimate.
- toxik 3y agoTo be clear, you just made up MoE details while MoE is actually well established and hails from decades old research?
- ShamelessC 3y agoFYI, George Hotz has been claiming to know this aspect for a couple of weeks now. > The fact that they're using MoE was news to me and very interesting. Maybe adds some legitimacy to the claim.
- krackers 3y agoInteresting on a meta point that the more clickbaity title "GPT-4 details leaked" won out over the more dispassionate but drier "GPT-4 Architecture, Infrastructure, Training Dataset, Costs".
- behnamoh 3y agoClickbait has its time and place. Despite my hatred towards it, sometimes it's really needed.
- H8crilA 3y agoIt is needed if you want people to click on your content more.
- swyx 3y agois it needed when you pay for the blogpost and then immediately chargeback the card like this dude did? https://twitter.com/untitled01ipynb/status/1678655012015071232 https://twitter.com/untitled01ipynb/status/16786550120150712... what a colossal asshole
- KaoruAoiShiho 3y agoWhy is it not okay to summarize? It's clearly transformative and not a copyright violation. Yes asshole but he should be in the clear legally.
- nicpottier 3y agoTo be fair the latter has the meat of it behind a paywall.
- Aachen 3y agoWhen choosing titles for my own submissions, yeah, the accurate title that HN says they desire gets no votes whatsoever. Any clickbait on here, people bring upon themselves (and this isn't even a clickbait-level title)
- hospitalJail 3y agoYeah the Mixture of Experts might have not been called out by name, but it was pretty obvious you were getting different models depending on the question. It goes to show how LLMs are nothing like AGI. I think combining it with a calculator is just a bandaid. A useful bandaid, but its not going to be able to do science ever.
- famouswaffles 3y agoSparse architectures are a way to theoritcally utilize only a small portion of a general models parameters at any given time. All "experts" are trained on the exact same data. They're not experts in the way you seem to think they are and they're certainly not wholly different models. The "experts" work at the token level. An expert for one token could be different from the expert chosen for the very next. GPT-4 isn't "nothing like AGI" any more than its dense equivalent would be.
- htss2013 3y agoI dont see how LLMs using many experts means it's very different from AGI. Why would anyone assume that human AGI isn't based on multiple models running in a similar architecture? At minimum humans are operating with a left and right brain, which process data very differently.
- pas 3y agoInterestingly Google was using ~2000 experts back in the first Trasnformer architecture (if I understand correctly) https://www.youtube.com/watch?v=9P_VAMyb-7k&t=6m42s https://www.youtube.com/watch?v=9P_VAMyb-7k&t=6m42s [sparsely-gated mixture of experts layer]