6 ms·
GPT-4 Architecture
- abtinsetyani 3y agosource: https://www.semianalysis.com/p/gpt-4-architecture-infrastructure https://www.semianalysis.com/p/gpt-4-architecture-infrastruc...
- version_five 3y agoA long intro with no real content, just tech bro "here's the thing" stuff trying to bait you into subscribing. Doesn't actually explain the architecture if you don't subscribe. Don't bother reading.
- famouswaffles 3y agoThe info is on the twitter link.
- version_five 3y agoI don't have a twitter account, I can only see the top level tweet that says "GPT-4's details are leaked. It is over. Everything is here:"
- cmcaleer 3y agohttps://nitter.net/Yampeleg/status/1678545170508267522 https://nitter.net/Yampeleg/status/1678545170508267522
- version_five 3y agoThanks!
- archiv 3y agonot found - is there an archived version I can take a look at?
- redox99 3y agohttps://i.4cdn.org/g/1689038229454107.png https://i.4cdn.org/g/1689038229454107.png
- brianjking 3y agohttps://archive.is/Y72Gu https://archive.is/Y72Gu
- vd1f386f3 3y agohttps://threadreaderapp.com/thread/1678545170508267522.html https://threadreaderapp.com/thread/1678545170508267522.html
- vd1f386f3 3y agohttps://web.archive.org/web/20230711002505/https://threadreaderapp.com/thread/1678545170508267522.html https://web.archive.org/web/20230711002505/https://threadrea...
- janalsncm 3y agoCurrently says “Tweet not found” for me
- famouswaffles 3y agoDown on twitter now so https://archive.is/Y72Gu https://archive.is/Y72Gu The reason companies/researchers haven't generally touched MoE for LLMs despite how good it sounds on paper is because they've typically sucked and underperformed their dense counterparts. assuming this is all true, Did Open ai do anything differently here or is it just scale ? I know this very recent paper shows MoE benefit far more from Instruct tuning - https://arxiv.org/abs/2305.14705 https://arxiv.org/abs/2305.14705 FLAN-MOE-32B comfortably surpasses FLAN-PALM-62B with a third of the compute. It goes from 25.5% to 65.4% on MMLU. In comparison, 55.1 to 59.6% for Flan-Palm 62b. That just kind of shows the underperformance you expect from sparse models. But from Open ai's technical report, it doesn't seem like they needed that. The Vision component seems to be just scale. Well all of it seems to be just scale. Seems like there's plenty scale left too as far as performance gains go.
- deleted 3y ago[deleted]