5 ms·
GPT-4: 8 x 220B experts trained with different data/task distributions
- euclaise 3y agoThe only paper that I could find using an approach with fully separated experts like this is https://arxiv.org/pdf/2208.03306.pdf https://arxiv.org/pdf/2208.03306.pdf
- swyx 3y agothe source podcast that this came from: https://news.ycombinator.com/item?id=36407269 https://news.ycombinator.com/item?id=36407269
- SheinhardtWigCo 3y agoAs a heavy user of GPT-4 (I'm working on a plugin), reading this felt like a puzzle piece being dropped into place. Maybe this is just confirmation bias, but yeah, trying to push the model's capabilities is like working with a committee of brilliant minds chaired by an idiot. Also, I can see why they kept this secret. Competitors just shaved months off their R&D timelines.
- adeon 3y agoIs this an actually confirmed detail or just something George Hotz speculated? How credible is it?
- kylewatson 3y agoI feel like we should try asking it directly. Maybe it would tell us.