3 ms·
Hi HN, Summary: The latest batch of language models can be much smaller yet achieve GPT-3 like performance by being able to query a database or search the web
by jayalammar 5y ago
Hi HN,
Summary: The latest batch of language models can be much smaller yet achieve GPT-3 like performance by being able to query a database or search the web for information. A key indication is that building larger and larger models is not the only way to improve performance.
Hope you find it useful. All feedback is welcome!
- changoplatanero 5y agoHow would you respond to the argument that the size of the database should be accounted for when computing the size of the total model? How does the latency of the database lookup compare to the extra latency from running the full size gpt3?
- mountainriver 5y agoI’ve been thinking on this too, the real driver is just that it’s hard to scale models to that size, I think with the MoE work happening like GLAM it will hopefully be easier to scale models in a distributed fashion.
- jayalammar 5y agoThe size of the database, the training set, the details of the architecture, as well as results on benchmark tasks should all be considered in the comparison. I'm also a fan of Behavioral Testing [1]. Parameter count is not very accurate measure of model performance. Mixture of Expert models like the Switch Transformer [2] can be 1 trillion parameters in size, but are not 5X the performance, for example. They clock the retrieval at 10 ms, unclear if that includes the BERT inference, however. My assumption is that it does not. [1] https://arxiv.org/abs/2005.04118 https://arxiv.org/abs/2005.04118 [2] https://arxiv.org/abs/2101.03961 https://arxiv.org/abs/2101.03961
- axpy906 5y agoThanks for you and all that you do Jay. Any plans to do one on MOE for LM?
- jayalammar 5y agoI'm intrigued by them and wanted to dig deeper into Switch Transformer. So hopefully yeah if they continue to show promise. Thank you!