3 ms·
1.5 years is actually not that bad. In fact, all changes and improvements to LLMs since the original Transformer paper is just the size -- tensor dimension, lay
by yvn1uo 3y ago
1.5 years is actually not that bad. In fact, all changes and improvements to LLMs since the original Transformer paper is just the size -- tensor dimension, layers, etc. GPT-3, which is still widely used today, was proposed more than 3 years ago.