2 ms·
I'd be curious to see this architecture trained to a comparable level to Llama3. 8B params, 15T tokens of training. Llama3 8B supercedes ChatGPT3.5. If Recurr
by TOMDM 2y ago
I'd be curious to see this architecture trained to a comparable level to Llama3.
8B params, 15T tokens of training.
Llama3 8B supercedes ChatGPT3.5. If RecurrentGemma scales, then it would be both faster and more capable than its Llama peer.