3 ms·
It is important to note the claims are made for similar dataset & params. If you like some independent verification outside eleuther circle. There is microsoft
by pico_creator 3y ago
It is important to note the claims are made for similar dataset & params.
If you like some independent verification outside eleuther circle.
There is microsoft paper on retnet: https://arxiv.org/abs/2307.08621 https://arxiv.org/abs/2307.08621
And the stanford heyena paper: https://arxiv.org/abs/2302.10866 https://arxiv.org/abs/2302.10866
Do note that both cited are competing alternatives for linear transformers, and in some sense competes with RWKV, in this upcoming "linear transfromer architecture" trend.
---
While all the above validates architecture. It's important to separate out architecture design from dataset.
The quality of an AI model answers, is equal parts as much about the dataset, than the architecture used. This is why it makes less sense comparing LLaMA2 directly with retnet/RWKV for example, due to the large differences in dataset size.
> Disclaimer: I am the same Eugene in the podcast