3 ms·
These models don't have a fixed context size and are progressively fine-tuned for longer and longer contexts. The context length also doesn't impact inference
by marmaduke 3y ago
These models don't have a fixed context size and are progressively fine-tuned for longer and longer contexts. The context length also doesn't impact inference cost.
Another aspect of performance is not just how well does the trained model perform, but is it data efficient (performance per token trained)? The comparison with Pythia (an open GPT) is shown in the article.
The rwkv4 paper is quite detailed and has examples of prompt and responses on the last few pages
https://arxiv.org/abs/2305.13048 https://arxiv.org/abs/2305.13048
And iirc rwkv5 is very similar to retnet which is detailed here
https://arxiv.org/abs/2307.08621 https://arxiv.org/abs/2307.08621
Edit now that I thought more about, the data efficiency seems like a highly important aspect given their noble goal to be fully multi lingual. This is fairly interesting theoretically as well and for other applications where abundance of data is not a given