4 ms·
Agreed that as a whole, we do need more public benchmarks (both transformers and RWKV), with much longer sequence length. I view this as a case of the "benchma
by pico_creator 3y ago
Agreed that as a whole, we do need more public benchmarks (both transformers and RWKV), with much longer sequence length.
I view this as a case of the "benchmarks" and "datasets" not keeping up with the current progress, as the vast majority of benchmarks presumes the limit of up to 2k tokens. Which was cutting edge last year, but isn't today.
Internally we do have some benchmarking for how the decay over time affects lossy memory, and we are making major improvements in this direction for v5, and aim for no loss within the 8k to 100k and beyond context sizes. And hopefully with time, there will be more formalised standard benchmarks for larger context sizes.
Edit/PS: I believe until we get a sponsor for the compute required to build on similar scale and dataset compared to LLaMA2, so that we can have that "showdown" - we would be stuck - in its non threatening status. As there will always be that shadow of doubt, if it can scale to the next milestone that transformers have already crossed.