Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
karalala
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
karalala
2y ago
Results of xlstm are promising but will need larger scale experiments. However they completely messed up benchmarking experiments for various RNN models which in their papers claim comparable and even better performance than base transforme
2.
▲
by
karalala
2y ago
True, but they normally arent this far off. HGRN claims that they outperform transformer for 1B parameter model trained on the pile. HGRN performing 8ppl worse suggests that its useless.
3.
▲
by
karalala
2y ago
Its xlstm contradicting existing peer reviewed papers lmao. Either xlstm should fix their benchmarks or existing peer reviewed papers should retract. RWKV-v6 > RWKV-v5 > RWKV-v4, not the other way round obviously. HGRN 8 ppl worse tha
4.
▲
by
karalala
2y ago
Already seeing major flaws in the paper. The benchmarking done in the table 1 is extremely questionable. Their table basically contradicts the results from multiple peer reviewed papers, especially for RNNs which report results much closer