3 ms·
I wish someone performed a large scale experiment to evaluate all these alternate architectures. I kind of feel that they get drowned out by new sota results fr
by rdedev 3y ago
I wish someone performed a large scale experiment to evaluate all these alternate architectures. I kind of feel that they get drowned out by new sota results from openai and others. What I wish is something that tries to see if emergent behaviors pop up with enough data and parameters.
Maybe vision is special enough that convnets and approache transformer level performance or it could be generalized to any modality. I haven't read enough papers to know if someone has already done something like this but everywhere I look on the application side of things, vanilla transformers seems to be dominating
- whimsicalism 3y agoS4 & the H3 paper are probably what you are looking for