3 ms·
I have no ML experience so tell me if I’m wrong here. The argument is that Transformers spend most of their compute finding this language subspace. Once the su
by brad0 4y ago
I have no ML experience so tell me if I’m wrong here.
The argument is that Transformers spend most of their compute finding this language subspace. Once the subspace is found, it’s very easy for it to add words/phrases/etc to the model.
What this is proposing is that we should try to find a better way to represent this subspace.
We have no idea how to represent it at the moment. But maybe transformers can help us with figuring that out.
- solarmist 4y agoNot a language subspace specifically, but a [learning maybe?] subspace for representing data and its relations to each other. A scaffolding that is highly transferrable between domains. His idea is that this structure might be similar to the brain's built-in structures, which make learning easy for children. Also, that language materials naturally surface the desirable properties, which is why large language models do so well on other types of problems.