4 ms·
> some hacks to get transformers to work with computer vision at a meaningful scale (splitting images into patches and convoluting the patches to produce featur
by fault1 5y ago
> some hacks to get transformers to work with computer vision at a meaningful scale (splitting images into patches and convoluting the patches to produce features to feed into the transformer).
sounds a lot like 'classical computer vision'. e.g, when I learned the subject (mid 2000s), topological features were all the rage: https://en.wikipedia.org/wiki/Digital_topology https://en.wikipedia.org/wiki/Digital_topology
- mirker 5y agoYeah. Even modern CV methods are hacky insofar as picking the “right” way to apply linear algebra. Convolution layers are hacked up matrix multiplications that are “inspired” by human vision. Of course, the real reason for the hacks is that form works in practice.