6 ms·
it's quite interesting how at least in ML, the transformer architecture has 'won out', at least for the time being, it appears to be everywhere these days: http
by fault1 5y ago
it's quite interesting how at least in ML, the transformer architecture has 'won out', at least for the time being, it appears to be everywhere these days: https://threadreaderapp.com/thread/1468370605229547522.html https://threadreaderapp.com/thread/1468370605229547522.html
the advantage of transformers (computationally) seems to be how little sophistication the attention mechanism needs from AD systems (and how well it appears to scale with data). it's also a very static architecture in terms of a data flow/control flow perspective.
as far as I understand, this is far different from systems needing to be modeled in continuous time, especially things like SDEs. I am curious if things like delay embeddings will ever be modeled in terms of mechanisms similar to attention however.
- jowday 5y agoOutside of research transformers are rarely used for computer vision problems and CNNs remain the go to architecture. And you actually need to do some hacks to get transformers to work with computer vision at a meaningful scale (splitting images into patches and convoluting the patches to produce features to feed into the transformer).
- xiphias2 5y agoIt will be interesting when Tesla and Waymo moves to transformer architecture, but as you wrote my guess is that it's not yet in production for vision tasks.
- jowday 5y agoI’m not sure they will, at least not with the research in the state it is presently. Researchers are interested in vision transformers because they’re competitive with CNNs if you give them enough training data - they don’t drastically outperform them. Right now switching over to them would require a ton of code changes, relearning intuitions, debugging, profiling, etc. for not a ton of benefit.
- xiphias2 5y agoSure, I think the same, but the tweets came from Andrew Karpathy, he's watching this space like an eagle.
- liuliu 5y agoTesla did, as mentioned in their AI Day. It is not full transformer (aka ViT). The use transformer decoder to synthesize data from different cameras and decode 3d coordinates directly (aka DETR).
- xiphias2 5y agoThanks, sounds great, I'll read the DETR paper
- joconde 5y agoI’ve looked into transformers for semantic segmentation, but the patching aspect seems to make it hard too. Do you have some sources that describe these hacks in detail?
- lowdose 5y agoYou could do a code search on GitHub. I’m pretty lazy in the aspect of coding. I always seem to find a repo that has implemented an MVP with what I already had in mind. There are some gold nuggets on GitHub like Googles DDSP implementation they have academically published anonymous.
- fault1 5y ago> some hacks to get transformers to work with computer vision at a meaningful scale (splitting images into patches and convoluting the patches to produce features to feed into the transformer). sounds a lot like 'classical computer vision'. e.g, when I learned the subject (mid 2000s), topological features were all the rage: https://en.wikipedia.org/wiki/Digital_topology https://en.wikipedia.org/wiki/Digital_topology
- mirker 5y agoYeah. Even modern CV methods are hacky insofar as picking the “right” way to apply linear algebra. Convolution layers are hacked up matrix multiplications that are “inspired” by human vision. Of course, the real reason for the hacks is that form works in practice.
- liuliu 5y agoLike other comments, CNNs and LSTMs are still in wide use today. If you dig deep enough, position encoding doesn't really capture time-based series information that well.
- blovescoffee 5y agoCould you elaborate? I've built some Causal CNN's but never used a transformer for time series data. What are the challenges?