7 ms·
X-Transformers: A fully-featured transformer with experimental features
- mrfusion 5y agoExplain like I’m a first year CS major?
- thesehands 5y agoTransformers suffer from a quadratic bottleneck when calculating attention. Much work has been done investigating where memory can be saved by being more explicit on which attentions to calculate. This repo implements transformers with noted improvements
- erik_seaberg 5y agoIt’s a kind of machine learning model: https://medium.com/inside-machine-learning/what-is-a-transformer-d07dd1fbec04 https://medium.com/inside-machine-learning/what-is-a-transfo...
- ericjang 5y agovanilla neural networks from the 1950's look like def f(x): for _ in range(3): x = g(Wx + b) return x Essentially it is a matrix multiplication, a vector addition, a non-linearity. Transformers are a modification to that architecture - using different multiplications, additions, and non-linearities. both of these are general in the sense that if you have enough of them, they can approximate any function. The ones used for transformers empirically do well on a lot of machine learning problems, particularly where data has a sequential nature.
- argvargc 5y agoUnfortunately for me, I genuinely thought this was going to be a DIY robot build that could disguise itself as something else.
- giords 5y ago*for us
- CamperBob2 5y agoI was hoping for some new ferromagnetic hotness, myself.
- themodelplumber 5y agoSame. Can you imagine some kind of epic, liberated Transformers trademark tech which makes use of a GitHub account, and is called "X-Transformers"? Well I just did for a few seconds... Actually maybe I'm not done yet
- shayankh 5y agoabsolutely fucking amazing
- fao_ 5y agoAs others have mentioned, anything obscure like this should literally come with a Wikipedia (or other such) link to explain what it is, what it does. This is the primary problem with small project READMEs, imo. They assume you're already familiar with them and know what the hell they are. Like, take Ironhide: https://github.com/MrMEEE/ironhide Optimus Support for Linux Through VirtualGL - PPA version also available That's... great. So it's doing something with GL, and it's running on Linux, but uhhh. my branch of the original bumblebee project.. What is Optimus? What is Bumblebee? The trick of it is that it links to a blog where neither of these terms are ever explained. Maybe it's to just look impressive on someone's CV? How could I even tell the difference? Likewise for this project, all you need in the README is one line that's like: X-Transformers is a re-implementation of Machine Learning Transformers that has been built based on experimental Arxiv papers It's a one-line fix but it'll stop people like me being confused as to whether or not you're implementing a new HTTP header
- nerdponx 5y agoI would agree if this were some kind of public release announcement. Would you say the same about an experimental programming language based on a bunch of recent PLT/CS research? That's pretty much what this is, but for machine learning. It's not meant to be "for the public", it's effectively research.
- enchiridion 5y agoWe're talking 2-3 sentences to explain what's going on. A researcher releasing work publicly on github is presumably doing so too spread the ideas.
- skybrian 5y agoSure, they have some audience in mind, but not necessarily us. There are a lot of documents that are public but are meant for a specialized audience. If they were meant for the general public, they'd be written very differently. Sharing a link on Hacker News is effectively taking it out of context and sometimes it's up to us to add that context back in. The author doesn't owe it to us.
- adontz 5y agoI have expected to see a 3D model for Optimus Prime.
- krick 5y agoThat's really cool. Now I need a bunch of pre-trained models for this...
- throwawaybbq1 5y agoFYI .. I work in deep learning and lucidrains is becoming a legend in my line of work. Seems someone who is obsessed about transformers (the deep learning ones, and rightly so, they are amazing). To the author (if you are reading this on HN), I want to thank you for the amazing work you have done! For the non-DL crowd, Transformers are a tsunami in deep learning for the past few years. They are topping benchmarks in many subfields. I do research professionally and this work is amazingly useful for people like me.
- mrfusion 5y agoSo they replace the perceptron?
- psychomugs 5y agoA perceptron is essentially a vanilla single-layer linear-activation neural network, like a lone Lego brick. Neural networks are Lego sets, and Transformers are the 1254-piece Millennium Falcon.
- gwern 5y agoActually... we're not sure! While it's true that Transformers work amazingly on more domains than they have any right to, it's unclear if they truly obsolete multi-layer perceptrons/fully-connected networks. You may have seen the recent splash of MLP-Mixer for image classification using perceptrons (https://arxiv.org/abs/2105.01601#google https://arxiv.org/abs/2105.01601#google) but FC-heavy or FC-only nets have popped up here and there, never quite going away, and slowly getting better initializations & residual layers making them ever more trainable. I have a bibliography of some links I've noted over the years about them: https://www.gwern.net/notes/FC https://www.gwern.net/notes/FC I've wondered if MLPs will be another example of the Bitter Lesson - the additional flexibility and power was their undoing until data/compute caught up and 'grad student descent' figured out how to use them right... If in the next 5 years everyone gets on the "MLP Is All You Need" train, I will be only mildly surprised.
- lucidrains 5y agoThanks for the kind words :) Hope you train something amazing with the code
- bratao 5y agolucidrains and Ice Cream are my references in terms of research, knowledge and productivity. Phil was always available to guide and hear me. One time I told him about an underground research in another language and he was kind enough to check if it had any merit. About X-Transformers, it is a very great piece of engineering that implemented almost of all possible improvements in transformers. But according to my experience and Phil himself, only the Feedforward GLU and RoPe (Rotary Positional Embeddings) works (or to be fair, they show improvements in more general use-cases)
- lucidrains 5y agoLol, thanks for mentioning Ice Cream
- bravura 5y agoWhat do you use for images that don’t have identical height and width? It seems the image transformer here expects square images.