5 ms·
Attention models? Attention existed before those papers. What they did was show that it was enough to predict next word sequences in a certain context. I'm cer
by freddealmeida 3y ago
Attention models? Attention existed before those papers. What they did was show that it was enough to predict next word sequences in a certain context. I'm certain they didn't realize what they found. We used this frame work in 2018 and it gave us wildly unusual behavior (but really fun) and we tried to solve it (really looking for HF capability more than RL) but we didn't see what another group found: that scale in compute with simple algorithms were just better. To argue one group discovered and changed AI and ignore all the other groups is really annoying. I'm glad for these researchers. They deserve the accolade but they didn't invent modern AI. They advanced it. In an interesting way. But even now we want to return to a more deterministic approach. World models. Memory. Graph. Energy minimization. Generative is fun and it taught us something but I'm not sure we can just keep adding more and more chips to solve AGI/SGI through compute. Or maybe we can. But that paper is not written yet.
- calepayson 3y agoI’m studying neuroscience but very interested in how ai works. I’ve read up on the old school but phrases like memory graph and energy minimization are new to me. What modern papers/articles would you recommend for folks who want to learn more?
- voiceblue 3y agoFor phrases, Google's TF glossary [0] is a good resource, but it does not cover certain subsets of AI (and more specifically, is mostly focused on TensorFlow). [0] https://developers.google.com/machine-learning/glossary https://developers.google.com/machine-learning/glossary
- peppertree 3y agoIf you are in neuroscience I would recommend looking into neural radiance fields rendering as well. I find it fascinating since it's essentially an over-fitted neural network.
- svachalek 3y agoSomeone put this link up on another discussion the other day and I found it really fascinating: https://bbycroft.net/llm https://bbycroft.net/llm I believe energy minimization is literal, just look at the size of that thing and imagine the power bill.
- arcen 3y agoHe is most likely referring to some sort of free energy minimization.
- voiceblue 3y ago> but we didn't see what another group found: that scale in compute with simple algorithms were just better The bitter lesson [0] strikes again. [0] http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- tracerbulletx 3y agoI kind of wonder if the reason this seems to be true is that emergent systems are just able to go a lot farther into much more complex design spaces than any system a human mind is capable of constructing.
- int_19h 3y agoI think this is overly general. A more accurate statement is that, on tasks where we don't actually understand how something works in precise detail, it's more effective to just throw compute at it until a system with "innate" understanding emerges. But if you do actually know how something works (rather than vague models with no clear supporting evidence), it's still more effective to engineer the system specifically based on that knowledge.
- fnordpiglet 3y agoIn fairness, an article that’s about “the Google engineers who incrementally advanced AI” wouldn’t sell many advertisements.
- acchow 3y agoI guess the same thing happened with special relativity and Poincaré was forgotten.
- autokad 3y agoI argue they definitely 'changed AI' but definitely agree they didn't 'invent modern AI'. Personally, I think both compute and NN architecture is probably needed to get closer to AGI.
- godelski 3y agoThis is a classic... This isn't the first time this piece has been posted before either. Here's one from FT last year[0] or Bloomberg[1]. You can find more. Google certainly played a major role, but it is too far to say they invented it or that they are the ones that created modern AI. Like Einstein said, shoulder of giants. And realistically, those giants are just a bunch of people in trench coats. Millions of researchers being unrecognized. I don't want to undermine the work of these researchers, but that doesn't mean we should also undermine the work of the many others (and thank god for the mathematicians who get no recognition and lay all the foundation for us). And of course, a triggered Yann[2] (who is absolutely right). But it is odd since it is actually a highly discussed topic, the history of attention. It's been discussed on HN many times before. And of course there's Lilian Weng's very famous blog post[3] that covers this in detail. The word attention goes back well over a decade and even before Schmidhuber's usage. He has a reasonable claim but these things are always fuzzy and not exactly clear. At least the article is more correct specifying Transformer rather than Attention, but even this is vague at best. FFormer (FFT-Transformer) was a early iteration and there were many variants. Do we call a transformer a residual attention mechanism with a residual feed forward? Can it be a convolution? There is no definitive definition but generally people mean DPMHA w/ skip layer + a processing network w/ skip layer. But this can be reflective of many architectures since every network can be decomposed into subnetworks. This even includes a 3 layer FFN (1 hidden layer). Stories are nice, but I think it is bad to forget all the people who are contribution in less obvious ways. If a butterfly can cause a typhoon, then even a poor paper can contribute to a revolution. [0] https://www.ft.com/content/37bb01af-ee46-4483-982f-ef3921436a50 https://www.ft.com/content/37bb01af-ee46-4483-982f-ef3921436... [1] https://www.bloomberg.com/opinion/features/2023-07-13/ex-google-scientists-kickstarted-the-generative-ai-era-of-chatgpt-midjourney https://www.bloomberg.com/opinion/features/2023-07-13/ex-goo... [2] https://twitter.com/ylecun/status/1770471957617836138 https://twitter.com/ylecun/status/1770471957617836138 [3] https://lilianweng.github.io/posts/2018-06-24-attention/ https://lilianweng.github.io/posts/2018-06-24-attention/
- a_wild_dandan 3y agoThis is an uncharitable and oddly dismissive take (i.e. perfect for HN, I suppose). Today's incredible state-of-the-art does not exist without the transformer architecture. Transformers aren't merely some lucky passengers riding the coattails of compute scale. If they were, then the ChatGPT app which set the world ablaze would've instead been called ChatMLP, or ChatCNN. But it's not. And in 2024 we still have no competing NLP architecture. Because the transformer is a genuinely profound, remarkable idea with remarkable properties (e.g. training parallelism). It's easy to downplay GPTs as a mostly derivative idea with the benefit of hindsight. I'm sure we'll perform the same revisionist history with state-space models, or whatever architecture eventually supplants transformers. Do GPTs build on prior work? Do other approaches and ideas deserve recognition? Yeah, obviously. Like...welcome to science. But the transformer's architects earned their praise -- including via this article -- which isn't some slight against everyone else, as if accolades were a zero-sum game. These 8 people changed our world and genuinely deserve the love!
- hn_throwaway_99 3y agoQuestion for you, as someone relatively new to the world of AI (well, not exactly new - I took many courses in AI, including neural networks, but in the late 90s... the world is just a tad different now!) Is there any good summary of the history of AI/deep learning from, say, late 00s/2010 to the present? I think learning some of this history would really help be better understand how we ended up at the current state of the art.
- Swiffy0 3y agoI've been following a data science course called "The Data Science Course 2023: Complete Data Science Bootcamp" at Udemy. The course starts all the way back from basic statistics and goes through things like linear regression and supposedly will arrive at neural networks and machine learning at some point. So I don't know if something like this is exactly what you're looking for, but I think that, in general, if one wants to learn about (the history) AI, then it might be a good idea to start from statistics and learn about how we got from statistics to where we are now.
- 3y ago
- j7ake 3y agoThey didn’t invent it, they advanced it. I agree. However, some advances can have huge consequences to the field compared to others, even if at the technical level they appear comparable. One example that comes to mind is CRISPR.