14 ms·
Ask HN: What were the papers on the list Ilya Sutskever gave John Carmack?
John Carmack's new interview on AI/AGI [1] carries a puzzle:
“So I asked Ilya Sutskever, OpenAI’s chief scientist, for a reading list. He gave me a list of like 40 research papers and said, ‘If you really learn all of these, you’ll know 90% of what matters today.’ And I did. I plowed through all those things and it all started sorting out in my head.”
What papers do you think were on this list?
[1] https://dallasinnovates.com/exclusive-qa-john-carmacks-different-path-to-artificial-general-intelligence/
- klaussilveira 4y agoFollowing: https://twitter.com/u3dcommunity/status/1621524851898089478?s=20 https://twitter.com/u3dcommunity/status/1621524851898089478?...
- sebkomianos 4y agoFollowing: https://news.ycombinator.com/item?id=34643510 https://news.ycombinator.com/item?id=34643510
- mritchie712 4y ago[flagged]
- KRAKRISMOTT 4y agoStart tweeting at him until he shares
- fnordpiglet 4y agoClearly do this by tweet storming him via LLM
- steveBK123 4y agoAs an AI LLM, I cannot decide which academic papers are "best" as the idea of "best" is subjective and there are many different factors that need to be considered.
- cwillu 4y agoI apologize for the oversight, you are correct. Let me know if there's anything else I can help you with.
- albertzeyer 4y ago(Partly copied from https://news.ycombinator.com/item?id=34640251 https://news.ycombinator.com/item?id=34640251.) On models: Obviously, almost everything is Transformer nowadays (Attention is all you need paper). However, I think to get into the field, to get a good overview, you should also look a bit beyond the Transformer. E.g. RNNs/LSTMs are still a must learn, even though Transformers might be better in many tasks. And then all those memory-augmented models, e.g. Neural Turing Machine and follow-ups, are important too. It also helps to know different architectures, such as just language models (GPT), attention-based encoder-decoder (e.g. original Transformer), but then also CTC, hybrid HMM-NN, transducers (RNN-T). Some self-promotion: I think my Phd thesis does a good job on giving an overview on this: https://www-i6.informatik.rwth-aachen.de/publications/download/1223/Zeyer--2022.pdf https://www-i6.informatik.rwth-aachen.de/publications/downlo... Diffusion models is also another recent different kind of model. Then, a separate topic is the training aspect. Most papers do supervised training, using cross entropy loss to the ground-truth target. However, there are many others: There is CLIP to combine text and image modalities. There is the whole field on unsupervised or self-supervised training methods. Language model training (next label prediction) is one example, but there are others. And then there is the big field on reinforcement learning, which is probably also quite relevant for AGI.
- hardware2win 4y agoI do wonder whether people behind Attention is all you need paper Will receive Turing Award It is being cited often
- mattcaldwell 4y agoCame here expecting a Haiku.
- maxbond 4y agoThe authors who wrote "Attention is all you need" - Turing candidates?
- 4y ago
- hexhowells 4y agoWhile not all papers, this list contains a lot of important papers, writings, and conversations currently in AI: https://docs.google.com/document/d/1bEQM1W-1fzSVWNbS4ne5PopB2b7j8zD4Jc3nm4rbK-U/edit https://docs.google.com/document/d/1bEQM1W-1fzSVWNbS4ne5PopB...
- dev_0 4y ago[dead]
- jimmySixDOF 4y ago>90% of what matters today Strikes me as the kind of thing where that last 10% will need 400 papers
- tikhonj 4y agoAlong with the kind of details and tacit knowledge that never makes it into papers...
- mindcrime 4y ago"The first 90% is easy. It's the second 90% that kills ya."
- kabdib 4y ago"All projects are divided into three phases, each consisting of 90% of the work." -- just about everything I've shipped :-)
- michpoch 4y agoFor the last 10% you'll need to write a paper yourself.
- swyx 4y agomaybe thats the part he intends to deviate. he just doesnt need to reinvent the settled science.
- cloudking 4y agoIlya's publications may be on the list https://scholar.google.com/citations?user=x04W_mMAAAAJ&hl=en https://scholar.google.com/citations?user=x04W_mMAAAAJ&hl=en
- username3 4y agoThey asked on Twitter and he didn’t reply. We need someone with a blue check mark to ask. https://twitter.com/ifree0/status/1620855608839897094 https://twitter.com/ifree0/status/1620855608839897094
- mirekrusin 4y agoAsk Elon to ask him.
- optimalsolver 4y agoCarmack says he's pursuing a different path to AGI, then goes straight to the guy at the center of the most saturated area of machine learning (deep learning)? I would've hoped he'd be exploring weirder alternatives off the beaten path. I mean, neural networks might not even be necessary for AGI, but no one at OpenAI is going to tell Carmack that.
- mindcrime 4y agoWouldn't it be fair to say that one has to know what the current path is and have some idea where it leads and what its issues are, before forging a new path? I mean, any idiot can go off-trail and start blundering around in the weeds, and ultimately wind up tripping, falling, hitting their head on a rock, and drowning to death in a ditch. But actually finding a new, better, more efficient path probably involves at least some understanding of the status quo.
- someweirdperson 4y agoTo walk a path no knowledge of the existing is needed. But to be able to claim it is new it is. Even more so to be able to claim that the new is better.
- fnordpiglet 4y agoBias and ignorance are two different things. No knowledge is ignorance. Bias is using knowledge to judge new knowledge. The goal isn’t to pursue things with raging ignorance but to pursue them with no bias and collecting knowledge without conclusion, then once you’re knowledgeable of what is there you can take off with raging ignorance in the direction no one has gone before. But you can’t do than holding bias any more than you can having ignorance of what directions have been gone before.
- agar 4y ago> probably involves at least some understanding of the status quo. Oh man, you had me going with such a vivid metaphor. I was really hoping for a payoff in the end, but you abandoned it. The easy close would be "probably involves at least some understanding of the existing terrain" but I was optimistic for something less prosaic.
- layer8 4y ago[flagged]
- caxco93 4y agoThis comment feels very ChatGPTy
- nathias 4y agoI got: Some of the highly influential papers in the field of AI that could have been on the list include "Generative Adversarial Networks" by Ian Goodfellow et al., "Attention is All You Need" by Vaswani et al., "AlexNet: ImageNet Classification with Deep Convolutional Neural Networks" by Alex Krizhevsky et al., "Playing Atari with Deep Reinforcement Learning" by Volodymyr Mnih et al., "Human-level control through deep reinforcement learning" by Volodymyr Mnih et al., "A Few Useful Things to Know About Machine Learning" by Pedro Domingos, among many others.
- siekmanj 4y ago"RL: A Deep Reinforcement Learning Framework" seems to have been hallucinated, does not exist.
- homarp 4y agohttps://arxiv.org/abs/1611.02779 https://arxiv.org/abs/1611.02779 is the closest - RL2: Fast Reinforcement Learning via Slow Reinforcement Learning
- polskibus 4y agoWhat about just asking Carmack on twitter?
- jranieri 4y agoI did, without success.
- belter 4y agoI asked him too. He said: - Who are you, and how did you get into my house?
- wincy 4y agoI wouldn’t advise this after seeing what Carmack did to that guy he got in a headlock. [0] “That was the tap part”, makes me laugh every time. [0] https://m.youtube.com/watch?v=X68Mm_kYRjc https://m.youtube.com/watch?v=X68Mm_kYRjc
- TigeriusKirk 4y agoIs anyone asking Ilya Sutskever?
- arbuge 4y agoOr, more directly, ask Sutskever...
- databroker 4y ago[dead]
- zomglings 4y ago[flagged]
- unixhero 4y agoI don't care. The cat was probably put down painlessly at a vet. I don't see the issue at all. Let's not do these cancelling attempts.
- ryanSrich 4y agoWhy not put it up for adoption? Having it intentionally killed is psychotic.
- icepat 4y agoHaving an animal killed for no good reason, other than it caused you problems, is a bit twisted. In situations like this, rehoming them is the ethical thing to do. This is literally saying 'I don´t care if he´s killing kittens'.
- unixhero 4y agoAsk any vet if this happens all the time.
- icepat 4y agoYeah, but that does not mean it's not wildly unethical.
- xnickb 4y agoStill completely irrelevant for the discussion
- hungryforcodes 4y agoI enjoyed my meatballs today, btw!I'm joking in the sense I didn't eat meatballs today, but we kill animals all the time. That's what humans do. Do you eat meat or foods cooked with animal fats? I think you get my point. A life is a life. I don't see why pets are somehow more important than other animals.
- winwhiz 4y agoI had read that somewhere else and this is as far as I got https://twitter.com/id_aa_carmack/status/1241219019681792010 https://twitter.com/id_aa_carmack/status/1241219019681792010
- Liberonostrud 4y ago[flagged]
- EvgeniyZh 4y agoAttention, scaling laws, diffusion, vision transformers, Bert/Roberta, CLIP, chinchilla, chatgpt-related papers, nerf, flamingo, RETRO/some retrieval sota
- deleted 4y ago[deleted]
- seydor 4y agowhat do you mean 'scaling laws'?
- EvgeniyZh 4y agoJ. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. and multiple follow-ups
- theusus 4y agolike papers are that comprehensible.
- adt 4y agohttps://lifearchitect.ai/papers/ https://lifearchitect.ai/papers/
- chrgy 4y agoFrom ChatGPT, although personally I think this list is bit old but should be at the 60% mark at the very least Deep Learning: AlexNet (2012) VGGNet (2014) ResNet (2015) GoogleNet (2015) Transformer (2017) Reinforcement Learning: Q-Learning (Watkins & Dayan, 1992) SARSA (R. S. Sutton & Barto, 1998) DQN (Mnih et al., 2013) A3C (Mnih et al., 2016) PPO (Schulman et al., 2017) Natural Language Processing: Word2Vec (Mikolov et al., 2013) GLUE (Wang et al., 2018) ELMo (Peters et al., 2018) GPT (Radford et al., 2018) BERT (Devlin et al., 2019)
- loveparade 4y agoYou are getting downvoted because this list if from ChatGPT, but as a researcher in the field, this list is actually really good, except for perhaps the SARSA and GLUE papers, which are less generally relevant. I would add WaveNet, the Seq2Seq paper, GANs, some optimizer papers (e.g. Adam), diffusion models, and some of the newer Transformer variants. I'm very confident that this is pretty much what any researcher, including Ilya, would recommend. It really isn't hard to find those resources, they are simply the most cited papers. Of course you can go deeper into any of the subfields if you desire.
- machiaweliczny 4y agoAs a hobbyst I would add Tsetlin Machines, DreamerV3, DiffusER and most RL from DeepMind.
- mgaunard 4y agoIn my experience, all deep learning is overhyped, and most needs that are not already addressable by linear regressions can be done so with simple supervised learning.
- ilaksh 4y agoMy guess is that multimodal transformers will probably eventually get us most of the way there for general purpose AI. But AGI is one of those very ambiguous terms. For many people it's either an exact digital replica of human behavior that is alive, or something like a God. I think it should also apply to general purpose AI that can do most human tasks in a strictly guided way, although not have other characteristics of humans or animals. For that I think it can be built on advanced multimodal transformer-based architectures. For the other stuff, it's worth giving a passing glance to the fairly extensive amount of research that has been labeled AGI over the last decade or so. It's not really mainstream except maybe the last couple of years because really forward looking people tend to be marginalized including in academia. https://agi-conf.org https://agi-conf.org Looking forward, my expectation is that things like memristors or other compute-in-memory will become very popular within say 2-5 years (obviously total speculation since there are no products yet that I know of) and they will be vastly more efficient and powerful especially for AI. And there will be algorithms for general purpose AI possibly inspired by transformers or AGI research but tailored to the new particular compute-in-memory systems.
- TimPC 4y agoWhy do you think multimodal transformers will get us anywhere near general purpose AI? Multimodal transformers are basically a technology for sequence-to-sequence intelligent mappings and it seems to me extremely unlikely that general intelligence is one or more specific sequence-to-sequence mappings. Many specific purpose problems are sequence-to-sequence but these tend to be specialized functionalities operating in one or more specific domains.
- RC_ITR 4y agoA lot of people don't really get that our brains are a bunch of specialized subcomponents that work in concert (Your pre-frontal cortex just cannot beat your heart, not matter how optimized it gets). This is unsurprising, as our brains are one of the most complex/hard to monitor things on earth. When an artificial tool that is really a point solution "tricks" us into thinking it has replicated a task that requires complex multi-component functioning within our brain, we assume the tool is acting like our brain is acting. The joke of course being that if you maliciously edited GPT's index for translating vectors to words, it would produce gibberish and we wouldn't care (despite being the exact same core model). We are only impressed by the complex sequence to sequence strings it makes because the tokens happen to be words (arguable the most important things in our lives). EDIT: a great historic metaphor for this is how we thought about 'computer vision' and CNN's. They do great at identifying things in images, but notice that we still use image-based captcha's (Even on OpenAI sites no less!)? That's because it turns out optical illusions and context-heavy images are things that CNN's really struggle at (since the problem space is bigger than 'how are these pixels arranged')
- Phil_Latio 4y agoNot in the list: https://arxiv.org/pdf/1805.09001.pdf https://arxiv.org/pdf/1805.09001.pdf
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- querez 4y agoA lot of other posts here are biased to recent papers, and papers that had "a big impact", but miss a lot of foundations. I think this reddit post on the most foundational ML papers gives a lot more balanced overview: https://www.reddit.com/r/MachineLearning/comments/zetvmd/d_if_you_had_to_pick_1020_significant_papers_that/ https://www.reddit.com/r/MachineLearning/comments/zetvmd/d_i...
- sillysaurusx 4y ago"The email including them got lost to Meta's two-year auto-delete policy by the time I went back to look for it last year. I have a binder with a lot of them printed out, but not all of them." RIP. If it's any consolation, it sounds like the list is at least three years old by now. Which is a long time considering that 2016 is generally regarded as the date of the deep learning revolution.
- moglito 4y ago"considering that 2016 is generally regarded as the date of the deep learning revolution" -- I thought it was 2012, when AlexNet took the imagenet crown?
- sillysaurusx 4y agoThat's probably fair. But you'd be hard-pressed to find a DL stack to try out your ideas with prior to 2016, since that's when Tensorflow launched. :) (Gosh, it's been less than a decade. Time sometimes doesn't fly, considering how much it's changed the world since then...)
- abrichr 4y agoTheano was first released in 2007.
- sillysaurusx 4y agoThat’s actually fascinating. Were there many experiments done in it back in the 00’s? I’m just trying to imagine the things you could do with it back then. 2007 had relatively fast gpus for the time, but certainly nothing compared to today. Yet it’d certainly be enough for MNIST training, which makes me wonder what else could be done.
- ladberg 4y agoFWIW in 2016 I was at an ML team at Apple that had been shipping production neural networks on-device for a while already. At the everyone used an assortment of random tools (Theano, Torch, Caffe). I worked on an internal tool that originally started as a Theano fork but was closer to a modern-day Tensorflow XLA (and has since been axed in favor of Tensorflow for most teams).
- sho_hn 4y ago> "You’ll find people who can wax rhapsodic about the singularity and how everything is going to change with AGI. But if I just look at it and say, if 10 years from now, we have ‘universal remote employees’ that are artificial general intelligences, run on clouds, and people can just dial up and say, ‘I want five Franks today and 10 Amys, and we’re going to deploy them on these jobs,’ and you could just spin up like you can cloud-access computing resources, if you could cloud-access essentially artificial human resources for things like that—that’s the most prosaic, mundane, most banal use of something like this." So, slavery?
- hosolmaz 4y agoRelated: https://qntm.org/mmacevedo https://qntm.org/mmacevedo
- sho_hn 4y agoI was quoting "Measure of a Man" :-) "Lena" is a bit of different case because it's not AGI. Probably ripe for the "forced prison labor" suggested by your sibling as the moral cop-out. Imagine being sentenced to being a cloud VM image!
- EamonnMR 4y agoIs there a good way to distinguish between the brain dumps in Lena and what you'd call an AGI?
- sho_hn 4y agoA brain dump has a history, and we ascribe meaning to the past. As mentioned the thread here has mentioned forced prison labor as a form of socially acceptable slavery, and society could convince itself that a given brain dump deserves its fate, even that it is a form of atonement. Artifical life on the other hand is presumably "pure at birth". Of course it's not that easy. You could discuss whether individual instances have unique sets of human rights, and value potential futures over pasts.
- 4y ago
- throwaway4837 4y agoWow, crazy coincidence that you all read this article yesterday too. I was thinking of emailing one of them for the list, then I fell asleep. Cold emails to scientists generally have a higher success-rate than average in my experience.
- dang 4y agoRecent and related: John Carmack’s ‘Different Path’ to Artificial General Intelligence - https://news.ycombinator.com/item?id=34637650 https://news.ycombinator.com/item?id=34637650 - Feb 2023 (402 comments)
- daviziko 4y agoI wonder what would Ilya Sutskever would recommend as an updated list nowadays. I don't have a twitter account, otherwise I'd ask him myself :)
- evc123 4y agohttps://arxiv.org/abs/2210.14891 https://arxiv.org/abs/2210.14891
- vikashrungta 4y agoI posted a list of papers on twitter, and will be posting a summary for each of them as well. here is the list https://twitter.com/vrungta/status/1623343807227105280 https://twitter.com/vrungta/status/1623343807227105280 Unlocking the Secrets of AI: A Journey through the Foundational Papers by @vrungta (2023) 1. "Attention is All You Need" (2017) - https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762 (Google Brain) 2. "Generative Adversarial Networks" (2014) - https://arxiv.org/abs/1406.2661 https://arxiv.org/abs/1406.2661 (University of Montreal) 3. "Dynamic Routing Between Capsules" (2017) - https://arxiv.org/abs/1710.09829 https://arxiv.org/abs/1710.09829 (Google Brain) 4. "Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks" (2016) - https://arxiv.org/abs/1511.06434 https://arxiv.org/abs/1511.06434 (University of Montreal) 5. "ImageNet Classification with Deep Convolutional Neural Networks" (2012) - https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf https://papers.nips.cc/paper/4824-imagenet-classification-wi... (University of Toronto) 6. "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" (2018) - https://arxiv.org/abs/1810.04805 https://arxiv.org/abs/1810.04805 (Google) 7. "RoBERTa: A Robustly Optimized BERT Pretraining Approach" (2019) - https://arxiv.org/abs/1907.11692 https://arxiv.org/abs/1907.11692 (Facebook AI) 8. "ELMo: Deep contextualized word representations" (2018) - https://arxiv.org/abs/1802.05365 https://arxiv.org/abs/1802.05365 (Allen Institute for Artificial Intelligence) 9. "Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context" (2019) - https://arxiv.org/abs/1901.02860 https://arxiv.org/abs/1901.02860 (Google AI Language) 10. "XLNet: Generalized Autoregressive Pretraining for Language Understanding" (2019) - https://arxiv.org/abs/1906.08237 https://arxiv.org/abs/1906.08237 (Google AI Language) 11. T5: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer" (2020) - https://arxiv.org/abs/1910.10683 https://arxiv.org/abs/1910.10683 (Google Research) 12. "Language Models are Few-Shot Learners" (2021) - https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165 (OpenAI)
- codeviking 4y agoThis inspired us to do a little exploration. We used the top cited papers of a few authors to produce a list that might be interesting, and to do some additional analysis. Take a look: https://github.com/allenai/author-explorer https://github.com/allenai/author-explorer