10 ms·
Richard Sutton and Andrew Barto Win 2024 Turing Award
- rvz 2y agoAbsolutely well deserved.
- darosati 2y agoHear hear
- ofirpress 2y agoGood time to re-read The Bitter Lesson: https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
- khaledh 2y agoIndeed a bitter lesson. I once enjoyed encoding human knowledge into a computer because it gives me understanding of what's going on. Now everything is becoming a big black box that is hard to reason about. /sigh/ Also, Moore's law has become a self-fulfilling prophecy. Now more than ever, AI is putting a lot of demand on computational power, to the point which drives chip makers to create specialized hardware for it. It's becoming a flywheel.
- anonzzzies 2y agoI am still hoping AI progress will get to the point where the AI can eventually create AI's that are built up out of robust and provable logic which can be read and audited. Until that time, I wouldn't trust it for risky stuff. Unfortunately, it's not my choice and within a scarily short timespan, black boxes will make painfully wrong decisions about vital things that will ruin lives.
- tromp 2y agoAI assisted theorem provers will go a bit in that direction. You may not know exactly how they managed to construct a proof, but you can examine that proof in detail and verify its correctness.
- anonzzzies 2y agoYes, I have a small team of (me being 1/3) doing formal verification in my company and we do this and it doesn't actually matter if how the AI got there; we can mathematically say it's correct which is what matters. We do (and did) program synthesis and proofs but this is all very far from doing anything serious at scale.
- InkCanon 2y agoWhat kind of company needs formal verification? Real time systems?
- anonzzzies 2y agoReal time / embedded / etc for money handling, healthcare, aviation/transport... And 'needs' is a loaded term; the biggest $ contributors to formal verification progress are blockchain companies these days while a lot of critical systems are badly written, outsourced things that barely have tests. My worst fear, which is happening because it works-ish, is vague/fuzzy systems being the software because it's so like humans and we don't have anything else. It's a terrible idea, but of course we are in a hurry.
- tasty_freeze 2y agoCompanies designing digital circuits use it all the time. Say you have a module written in VHDL or Verilog and it is passing regressions and everyone is happy. But as the author, you know the code is kind of a mess and you want to refactor the logic. Yes, you can make your edits and then run a few thousand directed tests and random regressions and hope that any error you might have made will be detected. Or you can use formal verification and prove that the two versions of your source code are functionally identical. And the kicker is it often takes minutes to formally prove it, vs hundreds to thousands of CPU hours to run a regression suite. At some point the source code is mapped from a RTL language to gates, and later those gates get mapped to a mask set. The software to do that is complex and can have bugs. The fix is to extract the netlist from the masks and then formally verify that the extracted netlist matches the original RTL source code. If your code has assertions (and it should), formal verification can be used to find counter examples that disprove the assertion. But there are limitations. Often logic is too complex and the proof is bounded: it can show that from some initial state no counter example can be found in, say, 18 cycles, but there might be a bug that takes at least 20 cycles to expose. Or it might find counter examples and you find it arises only in illegal situations, so you have to manually add constraints to tell it which input sequences are legal (which often requires modeling the behavior of the module, and that itself can have bugs...). The formal verifiers that I'm familiar with are really a collection of heuristic algorithms and a driver which tries various approaches for a certain amount of time before switching to a different algorithm to see if that one can crack the nut. Often, when a certain part of the design can be proven equivalent, it aids in making further progress, so it is an iterative thing, not a simple "try each one in turn". The frustrating thing is you can run formal on a module and it will prove there are no violations with a bounded depth of, say, 32 cycles. A week later a new release of your formal tool comes out with bug fixes and enhancements. Great! And now that module might have a proof depth of 22 cycles, even though nothing changed in the design.
- amelius 2y agoWell, take compiler optimization for example. You can allow your AI to use correctness-preserving transformations only. This will give you correct output no matter how weird the AI behaves. The downside is that you will sometimes not get the optimizations that you want. But, this is sort of already the case, even with human made optimization algorithms.
- cxr 2y agoCanonical URL: <http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html>
- kleiba 2y agoThis depends a little bit on what the goal of AI research is. If it is (and it might well be) to build machines that excel at tasks previously thought to be exclusively reserved to, or needing to involve, the human mind, then these bitter lessons are indeed worthwhile. But if you do AI research with the idea that by teaching machines how to do X, we might also be able to gain insight in how people do X, then ever more complex statistical setups will be of limited information. Note that I'm not taking either point of view here. I just want to point out that perhaps a more nuanced approach might be called for here.
- visarga 2y ago> if you do AI research with the idea that by teaching machines how to do X, we might also be able to gain insight in how people do X, then ever more complex statistical setups will be of limited information At the very least we know consistent language and vision abilities don't require lived experience. That is huge in itself, it was unexpected.
- kleiba 2y agoIs that true though given e.g. the hallucinations you regularly get from LLMs?
- probably_wrong 2y ago> At the very least we know consistent language and vision abilities don't require lived experience. I don't think that's true. A good chunk of the progress done in the last years is driven by investing thousand of man-hours asking them "Our LLM failed at answering X. How would you answer this question?". So there's definitely some "lived experience by proxy" going on.
- crabbone 2y agoI remember the article, and remember how badly it missed the point... The goal of writing a chess program that could beat a world champion wasn't to beat the world champion... the goal was to gain understanding into how anyone can play chess well. The victory in that match would've been equivalent to eg. drugging Kasparov prior to the match, or putting a gun to his head and telling him to lose: even cheaper and more effective.
- krallistic 2y ago"The goal of Automated driving is not to drive automatically but to understand how anyone can drive well"... The goal of DeepBlue was to beat the human with a machine, nothing more. While the conquest of deeper understanding is used for a lot of research, most AI (read modern DL) research is not about understanding human intelligence, but automatic things we could not do before. (Understanding human intelligence is nowadays a different field)
- crabbone 2y agoSeems like you missed the point too: I'm not talking about DeepBlue, I'm talking about using the game of chess as a "lab rat" in order to understand something more general. DeepBlue was the opposite to the desire of understanding "something more general". It just found a creative way to cheat at chess. Like that Japanese pole jumper (I think he was Japanese, cannot find this atm) who instead of jumping learned how to climb a stationary pole, and, in this way, won a particular contest. > most AI (read modern DL) research is not about understanding human intelligence, but automatic things we could not do before. Yes, and that's a bad thing. I don't care if shopping site recommendations are 82% accurate rather than 78%, or w/e. We've traded an attempt at answering an immensely important question for a fidget spinner. > Understanding human intelligence is nowadays a different field And what would that be?
- DavidPiper 2y agoThis describes Go AIs as a brute force strategy with no heuristics, which is false as far as I know. Go AIs don't search the entire sample space, they search based on their training data of previous human games.
- dfan 2y agoThe paragraph on Go AI looked accurate to me. Go AI research spent decades trying to incorporate human-written rules about tactics and strategy. None of that is used any more, although human knowledge is leveraged a bit in the strongest programs when choosing useful features to feed into the neural nets. (Strong) Go AIs are not trained on human games anymore. Indeed they don't search the entire sample space when they perform MCTS, but I don't see Sutton claiming that they do.
- signa11 2y ago> ... This describes Go AIs as a brute force strategy with no heuristics ... no, not really, from the paper >> Also important was the use of learning by self play to learn a value function (as it was in many other games and even in chess, although learning did not play a big role in the 1997 program that first beat a world champion). Learning by self play, and learning in general, is like search in that it enables massive computation to be brought to bear. important notion here is, imho "learning by self play". required heuristics emerge out of that. they are not programmed in.
- HarHarVeryFunny 2y agoFirst there was AlphaGo, which had learnt from human games, then further improved from self-play, then there was AlphaGo Zero which taught itself from scratch just by self-play, not using any human data at all. Game programs like AlphaGo and AlphaZero (chess) are all brute force at core - using MCTS (Monte Carlo Tree Search) to project all potential branching game continuations many moves ahead. Where the intelligence/heuristics comes to play is in pruning away unpromising branches from this expanding tree to keep the search space under control; this is done by using a board evaluation function to assess the strength of a given considered board position and assess if it is worth continuing to evaluate that potential line of play. In DeepBlue (old IBM "chess computer" that beat Kasparov) the board evalation function was hand written using human chess expertise. In modern neural-net based engines such as AlphaGo and AlphaZero, the board evaluation function is learnt - either from human games and/or from self-play, learning what positions lead to winning outcomes. So, not just brute force, but that (MCTS) is still the core of the algorithm.
- perks_12 2y agoThe Bitter Lesson seems to be generally accepted knowledge in the field. Wouldn't that make DeepSeek R1 even more of a breakthrough?
- currymj 2y agothat was “bitter lesson” in action. for example there are clever ways of rewarding all the steps of a reasoning process to train a network to “think”. but deepseek found these don’t work as well as much simpler yes/no feedback on examples of reasoning.
- Buttons840 2y agoOof. Imagine the bitter lesson classical NLP practitioners learned. That paper is as true today as ever.
- jdright 2y ago> In computer vision, there has been a similar pattern. Early methods conceived of vision as searching for edges, or generalized cylinders, or in terms of SIFT features. But today all this is discarded.Modern deep-learning neural networks use only the notions of convolution and certain kinds of invariances, and perform much better. I was there, at that moment where pattern matching for vision started to die. That was not completely lost though, learning from that time is still useful on other places today.
- abdullahkhalids 2y agoI was an undergrad interning in a computer vision lab in the early 2010s. During group meeting, someone presented a new paper that was using abstract machine learning like stuff to do vision. The prof was so visibly perturbed and agnostic. He could not believe that this approach was even a little bit viable, when it so clearly was. Best lesson for me - vowed never to be the person opposed to new approaches that work.
- kenjackson 2y ago> Best lesson for me - vowed never to be the person opposed to new approaches that work. I think you'll be surprised at how hard that will be to do. The reason many people feel that way is because: (a) they've become an expert (often recognized) in the old approach. (b) They make significant money (or something else). At the end of the day, when a new approach greatly encroaches into your way of life -- you'll likely push back. Just think about the technology that you feel you derive the most benefit from today. And then think if tomorrow someone created something marginally better at its core task, but for which you no longer reap any of the rewards.
- abdullahkhalids 2y agoOf course it is difficult, for precisely the reasons you indicate. It's one of those lifetime skills that you have to continuously polish, and if you fall behind it is incredibly hard to recover. But such skills are necessary for being a resilient person.
- blufish 2y agonice read and insightful
- PartiallyTyped 2y agoThis made my day! Well deserved!
- darkoob12 2y agoThey should have given it to some physicists to make it even.
- porridgeraisin 2y agoTheir book "Introduction to Reinforcement Learning" is one of the most accessible texts in the AI/ML field, highly recommend reading it.
- barrenko 2y agoI've tried descending down the RL branch, always seem way out of my depth with those formulas and star-this, star-that.
- porridgeraisin 2y agoYeah, the formalisations can be hard to crunch through (especially because of [1]). But this book in particular is quite well laid out. I'd suggest getting a math background on the (very) basics of "contraction mappings", as this is something the book kind of assumes you have the knowledge of. [1] There's a lot of confusing naming. For example, due to its historic ties with behavioural psychology, there are a bunch of things called "eligibility traces" and so on. Also, even more than the usual "obscurity through notation" seen in all of math and AI, early RL literature in particular has particularly bad notation. You'd see the same letter mean completely different things (sometimes even opposite!) in two different papers.
- zelphirkalt 2y agoYou mean "Reinforcement Learning: An Introduction"? Or did they write another one?
- porridgeraisin 2y agoYeah that one. Messed up the name.
- incognito124 2y agoWhat is your background? Unfortunately I did not find it very accessible.
- jxjnskkzxxhx 2y ago
- ignoramous 2y agoCongratulations to Prof Barto & Prof Sutton. I'm sure the late Harry Klopf is all smiles (: > The ACM A.M. Turing Award, often referred to as the "Nobel Prize in Computing," carries a $1 million prize with financial support provided by Google, Inc. Good on Google, but there will be questions if their mere sponsorship in any way influences the awards. If ACM wanted, could it not raise $1m prize money from non-profits/trusts without much hassle?
- j7ake 2y agoAmazing that Sutton (American) chooses to live in Edmonton, AB rather than USA. Shows he has integrity and is not a careerist focused on prestige and money above all else.
- Philpax 2y agoKeen is a fully remote outfit, so he can work wherever. It's pretty likely that his reputation would open that door for him no matter where he goes.
- j7ake 2y agoAt his level it is much more than just being able to do what he wants, it’s about attracting resources and talent to accomplish his goals. From that perspective location still matters if you want to maximise impact
- tbrockman 2y agoAs someone who grew up in Edmonton, attended the U of A, and had the good fortune of receiving an incredible CS education at a discount price, I'm incredibly grateful for his (and the other amazing professors there) immense sacrifice. Great people and cheap cost of living, but man do I not miss the city turning into brown sludge every winter.
- jp57 2y agoHe's been there since he left Bell Labs, in the mid 2000's, I think. The U of A is, or was, rich with Alberta oil sands money and willing to use it to fund "curiosity-driven research", which is pretty nice if you're willing to live where the temperatures go down to -40 in the winter.
- armSixtyFour 2y agohttps://nationalpost.com/news/canada/ai-guru-rich-sutton-deepmind https://nationalpost.com/news/canada/ai-guru-rich-sutton-dee... He gave up his US citizenship years ago but he explains some of the reasons why he left. I'll also say that the AI research coming out of Canada is pretty great as well so I think it makes sense to do research there.
- pklee 2y agoVery well deserved !! Amazing contributions !!
- mark_l_watson 2y agoNice! Well deserved. They make both editions of their RL textbook available as a free to read PDF. I have been a paid AI practitioner since 1982, and I must admit that RL is one subject I personally struggle mastering, and the Sutton/Barto book, the Cousera series on RL taught by Professors White and White, etc. personally helped me: recommended! EDIT: the example programs for their book are available in Common Lisp and Python. http://incompleteideas.net/book/the-book-2nd.html http://incompleteideas.net/book/the-book-2nd.html
- zackkatz 2y agoVery cool to see this! It turns out my wife and I bought Andy Barto’s (and his wife’s) house. During the process, there was a bidding war. They said “make your prime offer” so, knowing he was a mathematician, we made an offer that was a prime number :-) So neat to see him be recognized for his work.
- HPMOR 2y agoThis is a crazy story!! Hahaha wow. What was the prime number?
- deleted 2y ago[deleted]
- dustfinger 2y agoHa haa, that is fantastic. You should have joked and said - "I'd like to keep things even between us, how about $2?"
- grumpopotamus 2y ago> we made an offer that was a prime number $12345678910987654321?
- optimalsolver 2y agoSo 2025 really is the year of agents.
- jimbohn 2y agoWell deserved, RL will only gain more importance as time goes on thanks to its (and neural nets) flexibility. The bitter lesson won't feel so bitter as we scale.
- byyoung3 2y agothey deserve it. definitely recommend their book
- nextworddev 2y agoRL may prove to be the most important tech going fwd due to test time compute
- vicentwu 2y agoGreat!
- carabiner 2y agoWonder if he's still working in AGI with Carmack.
- deleted 2y ago[deleted]
- cxie 2y agoHuge congratulations to Andrew Barto and Richard Sutton on the well-deserved Turing Award! as a student, their textbook Reinforcement Learning: An Introduction was my gateway into the field. I still remembered that how Chapter 6 on ‘Temporal Difference Learning’ fundamentally reshaped the way I thought about sequential decision-making. a timeless classic that I still highly recommend reading today!
- vonneumannstan 2y agoGood time to remind everyone that Sutton is a human successionist and doesn't care if humans all die. He is not to be trusted nor celebrated: https://www.youtube.com/watch?v=NgHFMolXs3U https://www.youtube.com/watch?v=NgHFMolXs3U
- nycticorax 2y agoThis is so silly. Do you imagine temporal difference learning is some kind of human successionist plot?
- vonneumannstan 2y agoThe video is not about his technical work but rather his view that AI will or should take over the future.
- nycticorax 2y agoBut the Turing Award is for his technical work.
- kalkin 2y agoSure, and his other views - in the scope of his professional expertise but also quite relevant to, uh, other humans - seem relevant in an HN thread about the Turing award. This place isn't exactly restricted to technical discussion of the details of RL algorithms, and it's pretty fair for humans to have views on whether we ought to be replaced. It's not just one Youtube video, it's a repeatedly expressed view: https://x.com/RichardSSutton/status/1575619655778983936 https://x.com/RichardSSutton/status/1575619655778983936 Valuing technological advance for its own sake "beyond good and bad" is an admirably clear statement of how a lot of researchers operate, but that's the best I can say for it.
- nycticorax 2y agoThe statement I take issue with is that Sutton "is not to be celebrated or trusted". Which I can only interpret to mean that the speaker does not think that Sutton should be celebrated or trusted. (And they've chosen to state it in a kind of pompous way.) Which I think is too strong on both counts. I (and apparently the ACM) think that Sutton should be celebrated for his technical accomplishments. Also, I think he probably can be trusted on a lot of technical matters. Should he be trusted on matters of whether there need to be safeguards on AI research imposed by the state? Maybe not, but those are only a subset of all the matters.
- rhema 2y agoI used their RL book for a course I taught. It's beautifully written and freely available (http://incompleteideas.net/book/the-book-2nd.html http://incompleteideas.net/book/the-book-2nd.html)! I kept getting distracted by the beautiful writing that I would miss the actual content.
- textlapse 2y agoThis is a long time coming. To see through an idea from start to finish and make this span an entire field instead of a sub chapter in a dynamic programming book. I wish a lot more games actually ended up using RL - the place where all of this started in the first place - would be really cool!
- jamesblonde 2y agoBuilt a lot of my PhD on their work 20 years ago. It really stood the test of time.
- wegfawefgawefg 2y agoThese guys are great but unfortunately the ai sutton and barto book is really bad. You would do better with Grokking Machine Learning by trask, and then a couple months of implementing ml papers.
- Buttons840 2y agoI second this suggestion. Read Grokking Deep Reinforcement Learning before reading Sutton. Well, the Sutton book is free, so take a peak, but if the formulas scare you then read Grokking Deep Reinforcement Learning.
- 317070 2y agoThese books are about different topics? Sutton and Barto is about Reinforcement learning, and the other book you mention by Trask is on Deep Learning?
- wegfawefgawefg 2y agoThe sutton and barto book is often given as an introductory ai book to people with no experience in ai or rl. This is unfortunate as it functions as neither a good rl book nor a good ai book. Wheras the introductory book Grokking Deep Learning walks you through implementing your own pytorch, and has a portion about rl near the end, then has a follow up book on rl, and it is trivial to have your own from scratch model and framework playing tic tac toe, snake, even without any math skills beyond multiplication. This happens without just smacking the reader with the modified bellman equation, and a bunch of chain rule backwards, and padded paragraphs intended to sell additional versions to universities.
- Cheng2023 2y ago[dead]