32 ms·
Competitive Programming with AlphaCode
- jonas_kgomo 5y agoGenuine question, what are the reasons to be a software engineer without much ML knowledge in 2022. Seems like a wake up call for developers
- 0xdeadbeefbabe 5y agoI hope you are right, but just to answer the question: all those other AI winters.
- jonas_kgomo 5y agoThats a good meditation. I think the winters were more driven by research dichotomy, for example Marvin Minsky's critique of the perceptron really slowed the research by 10 years. Advances made thus far have too much commercial relevance that companies invested dont look like they are gonna stop soon. But its a valid point. Looks like there is more upside being in subsets of computing like quantum computing, web3, metaverse etc than being a regular front-end engineer
- eulers_secret 5y ago> what are the reasons to be a software engineer without much ML knowledge in 2022. I'm not quite sure what you're asking, but my reason is that I do not enjoy working on/with ML. I'd personally rather quit the industry. But I work in embedded/driver development. I do not worry about ML models replacing me yet, but if I were just gluing together API calls I would be a bit worried and try to specialize.
- slingnow 5y agoGenuine question: what are the reasons to be a carpenter without much robotics / automation knowledge in 2022. Seems like a wakeup call for carpenters.
- jonas_kgomo 5y ago7 months ago, I asked natfriedman the same question, of which he responded: "We think that software development is entering its third wave of productivity change. The first was the creation of tools like compilers, debuggers, garbage collectors, and languages that made developers more productive. The second was open source where a global community of developers came together to build on each other's work. The third revolution will be the use of AI in coding. The problems we spend our days solving may change. But there will always be problems for humans to solve." https://news.ycombinator.com/item?id=27676266&p=2 https://news.ycombinator.com/item?id=27676266&p=2
- qualudeheart 5y agoFind something that’s hard and interesting. Someone will probably have a business trying to solve it and will hire you.
- hmate9 5y agoBetween this and OpenAI's Github Copilot "programming" will slowly start dying probably. What I mean by that is that sure, you have to learn how to program, but our time will be spent much more on just the design part and writing detailed documentation/specs and then we just have one of these AIs generate the code. It's the next step. Binary code < assembly < C < Python < AlphaCode Historically its always been about abstracting and writing less code to do more.
- wittycardio 5y agoSolving competitive programming problems is essentially solving hard combinatorial optimization problems. Throwing a massive amount of compute and gradient descent at the problem has always been possible. If I'm not mistaken what this does is reduce the representation of the problem to a state where it can run gradient descent and then tune parameters. The real magic is in finding structurally new approaches. If anything I'd say algorithms and math continue to be the core of programming. The particular syntax or level of abstraction don't matter so much.
- chroem- 5y ago> Solving competitive programming problems is essentially solving hard combinatorial optimization problems. True, but if you relax your hard requirements of optimality to admit "good enough" solutions, you can use heuristic approaches that are much more tractable. High quality heuristic solutions to NP-hard problems, enabled by ML, are going to be a big topic over the next decade, I think.
- wittycardio 5y agoI should correct myself, this isn't even that. This is just text analysis on codeforces solutions, which makes it even worse than I thought. Very pessimistic about it's generalizability.
- jdlshore 5y ago> If anything I'd say algorithms and math continue to be the core of programming. I disagree; I think the core of programming is analyzing things people want and expressing solutions to those wants clearly, unambiguously, and in a way that is easy to change in the future. I'd say algorithms and math are a very small part of this work.
- algon33 5y agoHow suprising did you guys find this? I'd have said there was a 20% chance of this performing at the median+level if I was asked to predict things beforehand.
- hackinthebochs 5y agoI didn't find it very surprising, but then I tend to be more optimistic than average about the capabilities of transformer models and the prospect of general AI in the relatively near term.
- marcusbuffett 5y agoI would have guessed around the same chance, this was surprising to me after playing around with copilot and not being impressed at all.
- baobabKoodaa 5y agoI would have said there is a ~0% chance of this happening within our lifetimes.
- machiaweliczny 5y agoI am surprised, as recently OpenAI had ~25% of easy problems and ~2% in competitive problems. Seems like DeepMind is ahead in this topic as well. Actually I think Meta AI had some interesting discovery recently that could possibly improve NNs in genral, so probably this as well. I am not in field but wonder if some other approaches like Tsetlin machines would be more useful for programming.
- algon33 5y agoSomehow I have never heard of Tsetlin machines before this. Are you talking about this https://ai.facebook.com/blog/the-first-high-performance-self-supervised-algorithm-that-works-for-speech-vision-and-text/ https://ai.facebook.com/blog/the-first-high-performance-self... result by MetaAI?
- machiaweliczny 5y ago
- BoardsOfCanada 5y agoDo I understand it correctly that it generated (in the end) ten solutions that then were examined by humans and one picked? Still absolutely amazing though.
- thomasahle 5y agoNo human examination was done. But it generated 10 solutions which it ran against the example inputs, and picked the one that passed. Actually I'm not sure if it ran the solutions against the example inputs or the real inputs.
- aliceryhl 5y agoThey used the real inputs. The example inputs were used to filter out which candidates to submit for the 10 tries.
- deleted 5y ago[deleted]
- aliceryhl 5y agoNo, they gave the algorithm 10 tries and tested all of them, and said that it was solved if any one of them worked.
- 37ef_ced3 5y agoThe example problem (essentially, is T a subsequence of S with deletions of size N) is a classic problem with no doubt dozens of implementations in AlphaCode's training set. And yet, what a garbage solution it produces. To illustrate the difference between intelligence and regurgitation, someone tell me what CoPilot generates for this: // A Go function to swap the sixth bit and seventeenth bit of a 32-bit signed integer. Here is a human solution: func swap(x int32) int32 { const mask = 1 << 5 var ( xor1 = (x>>11 ^ x) & mask xor2 = xor1 << 11 ) return x ^ xor1 ^ xor2 } CoPilot cannot reason numerically like this (understand "seventeenth bit" and "sixth bit" and generate the right code for that combination). It needs to understand the size of the gap between the bits, i.e., 11, and that's too hard.
- quenix 5y agoHere is what CoPilot wrote in Python: # A function to swap the sixth bit and seventeenth bit of a 32-bit signed integer. def swap_bits(x): # Get the value of the sixth bit. bit6 = x & (1 << 5) # Get the value of the seventeenth bit. bit17 = x & (1 << 16) # Swap the values of the sixth and seventeenth bit. bit6 = bit6 >> 5 bit17 = bit17 >> 16 # Combine the swapped values of the sixth and seventeenth bit. x = x ^ (bit6 << 16) x = x ^ (bit17 << 5) return x It actually synthesised about 10 viable solutions. Some interesting ones: def swap_six_seventeen(x): # Get the binary representation of the integer. binary = bin(x)[2:] # Add zeros to the beginning of the binary representation. binary = '0' * (32 - len(binary)) + binary # Swap the sixth and seventeenth bit. binary = binary[:5] + binary[17] + binary[5:17] + binary[18:] # Convert the binary back to an integer. return int(binary, 2)
- gfd 5y agoRelevant blogpost on codeforces.com (the competitive programming site used): https://codeforces.com/blog/entry/99566 https://codeforces.com/blog/entry/99566 Apparently the bot would have a rating of 1300. Although the elo rating between sites is not comparable, for some perspective, mark zuckerberg had a rating of ~1k when he was in college on topcoder: https://www.topcoder.com/members/mzuckerberg https://www.topcoder.com/members/mzuckerberg
- baobabKoodaa 5y agoThe median rating is not descriptive of median ability, because a large number of Codeforces competitors only do one or a few competitions. A very small number of competitors hone their skills over multiple competitions. If we were to restrict our sample to competitors with more than 20 competitions, the median rating would be much higher than 1300. It's amazing that Alphacode achieved a 1300 rating, but compared to humans who actually practice competitive coding, this is a low rating. To clarify, this is a HUGE leap in AI and computing in general. I don't mean to play it down.
- gfd 5y agoYou can find the rating distribution filtered for >5 contests here: https://codeforces.com/blog/entry/71260 https://codeforces.com/blog/entry/71260 I am rated at 2100+ so I do agree that 1300 rating is low. But at the same time it solved https://codeforces.com/contest/1553/problem/D https://codeforces.com/contest/1553/problem/D which is rated at 1500 which was actually non-trivial for me already. I had one wrong submit before getting that problem correct and I do estimate that 50% of the regular competitors (and probably the vast majority of the programmers commenting in this thread right now) should not be able to solve it within 2hrs.
- the-smug-one 5y agoI'm trying to solve this for fun, but I'm stuck! I've got a recursive definition that solves the problem by building a result string. I think it's a dynamic programming problem, but right now I can't see the shared sub-problems so :). Some real sour cherries being experienced from not getting this one!
- msoad 5y agoThis seems to have a narrower scope than GitHub Copilot. It generates more lines of code to a more holistic problem vs. GitHub Copilot that works as a "more advanced autocomplete" in code editors. Sure Copilot can synthesize full functions and classes but for me, it's the most useful when it suggests another test case's title or writes repetitive code like this.foo = foo; this.bar = bar etc... Having used Copilot I can assure you that this technology won't replace you as a programmer but it will make your job easier by doing things that programmers don't like to do as much like writing tests and comments.
- ipnon 5y agoThe big question seems to be whether par with professional programmers is a matter of increasing training set and flop size, or whether different model or multi-model architectures are required. It does look like we've entered an era where programmers who don't use AI assistants will be disadvantaged, and that this era has an expiration date.
- sharemywin 5y agoTo me it's not about it's current capabilities. It's the trajectory. This tech wasn't even a thing 2 years ago. There's billions being poured into it and every time someone uses these tools there's more free training data.
- jxcole 5y agoI feel like you are very defensive here and I want to be sure we take time to recognize this as a real accomplishment. Seriously though, I do doubt I can be fully replaced by a robot any time soon, it may be the case that soon enough I can make high-level written descriptions of programs and hand them off to an AI to do most of the work. This wouldn't completely replace me, but it could make developers 50x productive. The question is how elastic is the market...can the market grow in step with our increase in productivitiy? Also, please remember that as with anything, within 5 years we should see vast improvements to this AI. I think it will be an important thing to watch.
- visarga 5y ago
- thomasahle 5y agoNext they can train it on kaggle, and we'll start getting closer to the singularity
- NicoJuicy 5y agoI would stop programming if all we needed to write was unit tests :p
- FartyMcFarter 5y agoTo compensate, lots of people would start programming if that happened though. Many scientists would be interested in solving their field's problems so easily - certainly maths would benefit from it.
- rmujica 5y agowasn't it this the motivation for Prolog?
- mwattsun 5y agoSeems to me that this accelerates the trend towards a more declarative style of programming where you tell the computer what you want to do, not how to do it
- deleted 5y ago[deleted]
- pretendscholar 5y agoI am a little bitter that it is trained on stuff that I gave away for free and will be used by a billion dollar company to make more money. I contributed the majority of that code before it was even owned by Microsoft.
- visarga 5y agoPaying it forward, it will help others in turn.
- pretendscholar 5y agoYes it will help the already powerful players disproportionately.
- alphabetting 5y agoThey opensourced alphafold for anyone to use commercially despite big financial incentive to keep it private and use in their new drug discovery lab. No idea how this works or differs from alphafold but imagine they'll do the same here if possible
- pretendscholar 5y agoOnly after another lab made their own open source one that was comparable.
- kzrdude 5y agoThe problem is not really that microsoft owns github, or that licenses allow corporations free use, but that the tech giants are so big and have so much power.
- Permit 5y agoCan you elaborate and give some history? What code did you contribute, and how did it end up being used by Microsoft and then DeepMind?
- prideout 5y agoIt is obvious to me that computer programming is an interesting AI goal, but at the same time I wonder if I'm biased, because I'm a programmer. The authors of AlphaCode might be biased in this same way. I guess this makes sense though, from a practical point of view. Verifying correctness would be difficult in other intellectual disciplines like physics and higher mathematics.
- thomasahle 5y agoJust make it output a proof together with the program.
- qayxc 5y agoThat won't work because the systems aren't trained on proofs and proper theorem provers don't work that way either.
- blt 5y agoI am always surprised by the amount of skepticism towards deep learning on HN. When I joined the field around 10 years ago, image classification was considered a grand challenge problem (e.g. https://xkcd.com/1425/ https://xkcd.com/1425/). 5 years ago, only singularity enthusiast types were envisioning things like GPT-3 and Copilot in the short term. I think many people are uncomfortable with the idea that their own "intelligent" behavior is not that different from pattern recognition. I do not enjoy running deep learning experiments. Doing resource-hungry empirical work is not why I got into CS. But I still believe it is very powerful.
- qayxc 5y agoThis scepticism shouldn't surprise you. Not being sceptical is just an indicator that you've not been in the field for long enough. 30 years ago, the end of programming was prophesised, because 5th generation languages (5GL) and visual programming would enable everybody to design and build software. 20 years ago, low-code and application builders were said to revolutionise the industry and allow people in business roles to build their applications using just a few clicks. End-to-end model-driven design and development (e.g. using Rational Rose and friends) were to put an end to bugs and maintenance problems. 10 years ago it was new programming languages (e.g. Rust, Go, Swift, ...) and a shift to functional programming that was advertised as being "the future". Today it's back to "no code", e.g. tool-(AI-)driven development that's all the rage. It's not so much being "uncomfortable" or clinging to the exceptionalism of the human mind. It's just experience. Every decade saw its great big hype and technological breakthrough, but the lofty promises didn't hold water. Note that this doesn't mean nothing changed - model driven development still has its niche, visual programming is widely used in video production, rendering and game development. Features of functional programming have been added to many "legacy" languages and many of the newly introduced programming languages have become mainstream. The same will happen with AI generated software. There a large portion of the "mechanical" process of programming will be done by AI. Large and complex software systems with changing requirements, however, will still be designed and implemented primarily by people. Programming is a conversation between humans and machines. AI will in many cases shift the conversation closer to the human side, but fundamentally it'll still be the same thing. I like to think of it as the difference between writing your program in assembly and writing it in Haskell; different approaches, same basic activity.
- agentultra 5y agoThis is kind of neat. I wonder if it will one day be possible for it to find programs that maintain invariant properties we state in proofs. This would allow us to feel confident that even though it's generating huge programs that do weird things a human might not think of... well that it's still correct for the stated properties we care about, ie: that it's not doing anything underhanded.
- timetotea 5y agoIf you want some video explanation https://youtu.be/Qr_PCqxznB0 https://youtu.be/Qr_PCqxznB0
- deleted 5y ago[deleted]
- londons_explore 5y ago> AlphaCode placed at about the level of the median competitor, In many programming contests, a large number of people can't solve the problem at all, and drop out without submitting anything. Frequently that means the median scoring solution is a blank file. Therefore, without further information, this statement shouldn't be taken to be as impressive as it sounds.
- deleted 5y ago[deleted]
- mirrorlake 5y agoI've been wondering this for a while: In the future, code-writing AI could be tasked with generating the most reliable and/or optimized code to pass your unit tests. Human programmers will decide what we want the software to do, make sure that we find all the edge cases and define as many unit tests as possible, and let the AI write significant portions of the product. Not only that, but you could include benchmarks that pit AI against itself to improve runtime or memory performance. Programmers can spend more time thinking about what they want the final product to do, rather than getting mired in mundane details, and be guaranteed that portions of software will perform extremely well. Is this a naive fantasy on my part, or actually possible?
- EVa5I7bHFq9mnYK 5y agoIt seems to me that writing an exhausting set of unit cases is harder than writing the actual code.
- aduitsis 5y agoOtherwise the AI will just over-fit the unit test case subset.
- phreeza 5y agoAnd a second AI to generate additional test cases similar to yours (which you accept as also in scope) to avoid the first AI gaming the test.
- machiaweliczny 5y agoFirst you need really good infra to make it easy to test working multiple solutions for AI but I think this will be bleeding edge in 2030. EDIT: with in-memory DBs I can imagine AI assisted mainframe than can solve 90% of business problems.
- qayxc 5y ago> Is this a naive fantasy on my part, or actually possible? Possible, yes, desirable, no. The issue I have with all these end-to-end models is that they're a massive regression. Practitioners fought tooth and nails to get programmers to acknowledge correctness and security aspects. Mathematicians and computer scientists developed theorem solvers to tackle the correctness part. Practitioners proposed methodologies like BDD and "Clean Code" to help with stability and reliability (in terms of actually matching requirements now and in the future). AI systems throw all this out of the window by just throwing a black box onto the wall and scraping up whatever sticks. Unit tests will never be proof for correctness - they can only show the presence of errors, not their absence. You'd only shift the burden from implementation (i.e. the program) to the tests. What you actually want is a theorem prover that proofs the functional correctness in conjunction with integration tests that demonstrate the runtime behaviour if need be (i.e. profiling) and references that link implementation to requirements. The danger lies in the fact that we already have a hard time getting security issues and bugs under control with software that we should be able to understand (i.e. fellow humans wrote and designed it). Imagine trying to locate and fix a bug in software that was synthesised by some elaborate black box that emitted inscrutable code in absence of any documentation and without references to requirements.
- mrsuprawsm 5y agoDoes this mean that we can all stop grinding leetcode now?
- qualudeheart 5y agoCalling it now: If current language models can solve competitive programming at an average human level, we’re only a decade or less off from competitive programming being as solved as Go or Chess. Deepmind or openAI will do it. If not them, it will be a Chinese research group on par with them. I’ll be considering a new career. It will still be in computer science but it won’t be writing a lot of code. There’ll be several new career paths made possible by this technology as greater worker productivity makes possible greater specialization.
- keewee7 5y agoAI is being aggressively applied to areas where AI practitioners are domain experts. Think programming, data analysis etc. Programmers and data scientists might find ourselves among the first half of knowledge workers to be replaced and not among the last as we previously thought.
- Der_Einzige 5y agoI'm already anticipating having the job title of "Query Engineer" sometime in the next 30 years, and I do NLP including large scale language model training. :(
- qualudeheart 5y agoOne of the big venture capitalists predicted “prompt engineering” as a future high paid and high status position. Essentially handling large language models. Early prompt engineers will probably be drawn from “data science” communities and will be similarly high status, well but not as well paid, and require less mathematical knowledge. I’m personally expecting an “Alignment Engineer” role monitoring AI systems for unwanted behavior. This will be structurally similar to current cyber security roles but mostly recruited from Machine Learning communities, and embedded in a broader ML ecosystem.
- jonas_kgomo 5y agoI like this descriptions better, considering that companies like Anthropic are working specifically on Alignment and AI Safety. Being that the team actually spun out of Deep Mind, it is interesting.
- jdrc 5y ago"And so in 2022 the species programmus programmicus went extinct"
- udev 5y agoI am thinking whether this result can create a type of loop that can self-optimize. We have AI to generate reasonable code from text problem description. Now what if the problem description text is to generate such a system in the first place? Would it be possible to close the loop, so to speak, so that over many iterations: - text description is improved - output code is improved Would it be possible to create something that converges to something better?
- machiaweliczny 5y agoI am actually trying this. Basically by asking questions to AI and teaching it to generate code / google when it doesn't know something. The other process checks if code is valid and either ask it to get more context or executes code and feeds back to file :)
- machiaweliczny 5y agoI think one can make problem "differentiable" via some heuristics and if you have NN trained to rate code quality and some understanding what should be used for type of problem, memory and speed and than can classify problem to group then rate solution it should be able to guide the process (in competitive programming).
- indiv0 5y agoDo you have a blog or a github or something? This sounds really neat.
- deleted 5y ago[deleted]
- FemmeAndroid 5y agoThis is extremely impressive, but I do think it’s worth noting that these two things were provided: - a very well defined problem. (One of the things I like about competitive programming and the like is just getting to implement a clearly articulated problem, not something I experience on most days.) - existing test data. This is definitely a great accomplishment, but I think those two features of competitive programming are notably different than my experience of daily programming. I don’t mean to suggest these will always be limitations of this kind of technology, though.
- jakub_g 5y ago100% agree. Someone (who?) had to take time and write the detailed requirements. In real jobs you rarely get good tickets with well defined expectations; it's one of most important developer's jobs to transform fuzzy requirement into a good ticket. (Side note: I find that many people skip this step, and go straight from fuzzy-requirement-only-discussed-on-zoom-with-Bob to code; open a pull request without much context or comments; and then a code reviewer is supposed to review it properly without really knowing what problem is actually being solved, and whether the code is solving a proper problem at all).
- ohwellhere 5y agoIs the next step in the evolution of programming having the programmer become the specifier? Fuzzy business requirements -> programmer specifies and writes tests -> AI codes
- buscoquadnary 5y agoThat's all we've ever been since we invented software. First we specified the exact flow of the bits with punch cards. Then we got assembly and we specified the machine instructions. Then we got higher level languages and we specified how the memory was to be managed and what data to store where. Now we have object oriented languages that allow us to work with domain models, and functional languages that allow us to work data structures and algorithms. The next level may be writing business rules, and specifying how services talk to each other, who knows, but it will be no different than it is now just a higher level.
- mcast 5y agoThe year is 2025, Google et al. are now conducting technical on-site interviews purely with AI tools and no human bias behind the camera (aside from GPT-3's quirky emotions). The interview starts with a LC hard, you're given 20 minutes -- good luck!
- jakey_bakey 5y agoI think Amazon already tried this and it had surprisingly racist results
- EGreg 5y agoTo me, coding in imperative languages are one of the hardest things to produce an AI for with current approaches (CNN’s, MCTS and various backpropagation). Something like Cyc would seem to be a lot more promising… And yet, I am starting to see (with GitHub’s Copilot, and now this) a sort of “GPT-4 for code”. I do see many problems with this, including: 1. It doesn’t actually “invent” solutions on its own like AlphaZero, it just uses and remixes from a huge body of work that humans put together, 2. It isn’t really ever sure if it solved the problem, unless it can run against a well-defined test suite, because it could have subtle problems in both the test suite and the solution if it generated both This is a bit like readyplayer.me trying to find the closest combination of noses and lips to match a photo (do you know any open source alternatives to that site btw?) But this isn’t really “solving” anything in an imperative language. Then again, perhaps human logic is just an approaching with operations using low-dimensional vectors, able to capture simple “explainable” models while the AI classifiers and adversarial training produces far bigger vectors that help model the “messiness” of the real world and also find simpler patterns as a side effect. In this case, maybe our goal shouldn’t be to get solutions in the form of imperative language or logic, but rather unleash the computer on “fuzzy” inputs and outputs where things are “mostly correct 99.999% of the time”. The only areas where this could fail is when some intelligent adversarial network exploits weaknesses in that 0.001% and makes it more common. But for natural phenomena it should be good enough !
- qualudeheart 5y agoCan you write more about how Cyc would help? The idea behind Cyc is cool but I don’t think I’ve seen anyone discuss using it for program synthesis.
- jakey_bakey 5y agoAt the risk of sounding relentlessly skeptical - surely by training the code on GitHub data you're not actually creating an AI to solve problems, but creating an extremely obfuscated database of coding puzzle solutions?
- ogogmad 5y agoWe validated our performance using competitions hosted on Codeforces, a popular platform which hosts regular competitions that attract tens of thousands of participants from around the world who come to test their coding skills. We selected for evaluation 10 recent contests, each newer than our training data. AlphaCode placed at about the level of the median competitor, marking the first time an AI code generation system has reached a competitive level of performance in programming competitions. [edit] Is "10 recent contests" a large enough sample size to prove whatever point is being made?
- deleted 5y ago[deleted]
- YeGoblynQueenne 5y agoThe test against human contestants doesn't tell us anything because we have no objective measure of the ability of those human coders (they're just the median in some unknown distribution of skill). There's more objective measures of performance, like a good, old-fashioned, benchmark dataset. For such an evaluation, see table 10 in the arxiv preprint (page 21 of the pdf), listing the results against the APPS dataset of programming tasks. The best performing variant of AlphaCode solves 25% of the simplest ("introductory") APPS tasks and less than 10% of the intermediary ("interview") and more advanced ones ("competition"). So it's not very good. Note also that the article above doesn't report the results on APPS. Because they're not that good.
- qualudeheart 5y agoThat’s been a common objection to Copilot and other recent program synthesis papers. The models regurgitate solutions to problems already encountered in the training set. This is very common with Leetcode problems and seems To still happen with harder competitive programming problems. I think someone else in this thread even pointed put an example of AlphaCode doing the same thing.
- wilde 5y agoOh sweet! When can skip the bullshit puzzle phone screens?
- errcorrectcode 5y agoAli Group CAPTCHA's or Android unlock?
- d0mine 5y agoIt reminds me that median reputation on StackOverflow is 1. All AlphaSO would have to do is to register to receive median reputation on SO ;) (kidding aside AlphaCode sounds like magic) Inventing relational DBs hasn't replaced programmers, we just write custom DB engines less often. Inventing electronic spreadsheets hasn't deprecated programmers, it just means that we don't need programmers for corresponding tasks (where spreadsheets work well). AI won't replace programmers until it grows to replace the humanity as a whole.
- falcor84 5y ago>AI won't replace programmers until it grows to replace the humanity as a whole. Yes, but after seeing this progress in the former, my time estimate of time remaining until the latter had just significantly shortened.
- d0mine 5y agoGiven close to zero chances of a safe AI, I'm optimistic that AI is a much tougher problem and we are not significantly closer to the solution than e.g., in 60s when computer vision was a summer project. There is a progress in certain domains (such as image recognition) but (outside specialized tasks) gigantic language models look like no more than impressive BS generators.
- qualudeheart 5y agoI don’t even think the “will AI replace human programmers” question is that interesting anymore. My prediction is that a full replacement won’t happen until we achieve general artificial intelligence, and have it treat programming as it would any other problem. Elsewhere ITT I’ve claimed that to fully automate programming you also need a model of the external world that’s on par with a humans. Otherwise you can’t work a job because you don’t know how to do the many other tasks that aren’t coding. You need to understand what the business goals are and how your program solves them.
- jdrc 5y agoI think it would be interesting the train a system end-to-end with assembly code instead of various programming languages. This would make it a much more generic compiler
- softwaredoug 5y agoI think CoPilot, etc will be revolutionary tools AND I think human coders are needed. Specifically I love CoPilot for the task of "well specified algorithm to solve problem with well-defined inputs and outputs". The kind of problem you could describe as a coding challenge. BUT, our jobs have a lot more complexity - Local constraints - We almost always work in a large, complex existing code base with specific constraints - Correctness is hard - writing lots of code is usually not the hard part, it's proving it correct against amorphous requirements, communicated in a variety of human social contexts, and bookmarked. - Precision is extremely important - Even if 99% of the time, CoPilot can spit out a correct solution, the 1% of the time it doesn't creates a bevy of problems Are those insurmountable problems? We'll see I suppose, but we begin to verge on general AI if we can gather and understand half a dozen modalities of social context to build a correct solution. Not to mention much of the skill needed in our jobs has much more to do with soft skills, and the bridge between the technical and the non technical, and less to do with hardcore heads-down coding. Exciting times!
- pedrobtz 5y agoWhat about finding bugs, zero-day exploits?
- throwaway5752 5y agoMost people here are programmers (or otherwise involved in the production of software). We shouldn't look at RPA and other job automation trends dispassionately. SaaS valuations aren't were they are (and accounting doesn't treat engineering salary as cost of goods sold) because investors believe that they will require armies of very well paid developers in perpetuity.
- countvonbalzac 5y agowhat?
- a-dub 5y ago> In our preprint, we detail AlphaCode, which uses transformer-based language models to generate code at an unprecedented scale, and then smartly filters to a small set of promising programs if you're using a large corpus of code chunks from working programs as symbols in your alphabet, i wonder how much entropy there actually is in the space of syntactically correct solution candidates.
- ahgamut 5y agoI find almost every new advance in deep learning is accompanied by contrasting comments: it's either "AI will soon automate programming/<insert task here>", or "let me know when AI can actually do <some-difficult-task>". There are many views on this spectrum, but these two are sure to be present in every comment section. IIUC, AlphaCode was trained on Github code to solve competitive programming challenges on Codeforces, some of which are "difficult for a human to do". Suppose AlphaCode was trained on Github code that contains the entire set of solutions on Codeforces, is it actually doing anything "difficult"? I don't believe it would be difficult for a human to solve problems on Codeforces when given access to the entirety of Github (indexed and efficiently searchable). The general question I have been trying to understand is this: is the ML model doing something that we can quantify as "difficult to do (given this particular training set)"? I would like to compute a number that measures how difficult it is for a model to do task X given a large training set Y. If the X is part of the training set, the difficulty should be zero. If X is obtained only by combining elements in the training, maybe it is harder to do. My efforts to answer this question: https://arxiv.org/abs/2109.12075 https://arxiv.org/abs/2109.12075 In recent literature, the RETRO Transformer (https://arxiv.org/pdf/2112.04426.pdf https://arxiv.org/pdf/2112.04426.pdf) talks about "quantifying dataset leakage", which is related to what I mentioned in the above paragraph. If many training samples are also in the test set, what is the model actually learning? Until deep learning methods provide a measurement of "difficulty", it will be difficult to gauge the prowess of any new model that appears on the scene.
- pedrosorio 5y ago> Suppose AlphaCode was trained on Github code that contains the entire set of solutions on Codeforces, is it actually doing anything "difficult"? They tested it on problems from recent contests. The implication being: the statements and solutions to these problems were not available when the Github training set was collected. From the paper [0]: "Our pre-training dataset is based on a snapshot of selected public GitHub repositories taken on 2021/07/14" and "Following our GitHub pre-training dataset snapshot date, all training data in CodeContests was publicly released on or before 2021/07/14. Validation problems appeared between 2021/07/15 and 2021/09/20, and the test set contains problems published after 2021/09/21. This temporal split means that only information humans could have seen is available for training the model." At the very least, even if some of these problems had been solved exactly before, you still need to go from "all of the code in Github" + "natural language description of the problem" to "picking the correct code snippet that solves the problem". Doesn't seem trivial to me. > I don't believe it would be difficult for a human to solve problems on Codeforces when given access to the entirety of Github (indexed and efficiently searchable). And yet, many humans who participate in these contests are unable to do so (although I guess the issue here is that Github is not properly indexed and searchable for humans?). [0] https://storage.googleapis.com/deepmind-media/AlphaCode/competition_level_code_generation_with_alphacode.pdf https://storage.googleapis.com/deepmind-media/AlphaCode/comp...
- FiberBundle 5y agoIt never ceases to amaze me what you can do with these transformer models. They created millions of potential solutions for each problem, used the provided examples for the problems to filter out 99% of incorrect solutions and then applied some more heuristics and the 10 available submissions to try to find a solution. All these approaches just seem like brute-force approaches: Let's just throw our transformer on this problem and see if we can get anything useful out of this. Whatever it is, you can't deny that these unsupervised models learn some semantic representations, but we have no clue at all what that actually is and how these model learn that. But I'm also very sceptical that you can actually get anywhere close to human (expert) capability in any sufficiently complex domain by using this approach.
- bricemo 5y agoWhat do you think then is the difference between going from 50th to 99.9th percentile in their other domains? Is there something materially different between ago, protein folding, or coding? (I don’t know the answer, just curious if anyone else does)
- FiberBundle 5y agoWell with respect to Go the fundamental difference afaict is that you can apply self-supervised learning, which is an incredibly powerful approach (But note e.g. that even this approach wasn't successful in "solving" Starcraft). Unfortunately it's extremely difficult to frame real-world problems in that setting. I don't know anything about protein-folding and don't know what Deepmind uses to try to solve that problem, so I cannot comment on that.
- cjbprime 5y ago> this approach wasn't successful in "solving" Starcraft) Why do you say that? As I understand it, AlphaStar beat pros consistently, including a not widely reported showmatch against Serral when he was BlizzCon champ.
- 5y ago
- erwincoumans 5y agoIt would be interesting if a future 'AlphaZeroCode' with access to a compiler and debugger can learn to code, generating data using self-play. Haven't read the paper yet, seems some impressive milestone.
- zmmmmm 5y agoHas nobody yet asked it to write itself?
- YeGoblynQueenne 5y ago>> AlphaCode ranked within the top 54% in real-world programming competitions, an advancement that demonstrates the potential of deep learning models for tasks that require critical thinking. Critical thinking? Oh, wow. That sounds amazing! Let's read further on... >> At evaluation time, we create a massive amount of C++ and Python programs for each problem, orders of magnitude larger than previous work. Then we filter, cluster, and rerank those solutions to a small set of 10 candidate programs that we submit for external assessment. Ah. That doesn't sound like "critical thinking", or any thinking. It sounds like massive brute-force guessing. A quick look at the arxiv preprint linked from the article reveals that the "massive" amount of prorgams generated is in the millions (see Section 4.4). These are "filtered" by testing them against program input-output (I/O) examples given in the problem descriptions. This "filtering" still leaves a few thousands of candidate programs that are further reduced by clustering to "only" 10 (which are finally submitted). So it's a generate-and-test approach rather than anything to do with reasoning (as claimed elsewhere in the article) let alone "thinking". But why do such massive numbers of programs need to be generated? And why are there still thousands of candidate programs left after "filtering" on I/O examples? The reason is that the generation step is constrained by the natural-language problem descriptions, but those are not enough to generate appropriate solutions because the generating language model doesn't understand what the problem descriptions mean; so the system must generate millions of solutions hoping to "get lucky". Most of those don't pass the I/O tests so they must be discarded. But there are only very few I/O tests for each problem so there are many programs that can pass them, and still not satisfy the problem spec. In the end, clustering is needed to reduce the overwhelming number of pretty much randomly generated programs to a small number. This is a method of generating programs that's not much more precise than drawing numbers at random from a hat. Inevitably, the results don't seem to be particularly accurate, hence the evaluation against programs written by participants in coding competitions, which is not any objective measure of program correctness. Table 10 on the arxiv preprint lists results on a more formal benchmar, the APPS dataset, where it's clear that the results are extremely poor (the best performing AlphaCode variant solves 20% of the "introductory" level problems, though outperforming earlier approaches). Overall, pretty underwhelming and a bit surpirsing to see such lackluster results from DeepMind.
- doctor_eval 5y agoI sometimes read these and wonder if I need to retrain. At my age, I’ll struggle to get a job at a similar level in a new industry. And then I remember that the thing I bring to the table is the ability to turn domain knowledge into code. Being able to do competitive coding challenges is impressive, but a very large segment of software engineering is about eliciting what the squishy humans in management actually want, putting it into code, and discovering as quickly as possible that it’s not what they really wanted after all. It’s going to take a sufficiently long time for AI to take over management that I don’t think oldies like me need to worry too much.
- atleta 5y agoThe thing is that we don't know. What I also have been seeing for a while (like for at least for a decade) that whatever profession seemed to be in danger, whichever profession came out on top on (guess) lists like "these will be replaced by AI soon", each and every one of them thought that it can't happen to them and they all had (and continue to have) explanations, usually involving how that jobs needs human ingenuity. (Unlike all the others, of course :) ) Now completely I agree with you that a significant part of our job is understanding and structuring the problem, but I'm not sure it can't be done in another way. We usually get taking in when we think about what machines will be able to do by thinking that just because we use intelligence (general/human intelligence) to solve the task it means that it's a requirement. Think chess. Or even calculating (as in, with numbers). Or go. Etc. The funny thing is that we don't know, until someone does it. I've been thinking for a while that a lot of what I do could be done by a chat bot. Asking clarification questions. Of course, I do have a lot of background knowledge and that's how I can come up with those questions, but that knowledge is probably easy to acquire from the internet and then use it as training data. (Just like we have an awful lot of code available, we have a lot of problem descriptions, questions, comments and some requirement specifications/user guides.) The hard part would probably be not what we have learned as a software developer, but the things we have learned while we were small kids and also the things that we have learned since, on the side. I.e. being a reasonable person. Understanding what people usually do and want. So the shared context. But I'm not sure it's needed that much. So yeah, I can imagine a service that will talk to a user about what kind of app they want (first just simpler web sites, web shops, later more and more complicated ones) and then just show them "here is what it does and how it works". And then you can say what you'd like to be changed. The color or placement of a button (earlier versions) or even the association type between entities (oh, but a user can have multiple shipping addresses).
- tasubotadas 5y agoI just hope that this shows how useless competitive programming is that it can be replace by the Transformer-model. Additionally, people should REALLY rething their coding interviews if they can be solved by a program.
- thorwawayrus53 5y ago
- aidenn0 5y ago> Creating solutions to unforeseen problems is second nature in human intelligence If this is true then a lot of the people I know lack human intelligence...
- knowmad 5y agoI agree with most of the comments I've read in this thread. Writing code to solve a well defined narrowly scoped problem isn't that hard or valuable. It's determining what the problem actually is and how software could be used to solve it that is challenging and valuable. I would really like to see more effort in the AI/ML code generation space being put into things like code review, and system observation. It seems significantly more useful to use these tools to augment human software engineers rather than trying to tackle the daunting and improbable task of completely replacing them. *Note: as a human software engineer I am biased
- ensan 5y agoWake me up when an AI creates an operating system on the same level of functionality as early-years Linux.
- errcorrectcode 5y agoThat will happen faster than you can conceive because you won't be aware of the progress until it is announced. And, have you tried polling? I hear it keeps the CPU warm in winter. Interrupts are so ... this just in, Nike's stock jump 3% ... Where was I? Did I save my task context properly? Did I reenable interrupts?
- xibalba 5y agoBetween developments like this (and Copilot [Is there a general accepted word for this class of things e.g. "AI Coders"?) and the move toward fully remote, I predict the mean software engineering salary in the United States will be lower in 10 years (in real dollars) than it is today.
- evouga 5y agoI think this is a safe bet, but I would make it with or without the presence of AI Coders. We're clearly in the middle of Tech Bubble 2.0 and it's sure to pop in the next 10 years (and probably much sooner, given the recent crypto and NASDAQ rumblings).
- hnfong 5y agoPeople have been talking about tech bubbles for years, there might be a small financial bubble due to the money printing in recent years but I'm not seeing a big bust coming like dot-com. Tech compensation is probably more influenced by the discrepancy between locations. Once people figure out how to properly handle remote workers and remote teams (which is happening due to Covid), global compensations level will probably level out.
- alasdair_ 5y agoThe interesting stuff happens once AlphaCode gets used to improve the code of AlphaCode.
- errcorrectcode 5y agoAnd this is how we reach the technological singularity and how programmers become as equivalently out-of-demand as piano tuners: self-programming systems. AI will eat any and all knowledge work because there's very little special a human can do that a machine won't be able to do eventually, and much faster and better. It won't be tomorrow, but the sands are inevitably shifting this way.
- rabbits77 5y agoWhat I always find missing from these Deep Learning showcase examples are an honest comparison to existing work. It isn’t like computers haven’t been able to generate code before. Maybe the novelty here is working from the English language specification, but I am dubious just how useful that really is. Specifications are themselves hard to write well too. And what if the “specification” was some Lisp code testing a certain goal, is this any better then existing Genetic Programming? Maybe it is better but in my mind it is kind of suspicious that no comparison is made. I love Deep Learning but nobody does the field any favors by over promising and exaggerating results.
- wantsanagent 5y agoSpecs are hard to write because the person interpreting them may not understand you, but it's the iteration time and cost that kills you. Make me a sandwich -> two weeks and $10k isn't viable Make me a sandwich -> 2 seconds and free, totally viable
- thephyber 5y agoI have fiddled with genetic programming. I don't think there is a good solution for a useful metric for comparing one code generator against another, so I don't think DeepMind should care. Most of the genetic programming results code generated by my algos doesn't compile. Very occasionally the random conditions exist to allow it to jump over a "local maxima" and come up with a useful candidate source code. Sometimes the candidates compile, run, and produce correct results. The time it takes to run varies vastly with parameters (like population, how the mutation function works, how the fitness function weights/scores, etc). Personally I really like that these DeepMind announcements don't get lost in performance comparisons, because inevitably those would get bogged down in complaints like "the other thing wasn't tuned as well as this one was". Let 3rd party researchers who have access to both do that work, independently.
- rabbits77 5y agoIf you think tuning GP parameters are challenging wait until you try tuning hyper parameters for a DL model! It is just a press release, to be fair to DeepMind, and I guess they can promote themselves however they wish. My original comment was more from the context of seeing neural network models in practice perform barely any better, if at all, then classic ML models. Just as those comparisons were revealing similarly I was suspecting this use case may be the same to another classic technique. GP is certainly not the shining star of AI right now but it is actively researched and perusing Google scholar on the subject will show you plenty of interesting, but less heralded, results. There are probably several meaningful metrics for this problem that can be examined. If nothing else it is a simple matter of grading the solutions of each, like a university assignment. Also, typically classical techniques are less resource intensive then any neural network methods; the energy savings alone when considered at production scales would be significant.
- dantodor 5y agoGreat. Now the only thing remaining is POs being able to come with a clear spec and I'm out of job
- derelicto 5y agoHey, honest question: how does one get into competitive programming? I imagine it goes far beyond just leetcoding but honestly i don't even know where to start.
- nsikorr 5y agoI suspect these code generating AIs will bring the singularity at some point in the future. Even if we don’t manage to create an artificial general intelligence, they will. I imagine they will learn to code on super human levels through self play just like AlphaGo and AlphaZero did. This will be awesome.
- deepbream 5y agoThis result is well worth a meme. https://opensea.io/assets/0x495f947276749ce646f68ac8c248420045cb7b5e/38800416672363094847602926489336820944788560867702800329357993734324516552705/ https://opensea.io/assets/0x495f947276749ce646f68ac8c2484200...
- thorwwaskeas 5y agoSince they used the tests this is not something you can do if you don't have a rich battery of tests. Perhaps many problems are something like finite automata and the program discover the structure of the finite automata and also an algorithm for better performance.