15 ms·
OpenAI Codex
- temp8964 5y agoCan this read existing code and fix one missing piece? That will be cool. Say I have a question I can't solve by searching through stackoverflow. If the AI can solve a problem like that, it will be great.
- priyanmuthu 5y agoProgram Synthesis can do some rudimentary fixes. But I would love to explore this problem of program correction using AI.
- refulgentis 5y agoI'm trying to extract some signal from this link...lots of upvotes, no comments, 30 min old, top 3 on HN...I'm worried this will be read as negative, but it's not, just learning, and enough time has passed I'm itching to jump in and ask: - Is the significance here exactly what it says on the tin: the model behind GitHub's AI code completion will be shared with people on an invite basis? Or am I missing something? - What is the practical import of the quote at the end of this comment? "can now" makes me think its a new feature over Github's implementation, which would then indicate the "simple commands" could be general UI, or at least IDE UI, navigation. If "can now" means "it is currently capable of, but will be capable of more", then I'd expect it to be the same as the current implementation on Github. Quote: "Codex can now interpret simple commands in natural language and execute them on the user’s behalf—making it possible to build a natural language interface to existing applications."
- sbierwagen 5y agoTake a look at the video demo. It takes natural text in a box and generates code. Copilot was super-autocomplete, so the interface was writing code in an IDE that it filled out for you. Natural language interface will be a little easier for non-programmers. (Though, how would you read the code to make sure it does what you meant...)
- polyanos 5y ago>Take a look at the video demo. It takes natural text in a box and generates code. Copilot was super-autocomplete, so the interface was writing code in an IDE that it filled out for you. No it wasn't, you can literally describe, in natural text, what you want in a comment and CoPilot will do its best to generate a complete method based on that comment. It seemed like it was so auto-compltely because that focussed on the "helping the developer" part. I'm fairly sure CoPilot could have shown something similar if they had a demo where you could make something visual easily, like HTML + Javascript/Typescript/whatever scripting language. They're using exactly the same model (Codex) after all.
- am17an 5y agoI really want to just play with this tech- it’s frightening but also the future, but I’m still waiting to be accepted on the GitHub copilot waitlist. I wonder how long this will take for people who don’t know someone who knows someone…
- deleted 5y ago[deleted]
- andyxor 5y agothat's not the future, these large language models have no understanding of language, they repeat the most frequently occurring patterns like parrots. They miss this whole thing called semantics.
- febrilian 5y agoUhh... I'm literally no one but got the access for like a week or so. I got 134 repos and 12,060 contributions in the last year. Idk if that mattered.
- northfoxz 5y agoA new way to talk to the computer I guess.
- z77dj3kl 5y agoI thought OpenAI was originally supposed to be some kind of for-the-good, non-profit institution studying AI and its safe use in particular with an effort to make it more accessible and available to all through more open collaboration. This is cool research, sure; but what happened to making models available for use by others instead of just through some opaque APIs? Maybe I'm just remembering wrong or conflating OpenAI with some other entity? Or maybe I bought too much of the marketing early on.
- Buttons840 5y agoNo, they did some good, they've done a few things to personally help me. They created OpenAI Gym which is a great help when doing reinforcement learning research and defined the standard interface for reinforcement learning libraries for a generation. But they not longer maintain OpenAI Gym. They also created Spinning Up [0], one of the best resources I've found for learning reinforcement learning. Their teaching resources are detailed but relatively brief and are focused on implementing the algorithms, even if some of the "proofs" are neglected. But they no longer maintain Spinning Up. So yes, originally they were for-the-good, but lately I've noticed them moving away from that in more ways than one. It seems they learned one cool trick with language sequence modelling, and they have a lot of compute, and this is all they do now. [0]: https://spinningup.openai.com/en/latest/ https://spinningup.openai.com/en/latest/
- blt 5y agoThat was the marketing message. They became for-profit in 2019 and took investment from Microsoft. Many people were skeptical before that because the main investors were mostly known for for-profit ventures.
- webmaven 5y agoYou're remembering correctly. OpenAI transitioned from non-profit to for-profit in 2019, took about $1 billion from Microsoft (there has been speculation that this was mostly in the form of Azure credits), and announced that Microsoft would be their preferred partner for commercializing OpenAI technologies: https://openai.com/blog/microsoft/ https://openai.com/blog/microsoft/
- 5y ago
- amrrs 5y agoI feel that OpenAI Codex could become like Webflow for coding. It might sound ironic, but what tools like Webflow in the world of Web programming is to give the power of creators to build something fast that can long last (without the speciality of a decent web programmer). If the same thing can happen in the world of programming, I guess evaluations like LeetCode and Whiteboarding can go away and bring in a new of logical thinking evaluation which could ultimately be a more realistic method of some programmer rising above the chain.
- Vermeulen 5y agoA warning to devs building on OpenAI APIs: We spent months developing a chatbot using GPT3 for our game and released a video showcasing it: https://www.youtube.com/watch?v=nnuSQvoroJo&t=264s https://www.youtube.com/watch?v=nnuSQvoroJo&t=264s Afterwards OpenAI then added GPT3 chatbot guidelines disallowing basically anything like this. We were in communication with them beforehand, but they decided later that any sort of free form chatbot was dangerous. What they allow changes on a weekly basis, and is different for each customer. I don't understand how they expect companies to rely on them
- option_greek 5y agoThat's a really interesting demo. What makes the responses so laggy? Does the model take that long to generate text? You can also experiment with things like repeating the user question or adding pauses like "hmm let's see" to make it less noticeable at least some of the time. Too bad they asked you to pull it. What's the danger they are worried about? Annoying thing from their press releases is how seriously they take their GPT3 bot impact on humans. Despite all the hype, it's difficult to see the end of humanity by GPT3 bots any time soon. Honestly they need to rename themselves - can't see what's open about openai.
- Vermeulen 5y agoIt's laggy since it needs to do speech to text, gpt3 text response, then text to speech. Not sure what adds the most latency actually. They only allow gpt3 chatbots if the chatbot is designed to speak only about a specific subject, and literally never says anything bad/negative (and we have to keep logs to make sure this is the case). Which is insane. Their reasoning to me was literally a 'what if' the chatbot "advised on who to vote for in the election". As if a chatbot in the context of a video game saying who to vote for was somehow dangerous I understand the need to keep GPT3 private. There is a lot of possibility for deception using it. But they are so scared of their chatbot saying a bad thing and the PR around that they've removed the possibility of doing anything useful with it. They need to take context more into account - a clearly labeled chatbot in a video game is different than a Twitter bot
- f0e4c2f7 5y agoThey just finished a demo on twitch. Pretty crazy! https://www.twitch.tv/videos/1114111652 https://www.twitch.tv/videos/1114111652 Starts at 15:45.
- raidicy 5y agolmao; copyright muted so you can't even hear them speaking.
- deleted 5y ago[deleted]
- karmasimida 5y agoIt is simultaneously impressive and underwhelming for me. I mean yes this is a super impressive demo, but it didn't go beyond my expectation. I really want to see whether this model can write a correct binary search method without seeing one before. Or even correctly using the binary search, does it understand concept like index boundaries?
- stavros 5y ago> I really want to see whether this model can write a correct binary search method without seeing one before. I don't believe the model was trained on Google interview answers, sadly.
- polyanos 5y agoI found the whole UI/sandbox they created the most interesting part. Now don't get me wrong, the tech is certainly great and all, but I really didn't had the feeling I watched/learned more than I already knew from what was shown with Github CoPilot, although I was kinda impressed, if it really is as simple as they stated, at how it is able to adapt to new apis. It's a shame they only limited the demo to relatively simple instructions.
- wrinkl3 5y ago> I really want to see whether this model can write a correct binary search method without seeing one before. It has almost definitely seen a lot of coding problems so I would expect "write a function to binary search a sorted array" to output the intended result. I don't think anybody expects it to come up with algorithms it hasn't encountered.
- throwaway128327 5y agoI don't understand what is going on, why are people even spending time on this? I think this and copilot and etc are solving a non problem of "we will remove the boring part of programming" by generating a bunch of code, so now it's even more boring to read it and check if it actually does what you want. In the same time zero of the developers I interviewed know how a linked list is laid out in memory, or what is the pro/con of continuous memory layouts, or even how a cpu works actually. Maybe those things are not needed anymore, but I see their code... I think it will be better if they know them.
- parksy 5y agoThis is just nascent technology leading toward something like this: "Computer, I want to play a game." "Okay, what will the game be?" "I want to be a starship captain, give me a cool space ship I can explore the galaxy with" "Okay... like this?" "Not quite, make the galaxy more realistic, with real stars and planets. Also make it 3d. I want to be the captain inside the ship." "How about now?" "Cool, and there should be space stations I can visit near planets, and I can fly my ship to stars with hyperspace. Make it so I have to trade for fuel at the space stations, maybe I need to mine asteroids or search derelict space ships for treasure. I want to play with my friends too, they can have their own ships or walk around my ship." "Done, was there anything else?" "Yes, add different alien races to some of the star systems, and make some of them have alliances. I want to talk to the aliens about their history and culture. Sometimes aliens are unfriendly and we'll have space battles if talking doesn't work. Make it so I can command a fleet and call for reinforcements." "Processing... Done. Anything else?" "Actually this is boring, can we start over?" "Game erased. Please provide new prompt."
- throwaway128327 5y agoOh! this will be so cool! do you really think it could lead in that direction? To me it seems more like a metaphysical cargo cult. I think I am too pessimistic, I should shake it off, nothing good comes out of being pessimistic (by definition). Thanks for the inspiration!
- 5y ago
- abeppu 5y agoI'm still surprised by the approach. I mean, great that it works this well -- but program synthesis is one of those rare domains where you can observe exactly what the outcome is after you generate something. You can see execution traces, variable values, what the JIT produced, etc. And all of this is relatively cheap -- often executing a code snippet should be far cheaper than an extra pass through a giant DNN right? So it's fascinating to me that they train entirely from dealing with code as text. Imagine learning to develop recipes, not by ever cooking or eating or even seeing food, but only reading a giant library of cookbooks. Or learning to compose music but never hearing or playing anything -- only seeing scores.
- deleted 5y ago[deleted]
- wantsanagent 5y agoFWIW execution guided code synthesis is a thing. Get a few possible outputs and ditch those that don't pass a parser as an example. At least in the SQL generation realm this is well worth the time it takes to tack onto a large language model.
- maxwells-daemon 5y agoThe "language models don't really understand anything" corner is getting smaller and smaller. In the last few months we've seen pretty definitive evidence that transformers can recombine concepts ([1], [2]) and do simple logical inference using contextual information ([3], "make the score font color visible"). I see no reason that this technology couldn't smoothly scale into human-level intelligence, yet lots of people seem to think it'll require a step change or is impossible. That being said, robust systematic generalization is still a hard problem. But "achieve symbol grounding through tons of multimodal data" is looking more and more like the answer. [1] https://openai.com/blog/dall-e/ https://openai.com/blog/dall-e/ [2] https://distill.pub/2021/multimodal-neurons/ https://distill.pub/2021/multimodal-neurons/ [3] https://openai.com/blog/openai-codex/ https://openai.com/blog/openai-codex/
- karmasimida 5y ago> The "language models don't really understand anything" This is still true. By all account, human doesn't need to read 159GB of Python code to write Python, or we simply can't. But it doesn't necessarily indicate language models aren't useful.
- maxwells-daemon 5y agoI would argue humans ingest a lot more than 159GB before they can write code. Most of it isn't Python, and humans currently transfer knowledge a lot more efficiently than NNs, but I suspect that'll change as incorporating more varied data sources becomes feasible.
- pnt12 5y agoWe generalize pretty well. One could say: "it took you 20 years to learn python!", but actually I learned python, Java, c#... Software engineering, machine learning... How to play guitar, how to cook.. .How to speak Portuguese, how to speak English... And thousands and thousands of different things which build on each other. You can give a programmer a few kb of code in a new language and that will give him a small grasp of how it works.
- GistNoesis 5y agoCan I use this to write solidity contracts ?
- mxwsn 5y agoThat has got to be one of the worst possible use cases one could imagine. In page 33 of the appendix, the authors note that nearly 40% of RSA encryption keys created by Codex are clearly insecure.
- woah 5y agoWhat does that have to do with anything?
- GistNoesis 5y agoOnly if tokens have value. If codex is able to handle a generic api from reading the doc, it maybe could use a python library for solidity contracts like https://web3py.readthedocs.io/en/stable/contracts.html https://web3py.readthedocs.io/en/stable/contracts.html As a contract user, I'd probably have more trust in a contract written by an independent AI from a short natural language specification which can't hide intent, than a contract with hidden backdoor, or a subtle bug. Also the AI will probably improve with usage. You probably can generate multiple version of your contract, and maybe a high level bug correction scheme like taking the median action between those version can increase bug robustness and find those edge cases when action differ.
- wrinkl3 5y agoThat's about as advisable as asking it to write firmware for a pacemaker. Smart contracts are some of the most delicate codebases - even a tiny bug can cause you to lose a lot of money. With a model like this bugs are very likely, especially in such a niche domain.
- GistNoesis 5y agoI agree that at first it seems to be a bad idea, but the more I think about it, the more I think it makes for a great test case to prove the AI is robust enough and that it can be trusted for real life scenarios. It also feels kind of inevitable. It seems like a safe playground, maybe it will lose a few tokens, but if they are no legal consequences, who cares ? The best place to move fast and break things, and with great reward potential. If I take some freelancer to write them what's telling me that they are not already using something like codex or copilot. It's like when a factory release its used water in the river nearby. Maybe we shouldn't drink from the river anymore, but it would be better to test the water the factory release to make sure it's OK.
- vincnetas 5y agoWill really be impressed when one could say: “here is this codebase, modify this function so that it would preduce [insert desired efect]” and also other functionality of project would not crash thumbling down… Because writing code from scratch now is i think much rearer than improoving existing codebases. Aka bugfixing.
- vincnetas 5y agoAlso curious what this ai would produce when provided with contradictory requests. Because often there are multiple requirements which on theyr own sounds reasonable but when you try to fit all requirements in one system, things get nasty.
- polyanos 5y agoIt is only able to translate small instructions into code. I think it will take a while to get to a situation where you can just give it a list of requirements and it spits a working program. Hell it messed up when they gave it the instruction "make every fifth line bold" in their Word api part of the demo, where it made the first line of every paragraph (which is only 4 lines long in total) bold instead of every fifth line.
- librexpr 5y agoIt didn't mess up the instruction "make every fifth line bold". The blank spaces between each "paragraph" are empty lines, so it counted them too. I think this is perfectly reasonable behavior, it's what I would have done absent further instructions too. You can see it in the generated code on the bottom right during that part of the demo. It loops over the lines and bolds them when index % 5 == 0. Edit: I guess with the 1-based indexing of natural languages, the code actually bolds lines number 1, 6, etc. So arguably it should have done index % 5 == 4 instead, to bold lines number 5, 10, etc. But funnily enough, if it had done that, it would have bolded all the empty lines, so it would have seemed like it didn't do anything.
- mark_l_watson 5y agoI watched their 30 minute demo on Twitch this morning, really good! I use their OpenAI beta APIs as a paying customer, I am still waiting for access to Codex.
- jmportilla 5y agoVery cool, will be interesting to see if this is ever added in to VisualStudio as some sort of "super" auto-complete.
- 3wolf 5y agoI think integrations like the MS Word example they show off at the end of the live demo have the potential to be even more impactful than just generating code for programmers.
- polyanos 5y agoThat still needs work though, it messed up the "Make every fifth line bold" pretty bad. Still, it showed it could adapt to a new API pretty well.
- 3wolf 5y agoYeah, definitely. I guess my point was that converting natural language to source code can be even more valuable for people who don't know how to code, but want to perform actions more complicated than a simple button press. For example, I often find myself doing regex based find-and-replace-alls in text files, and even that feels inefficient while also being over the head of the vast majority of users. I'd imagine there are a lot of people out there spending many hours manually editing documents and spreadsheets.
- jonesn11 5y agoHow did it mess up the "make every fifth line bold" prompt? Also, to follow up on the original comment, AI demos are nice, but being a student of history there are still fundamental challenges with these systems. My skepticism is in how much prompting is really required and how can it understand higher level semantics like code refactoring, reproducible examples, large scale design patterns etc. This synthesis of sequential symbolic processes and probabilistic neural generation is really exciting though. When the amount of human code edits and tweaking for complex programs goes down from hours to seconds then that's when I'll be impressed and scared.
- amalive 5y agoWould like to say "Fix that something of undefined error" some day.
- dmurray 5y agoThey should have released this first instead of GitHub Copilot. The focus would then have been much more on "look at the cool stuff they can do" rather than "Microsoft is releasing a product that plagiarizes GPL code". Once people had digested that and there had been a few other proof-of-concept business ideas around turning Codex into a SaaS (because some people will always queue to build their product on your API), announce the evil version. Not that I really think Copilot is evil, but the IP concerns are legitimate.
- pnt12 5y agoYes, its very strange to announce the product first and then the research.
- leesec 5y agoThe Writing On The Wall
- mensetmanusman 5y agoIf this actually worked, wouldn’t that be amazing? If you could break down a software idea into a blue print of concepts that need to be accomplished, and then dictate what should be done… I doubt it works, but I wonder how many decades from now we will be able to walk through a finite number of simple requests and wrap them together as working software. Then people will be able to convert their blueprint into action!
- delsarto 5y agoIn the demo video on https://openai.com/blog/openai-codex/#spacegame https://openai.com/blog/openai-codex/#spacegame it seems like it goes and does an image search for a picture of an asteroid, and them embeds without attribution a direct link to a image hosted on "d.newsweek.com". Not sure I'd call that a resounding example of generating good code...
- marstall 5y agoHow does Codex make the connection between natural language and code? Comments? variable names? file names? docs?
- tyingq 5y ago"Converting Python to Ruby with OpenAI Codex" Oof. Looking forward to maintaining some future ports done with this tool.