26 ms·
OpenAI's Codex sure knows a lot about HN [video]
- 8eye 5y agoopenai as a compiler in the browser would be interesting
- cxr 5y agoHow about just starting with "a compiler in the browser"? From [1]: > the web was first built in the 90s to share complicated academic work People complain a lot about the results of research not being replicable because people withhold their code when they publish, but the fact is that even then it's not guaranteed that anyone will be able to get it to work. Heck, there are plenty of run-of-the-mill software projects (not associated with research) with build processes that aren't replicable without substantial effort in making sure the appropriate toolchain is available and configured for your system. apt-get build-dep is nice and all, but it only goes so far. You'd think that we would have recognized by now that in addition to it being good hygiene to include a project README, a tremendous boon to productivity would result if everyone got on board with also including a document that captured the _exact_ process for transforming source into a binary (or whathaveyou), so you could just drop it into a UVC[2] and get said binary out. Not even mainstream JS programmers (largely writing software that is meant to be interacted with from a web browser!) get this right[3]. Modern JS has managed to grow its own body of implicit knowledge centered around SDKs and setup rituals[4] just like everyone else. 1. http://benschmidt.org/post/2020-01-15/2020-01-15-webgpu/ http://benschmidt.org/post/2020-01-15/2020-01-15-webgpu/ 2. https://scholar.google.com/scholar?hl=en&as_sdt=0%2C44&q=universal+virtual+computer&btnG= https://scholar.google.com/scholar?hl=en&as_sdt=0%2C44&q=uni... 3. https://www.colbyrussell.com/2019/03/06/how-to-displace-javascript.html https://www.colbyrussell.com/2019/03/06/how-to-displace-java... 4. https://news.ycombinator.com/item?id=24495646 https://news.ycombinator.com/item?id=24495646
- Waterluvian 5y agoThis was cute and neat until I connected the dots: natural language means voice APIs for cheap.
- visarga 5y agoText or voice. For voice you need another model. But I watched the demos and I can't wait for my invite. It's even better than GPT3 because this time there is a direct application of the model. I was surprised about how OpenAI sees it: a model learning code as recipes for solving problems. Code is much more exact than natural language, the mix of both is the main advantage. https://www.youtube.com/watch?v=CvgfxH0UZa4 https://www.youtube.com/watch?v=CvgfxH0UZa4
- whazor 5y agoI think voice would have a too high error rate, as you are multiplying voice recognition error rate * codex error rate. However, codex/gpt3 could generate intents and that would be quite cool.
- aardvarkr 5y agoThat’s incredible to watch and really does go to show that a picture (or video) is worth a thousand words.
- andybak 5y agoIn bed listening to a podcast with my partner so unless i remember this post tomorrow I’ll never know.
- MagicWishMonkey 5y agoAny tips on getting this to run as an extension?
- tectonic 5y agoIt's not currently open source, but I might release it if I can get it cleaned up.
- dang 5y agoThe submitted URL was https://twitter.com/tectonic/status/1426980192317177859 https://twitter.com/tectonic/status/1426980192317177859 but the video seems like the real submission here, so I changed it. I also changed the title to a nice representative phrase from the video.
- deleted 5y ago[deleted]
- leereeves 5y agoHeh, codex has a sense of humor. When asked to add "a url for the video on YouTube", codex added the url below. I won't spoil the surprise, but it's not the video linked in the OP: https://www.youtube.com/watch?v=dQw4w9WgXcQ https://www.youtube.com/watch?v=dQw4w9WgXcQ
- leppr 5y agoSo the question is whether this is real or just a troll.
- tectonic 5y agoIt's real. I was totally surprised when that was the URL it picked.
- YeGoblynQueenne 5y agoYou asked it to do something with "the video on youtube" but what does "the video" refer to? It seems the most likely url associated with the phrase "the video on youtube" is, well, that. So basically it failed at anaphora resolution. Seen another way, you asked it for "the video" and so it gave you the video.
- riveducha 5y agoI made a similar little GPT-3 toy (for Linux/bash) that also ended up generating a Rickroll URL.[0] In vague cases, GPT-3 functions approximately as well as a markov chain - it will give you the most “probable” sequence of tokens. There’s some implementation details regarding how it deals with tokenization, but unless you really increase the randomness (“temperature”) you’re likely to get very popular YT video IDs. In other cases, it will “hallucinate” things like URLs, UUIDs, hashes, and other things that are basically a random string of characters. In my experience it will make UUIDs that have the right number of characters but seem suspiciously non-random and don’t fit any of the defined UUID formats. Fun stuff. [0] https://youtu.be/j0UnS3jHhAA https://youtu.be/j0UnS3jHhAA
- monkeydust 5y agoBeen playing around with codex over the weekend as a on developer. Certainly impressive and also occasionally frustrating when you push it. The natural language to SQL are still the best and most consistent demos.
- mritchie712 5y agoAny SQL demos you can point me to?
- monkeydust 5y agoI created one using Codex here, hope it helps illustrate what is possible: https://www.loom.com/share/7a6ed39dcd0749aa87e93067624506d9 https://www.loom.com/share/7a6ed39dcd0749aa87e93067624506d9
- IdiocyInAction 5y agoKind of ironic, AFAIK SQL was first envisioned as being so natural language-like that it could be used by non-programmers.
- ineedasername 5y agoFor simpler queries I suppose it is. Select, join, filter, group sort... maybe a secondary sort on an aggregate (HAVING). Full on DBA work is much more complicated, as are complex queries, especially against production databases. But against denormalized data and fact tables it can still be pretty accessible from a natural language standpoint, at least in my limited experience teaching a few people.
- tectonic 5y agoHere's the entirety of the prompt: <|endoftext|>/* This code is running inside of a bookmarklet. Each section should set and return _.*/ // The bookmarklet is now executing on example.com. // Command: The variable called _ will always contain the previous result. let _ = null; /* Command: Add a new primary header "[PAGE TITLE]" by adding an HTML DOM node */ (() => { let newHeader = document.createElement('h1'); newHeader.innerHTML = '[PAGE TITLE]'; document.body.appendChild(newHeader); _ = newHeader; return newHeader; })() /* Command: Find the first node containing the word 'house' */ (() => { let xpath = "//*[contains(text(), 'house')]"; let matchingElement = document.evaluate(xpath, document, null, XPathResult.FIRST_ORDERED_NODE_TYPE, null).singleNodeValue; _ = matchingElement; return matchingElement; })() /* Command: Delete that node */ (() => { _.parentNode.removeChild(_); return _; })() /* Command: Change the background color to white */ (() => { document.body.style.backgroundColor = 'white'; _ = document.body; return document.body; })() /* Command: Select the contents of the first pre tag */ (() => { let node = document.querySelector('pre'); let selection = window.getSelection(); let range = document.createRange(); range.selectNodeContents(node); selection.removeAllRanges(); selection.addRange(range); _ = selection; return selection; })() // The bookmarklet is now executing on [PAGE URL]. It is customized for [PAGE TITLE] and knows the correct CSS selectors and DOM layout. let _ = null; /* Command: [USER INPUT] */
- tvirosi 5y agoThis might totally work and it's kind of impressive if it does. I'm still biased towards ultra skepticism towards all of this since the trustworthiness of all demos like this is completely corrupted at this point due to cherry picking and other deceptive tricks.
- tectonic 5y agoI had to try a few times to get the prompt right, but that's the limit of the cherrypicking. You're correct that it doesn't work nearly as well on more complex, less temporally stable sites like Reddit.
- sxp 5y agoThe skepticism is warranted for any bleeding edge technology. I wonder if there's another version of a Turing test when a technology can be considered sufficiently advanced when it's indistinguishable from a fake version you've seen in sci-fi. E.g, the Boston Dynamics' dancing robot video (https://www.youtube.com/watch?v=fn3KWM1kuAw https://www.youtube.com/watch?v=fn3KWM1kuAw) still looks fake to me because it's at the level that I would expect to see from Hollywood CGI rather than a real tech demo. If I saw the video anywhere else but on the BD page, I would have enjoyed it and forgotten about it since it's an average CGI video.
- OnlineGladiator 5y agoI genuinely don't understand your position. Are you saying a tech demo is only impressive if it can do things that can't be simulated? What can't be shown via simulation or CGI with enough time and money today? If we're limiting ourselves to video there's no interactive component. Even though that dancing video likely had hundreds of takes, the part that makes it impressive is that it's real. I swear I'm not trying to be disagreeable here - I honestly don't understand your perspective.
- MrOrelliOReilly 5y agoI think what the author is trying to say is that if a technology is sufficiently advanced it seems like it can’t be real, meaning it’s something only possible with CGI. So we see these dancing robots, think “just more CGI”, then are astounded when we find out it’s real
- archibaldJ 5y agothanks for the info! great stuff! gpt3's generalization-by-description never ceases to amuse me; but the difficult thing here is to get the right abstraction layers layered nicely in the conceptual lasagna. This is where category theory becomes extremely powerful. It has occured to me that codex-davinci has an intuitive "understanding" of constructs like monads, or something along that line.
- tectonic 5y agodebuild.co looks cool. Using Codex yet?
- archibaldJ 5y agoyes; here is our latest demo https://twitter.com/sharifshameem/status/1425185575645024256 https://twitter.com/sharifshameem/status/1425185575645024256
- Y_Y 5y agoCan you expand on the utility of categories here? There's a lot of space between knowing what defines a monad, when something might be a monad, what you can do with monadic structure etc. Of course if an AI truly understood monads I it would be a bright line marking where the machines have finally surpassed the human mind. Cool.
- archibaldJ 5y agoI think it's closely linked to the notions of semantics as approached in CS (ie. PLT) v.s. in linguistics where we are mostly concerned with the "micro-structures" and "meso-structures" that gave rise to qualia we humans experience (e.g. therefore languages with different structural systems such as English vs Chinese encode concepts (as well as intentions) very differently; (for an illustration, see Interality as a Key to Deciphering Guiguzi: A Challenge to Critics [1])), and not so much about how evaluation and execution came about (e.g. as studied from a compiler's persective in denotational semantics, or a more functional perspective in operational semantics, where things like natural transformations are ubiquitous) [1]: https://cjc-online.ca/index.php/journal/article/view/3187/3258 https://cjc-online.ca/index.php/journal/article/view/3187/32... And so categories naturally come in as a way to bridge and compose these two worlds, and that's just the beginning. There're so much more we don't understand yet, such as what understanding really is given a certain set of contexts and constraints as well as in regards to their relaxations. What is understanding if there are no doings? And what is doing with no understandings? How do things compose? These are great mysteries.
- nathan_phoenix 5y agoDoesn't this only work so well on HN only because HN uses really simple html and css? What about more complex sites?
- tectonic 5y agoIt's much less reliable on sites like Reddit, although it can usually handle "click on the profile link" or "delete all images" and stuff.
- nathan_phoenix 5y agoOkay, thanks for the info.
- 37ef_ced3 5y agoHow does Codex learn the relationship between English and code? Is it purely through the comments in the training corpus?
- mediumdeviation 5y agoIt's really interesting. HN's HTML is very un-semantic and is actually quite hard to work with. <tr class="athing" id="28191639"> <td class="title" valign="top" align="right"><span class="rank">9.</span></td> <td class="votelinks" valign="top"><center><a id="up_28191639" onclick="return vote(event, this, "up")" href="vote?id=28191639&how=up&auth=****&goto=news"><div class="votearrow" title="upvote"></div></a></center> </td> <td class="title"> <a href="http://be-n.com/spw/you-can-list-a-million-files-in-a-directory-but-not-with-ls.html" class="storylink">You can list a directory containing 8M files, but not with ls</a> <span class="sitebit comhead"> (<a href="from?site=be-n.com"><span class="sitestr">be-n.com</span></a>)</span> </td> </tr> In the video Codex picks up tr.athing as a news item. I wonder if this is actually generalized learning, or if it just picked the selector up from eg. a userscript that appeared in its training corpus. Another thing that's kind of scary (and makes it worrying if this is used for Copilot) is the second prompt to make the text uppercase results in code that is superficially correct, but is very semantically wrong - innerHTML.toUpperCase() is dangerous because it not only makes the content uppercase, it also modifies the attributes on the HTML elements inside. This definitely broke the vote button, which uses inline JS which is case sensitive. It also destroys any attached event handler since the elements are basically deleted then re-created. The correct way to do this is to either use CSS text-transform: uppercase, or if it is important to update the DOM itself, recursively descend and update childNodes with nodeType == text's nodeValue to uppercase.
- goatlover 5y agoI wonder why innerHTML has a toUpperCase method. It makes sense for innerText of course, but case sensitivity in the html can definitely matter for JS and CSS. I'm guessing because both are just treated as JS string objects. But there is a special NodeList collection, so why not a special HtmlString?
- astrea 5y agoWelp, where will all of us end up when this gets sufficiently complex?
- deleted 5y ago[deleted]
- 37ef_ced3 5y agoCode writers and prose writers will be reduced to operating the AI (checking its output, trying various inputs to elicit the desired language text). At least we won't be completely obsolete like the taxi drivers and Lee Se-dol: The South Korean Go champion Lee Se-dol has retired from professional play, telling Yonhap news agency that his decision was motivated by the ascendancy of AI. “With the debut of AI in Go games, I’ve realized that I’m not at the top even if I become the number one through frantic efforts,” Lee told Yonhap. “Even if I become the number one, there is an entity that cannot be defeated.” To speed your obsolescence, make sure you use Codex in your work, so it can learn you completely. Remember, you won't be able to compete with people who use Codex, so you have to feed the machine, whether you like it or not.
- amrrs 5y ago05:39 https://youtu.be/tNcBQBTeyf4 https://youtu.be/tNcBQBTeyf4 You can see how OpenAI Codex misses some details about HN scraping. What's impressive that you might notice is the variable names it chooses which seems to show the nature of HN scraping codes on the internet
- muzster 5y agoif you listen carefully you can hear the music...
- Zenst 5y agoI've looked at some demo's of OpenAI Codex and it's pretty impressive start for sure. Something like this tied into R and a whole level of data analysis would become far more accessible to those with business knowledge who don't really want to learn the nuances of tools. But I must say, having lived thru the 80's fad of code generating sudo 4gl's, the code this produces is pretty darn good indeed. Now when something like this can handle a Google coding exam - that's going to be an epic milestone. Though old coding exam questions would equally offer up some great material to push this thru it's paces.
- jimmySixDOF 5y agoMorgan McGuire, the Chief Scientist for Roblox, was on a panel at SIGGRAPH last week where he described one goal of their R&D to be basically "taking a half page written description from the user and autogenerating their desired 3D game experience". [1] https://twitter.com/CasualEffects/status/1425152593945321476 https://twitter.com/CasualEffects/status/1425152593945321476
- cush 5y agoImagine trying to do this on a normal site where an input is controlled and nested in 300 divs
- aaron695 5y agoThis demo wouldn't have been out of place in the 80's. Maybe everyone is smarter now and is is looking at some sort of underlying process. Or maybe it's just more of the same. It make no sense to auto fill 'the video' The correct answer is I don't understand. That was a mistake. It also bold'ed the (site) which is not correct. It's a short demo that clearly would have had many test runs. The fact it 'learned' to do a bad Behat is amazing. But there's no reason to think it can equal Behat in 10 years time. Chess AI had a way forward, it's not clear this does.
- naveen99 5y agoopen source attempt at a clone,(not by me…) https://github.com/CodedotAl/gpt-code-clippy https://github.com/CodedotAl/gpt-code-clippy
- deleted 5y ago[deleted]
- zwright 5y agoholy shit
- derac 5y agomaybe see if it can beat a simple forum captcha of the 1+1= variety
- cortexio 5y agoin before it starts posting swastikas on /pol.