6 ms·
Amazing if this is only a 12B model. If this already increases coding productivity by up to 50% (depending on kind of work), imagine what a 1T model will be cap
by peaslock 4y ago
Amazing if this is only a 12B model. If this already increases coding productivity by up to 50% (depending on kind of work), imagine what a 1T model will be capable of! I do wonder if some programmers at FAANG are already having access to a way more powerful coding assistants, and whether they code much at all at this point, or only make high level code specifications and then fix up the automatically generated code.
- gear54rus 4y ago'fix up generated code' but do you agree that finding a mistake (without even knowing if it's there) might be even harder than writing from scratch?
- PetahNZ 4y agoWe are doing this all the time anyway during code reviews.
- jrockway 4y agoIt's likely that programmers have this skill somewhere. We all make mistakes when typing in code, and many of them do get found. Some of them don't, that's what we call a bug. So AI isn't exactly breaking any ground here. I played with ChatGPT and asked it interview questions, and I thought it was a pretty interesting exercise to find its mistakes and get it to fix them. Good tool for training interviewers, perhaps.
- whazor 4y agoIn my eyes, the limitation of these models is that they only fit a limited amount of context. Not the complete API of your code base, or the latest version of the libraries you are using. I also don't believe a bigger model would resolve these limitations. However, I do believe there could be a meta model that can query code and libraries.
- IshKebab 4y agoPresumably if you had access to them you could fine tune them on your codebase.
- peaslock 4y agoYeah, continuous online learning by fine-tuning seems like an obvious way of making these models recall information from outside the perceptible context. One could also prompt the model to (recursively) summarize code and prepend this summary to each prompt, and/or enable the model to interactively query function definitions or code summaries before outputting a final answer (trained by RLHF). But any such tricks might also quickly be outcompeted by an even more general model, e.g. one that directly controls the GUI and can communicate with coworkers...
- karmasimida 4y agoMicrosoft is FAANG level and beyond.
- easygenes 4y agoMicrosoft paid for early exclusive access to GPT-3 internals. They're using it to develop things like Power Apps. FAANG are all doing similar and Google in particular at least purports to have models that outperform what OpenAI is doing.
- tjoff 4y ago> If this already increases coding productivity by up to 50% (depending on kind of work) Does anyone believe that? edit: I'm surprised to see that (so far) 3 replies actually agree with the statement. Is there a video that you'd recommend that shows realistic usage and gain from copilot? Maybe a livestream or something.
- insanitybit 4y agoSure. I'm way more productive with Copilot. I haven't been coding much lately but I could imagine it would double my productivity with regards to the actual "get an implementation of a thing done" bit of the work. In terms of design, I had a long conversation with ChatGPT the other day about designing a database, including optimizations that could be made given certain requirements and constraints, etc. It was a big productivity boost, like rubber ducking on steroids.
- BonoboIO 4y agoCan you give us an example how it helped to design the database? I could not think how it would have helped me, but maybe I m limited in my imagination or don’t know how to ask.
- insanitybit 4y agoI told it I was designing a database. I told it that my database could tolerate failure levels where more than a quorum of nodes failed at a given time. I then asked it about different algorithms for consensus; RAFT, Paxos, swarm based, etc. It described algorithms for me. I told it that in my database I could guarantee certain things, like that every operation commutes, and I asked how that would let me optimize things - it explained that I could paralellize certain parts of those algorithms. At one point I told it to name the algorithm we had been discussing something like "OptSwim" and we just kept iterating on the idea.
- dboreham 4y agoWell then. The singularity is here. Almost no humans understand these things.
- varunkmohan 4y agoA 1T model would be capable of much more than what the current version of Copilot in terms of autocompletion and even code correction. However, at that point, even with a lot of model parallelism to speedup inference, it's likely to be atleast 10x slower on the generation side. From my experience working on Codeium, a Copilot alternative, this would be too frustrating for users. It could be useful as a tool that runs asynchronously that modifies all your code at scale.
- google234123 4y agoIt could be interesting if it was an alternative that a user could query. I could imagine someone starting to write a new function might be willing to wait 10x more time to get something better.
- varunkmohan 4y agoVery true, I think the issue though is unless that output is very likely to be 100% correct, a user would always prefer something that is incomplete but quicker to iterate on. It would be interesting to see if we can get to a paradigm like that.
- csomar 4y agoGiven how fast Copilot is (a few seconds), I wouldn't mind waiting 10x. I also wouldn't mind letting it run overnight for some tasks (ie: write documentation, write tests, suggest bug fixes, etc...). Will check on my buddy on the next morning.
- throwaway888abc 4y agoThat sounds like modern day outsourcing
- thakkarparth007 4y agoI think the UX of large suggestions will require a lot of thinking and experimentation. That's because the longer the output of such model, higher the risk of it making some mistake. For short completions, it's often easy to identify mistakes from useful suggestions (though sometimes subtle bugs slip in). But for longer completion, it'll get tedious and we might start accepting wrong suggestions.
- zone411 4y agoIt doesn't work like this. A 1T model without architectural changes would not perform substantially better unless it has been trained on a lot more code. The original Codex was trained on 100B tokens, so you could possibly get some gains by increasing the model size but only up to a point. See the Chinchilla paper for reference.
- peaslock 4y agoNot necessarily: https://arxiv.org/abs/2206.14486 https://arxiv.org/abs/2206.14486 Also, even with "Chinchilla laws", you still gain performance in a larger model, you just need a lot more data (if just as noisy) to reach the same level of convergence, but a larger model will have already partially converged to a superior model with the same amount data.
- zone411 4y agoI've actually seen this paper before, but I don't think it's helpful. If the entire GitHub is 100B tokens and your prune it down properly, then fine, you can get equal performance with fewer tokes. However, if you want improved performance, you still need more data, not just a larger model size, and that's hard to obtain. I don't think it's a lost cause and we will be be stuck with current performance by any means though - there are other ways to go.
- peaslock 4y ago> if you want improved performance, you still need more data Not true. See figure 2: https://arxiv.org/pdf/2203.15556.pdf#page=5 https://arxiv.org/pdf/2203.15556.pdf#page=5 The loss decreases with greater model size at the same compute budget (i.e. stopping sooner regarding training data). Also some rehearsal/multi-epoch training improves the forgetting rate (thereby improving performance substantially), which hasn't been taken into account by Chinchilla et al. because they train <1 epoch. https://arxiv.org/abs/2205.12393 https://arxiv.org/abs/2205.12393
- zone411 4y ago