15 ms·
Diff Models – A New Way to Edit Code
- indeyets 4y agoRelated discussion: https://news.ycombinator.com/item?id=33271750 https://news.ycombinator.com/item?id=33271750
- return_to_monke 4y agowhile from the same company, same lab, I don't think this is what the article you linked is about. To me, that seems like a general purpose LLM and this just for code.
- indeyets 4y agoSo, it is loosely the same as copilot? I understand that approach is a tad different, but result of converting natural language descriptions into code-changes should be comparable. And both are trained on large corpus of github sources Is there a way to test it somehow? Public API maybe?
- Kiro 4y ago> converting natural language descriptions into code-changes Do people actually use Copilot for that? I just let it work its magic uninstructed. I guess it sometimes uses comments and function/variable names for its suggestions but that's about it. 99% of the time it just looks at my code, the context and neighboring files to predict what I'm trying to do.
- elcapitan 4y agoI use it most of the time as smart auto-completion as well, but sometimes for boilerplate it helps to just write a comment what you want to achieve, basically like a ChatGPT prompt.
- indeyets 4y agoI use both. Sometimes it feels easier to write five words of text than starting to write code.
- bil7 4y agofor my day job, no, not frequently. When I'm writing in an unfamiliar language like bash or something, I'll do a little # implement a function that does x, y and z
- noncovalence 4y agoI've found writing a temporary comment can be particularly useful when working with Unicode. For example, something similar to //insert a unicode dot between each character in the string, and convert the numbers to subscript saved me a lot of copy-pasting.
- pavlov 4y agoSomehow these GitHub-trained ML code assistants sadden me. My idea of enjoyable high-quality programming isn’t to dip a spoon into an ocean of soup made of other people’s random design decisions and bugs accumulated over fifteen years, hoping to get a spoonful without hidden crunchy insect bits. I know the soup is nutritious and healthy 98% of the time, and eating it saves so much time compared to preparing a filet mignon myself. But it’s still brown sludge.
- indeyets 4y agoWell, these are not tools for the art-level programming. But it helps to improve productivity of commercial programming a lot. Different genre
- carlbarrdahl 4y agoWhat if you could inspire the assistant with code you like and it would generate in that style? For example choose a few repos with code-bases you want to mimic, give it a set of instructions (and perhaps structure), and it generates code for it. Maybe something like GPT, style transfer, and OpenAPI combined.
- neximo64 4y agoAnd yet it is so useful. It is just an assistant. It's quite unlike soup, since you can easily alter it.
- indeyets 4y agoI think it might be compared of "cook it yourself" kits of ingredients. Good base but you can alter to your liking
- deleted 4y ago[deleted]
- fijiaarone 4y agoLike hamburger helper and box macaroni and cheese.
- RjQoLCOSwiIKfpm 4y agoPrepare for household appliances - washing machines etc. - doing strange things randomly. Prepare for the same thing with electronics which you didn't consider as containing much software before - central heating units, AC units, fridges, stoves, light switches, LED light bulbs, vacuum cleaners, electric shavers, electric toothbrushes, kids toys, microwave ovens, really anything which consumes electricity. Prepare for the support of the vendors of those appliances not taking phone calls anymore, only text communication. Prepare for the support not understanding the random problems you encounter. Prepare for the answers you get from support being similarly random. And maybe, with an unknown probability, prepare for your house burning down and nobody can tell you why.
- indeyets 4y agoYou imply, that such tools would lead to lower quality of code. I actually hope for the opposite. This is not a tool for generating applications using statistical methods (we have a lot of tools which do that already), but a tool for assisting human persons by taking boring/repetitive tasks from them and letting us focus on the meaning, the goal
- RjQoLCOSwiIKfpm 4y agoIf my house burns down due to random bugs in a big appliance, do you think the random underpaid 3rd world developers which will be used do care about that? I think this will lead to extreme cost cutting measures in choice of the developers which are used. People who would have previously been totally ineligible to develop software will happily be chosen. And they won't care about the garbage code they produce as long as it somehow seems to work from the outside. They'll care about feeding their families in the dire situation they are in, not more.
- indeyets 4y agoThere's a good chance this will produce BETTER code than "eligible" low-grade coders which do this now. You're overly optimistic about them
- startupsfail 4y agoFrom the safety perspective (may get important soon), it is perhaps a very bad idea to allow easy execution/injection of arbitrary code into random places with little review. One of the first steps of a misaligned/unhelpful/virus type of a system, attempting to secure its presence would likely be inference/GPU/TPU compute access. And code injection is a vector. There are multiple other vectors. When designing such systems, please do keep that in mind. Make sure code changes are properly signed and the originating models are traceable. Same applies to datasets generated by models.
- shireboy 4y agoIf this thing is trained on my commit messages we’re all doomed. Or else we’ll be able to type “fixed the thing” and have a whole app written.
- DominikPeters 4y agoIt would have been helpful to show some example generations of the model, unless I've missed them.
- bogwog 4y agoThis is all I could find: https://twitter.com/carperai/status/1619082410213404672 https://twitter.com/carperai/status/1619082410213404672
- jerpint 4y agoYes I agree. My understanding is you have in training dataset the original code, diff + commit message. So you train the LM to: Input: code+commit output: diff
- abdnafees 4y agoWhy now? I mean it's been only 20 odd years or so since modern programming became popular. And, it's not a lot. Let people learn how to code, make mistakes and then learn from those mistakes. Pre-cooked meals are not as good as home cooked goodness.
- parasti 4y agoI skimmed the post, but it seems not much was said about how the original diffs are generated. Git generates diffs only on request with varying levels of accuracy depending on the options given. Sometimes the diff completely fails to capture the intent of the change - it shows the path from A to B but not in any semantically meaningful way.
- moconnor 4y agoAll that to end with “no meaningful improvement over the salesforce codegen model” is a bit disappointing. Negative results are interesting in their own right. I’d rather read about why this isn’t better at the 6B parameter level than e see a hand wave that, well, the samples are more diverse and look the 350M model is better.
- youssefabdelm 4y agoYeah I felt the same way. Although perhaps at a higher scale the fine-tuning can make a bigger difference? The results go against this hypothesis but at least OpenAI states that GPT-3 only needs 200 examples, so who knows. In fact I wonder how well GPT-3 would do against this when fine-tuned on just 200 examples.
- spapas82 4y agoI'd really like to see how this would work with my commits... 99% of the messages on my commits are single word, similar to: - ok - fix - done - test - nice
- prettyStandard 4y agoGarbage in garbage out. You should fix that.
- alchemist1e9 4y agoYeah I bet the people working with them really love the commit messages /s Or more likely they are working alone.
- Taywee 4y agoI used to do that, until I had to go through my history to find a specific commit that I couldn't just diff for. It's like good comments in software. Half the time, you're doing it for your future self.
- manmal 4y agoAdditionally, people often write long PR descriptions while keeping commit messages to 80 chars. But when you think about it, PR descriptions are more or less ephemeral, while commit messages are persisted forever. There should be an emphasis on the latter.
- prettyStandard 4y agoThat's what my team and I do. We put effort into the PR, then squash it down. It works great.
- dizhn 4y agoThere's a character limit to commit messages in their training data.
- wslh 4y agoVery opportune. I am working on security diffs before and after security audit commits [1] reading the whole piece. [1] https://news.ycombinator.com/item?id=34360102 https://news.ycombinator.com/item?id=34360102
- lettergram 4y agoI view programming as a trade. I’ve spent years honing my skills, I pass wisdom to junior engineers as I can. I review code and provide detailed alternatives. My concern with AI across all fields are that people won’t gain the fundamental skills necessary for moving the bounds of what’s possible. Certainly, tools like this AI could produce good results. However, the underlying human is still providing the training data. More importantly, humans are producing the trajectory of development. If humans are no longer capable of pushing the AI systems. Then the AI systems will either cease to improve, or the AI systems will learn to play off each other. In highly complex systems like many programs, I suspect they’ll play off each other and achieve local minimum/maximum locations. Ie because the “game” (program development) can be iterative they’ll constantly improve code. However, because the AI systems don’t interact with all data (particularly real-world data) when a customer shows a sad face at some UI/UX, it won’t completely develop a new feature that matches the desires of the customer. Where I fear this will leave us is a class of less-skilled engineers and overly optimized AI. Basically, stuck in development.
- abhijeetpbodas 4y agoOn a philosophical level, AI for writing code has always seemed redundant to me. Here's why: 1. Humans create programming languages which machines can understand. OK. 2. Humans build tools (LSP, treesitter, tags, type checkers and others) to help humans understand code better. OK. 3. Humans build (AI) programs which run on machines so that the computer can understand... computer programs??? Aren't computers supposed to be able to understand code already? Wasn't the concept of "computer code" created so as to have something which the computer could understand? Isn't making a (AI) program to help the computer understand computer programs re-inventing the wheel? (Of course, I get that I use the terms "understand" and "computer programs" very loosely here!)
- semitones 4y agoThe benefit here is that the machine can execute what the AI produces, and humans can understand it / modify it if they need to.
- wankle 4y agoIt can be seen as a benefit or a cautionary tale. Earlier in the comments, someone claimed ChatGPT gave a recipe for a omelet made with 2 to 3 cow eggs. If instead of putting a recipe on the screen, the AI was connected to a cow and a frying pan...OW!
- semitones 4y agoI'm not making a judgment or claim as to whether the technology is beneficial overall. I was just explaining the benefit of the choice of output being source code rather than an executable.
- manmal 4y agoAs long as we don’t have „level 5“ code generation (no human oversight necessary), we need the code to be human readable. Afterwards, sure, why not produce assembly directly. Still it might be more practical to produce platform independent code instead - you‘ll only need to train one model instead of one per platform.
- jakear 4y agoExcellent. This is the beginning of the end for the cohort of people writing clear, descriptive commit messages. All your knowledge is soon to be acquisitioned and commodified by the Man with the GPU. I on the other hand will survive: what sense is an AI to make of such classic messages as David Bowie's excellent "ch-ch-changes!", the five "fix CI maybe???"s in a row, or the eternal "fuck this shit"?
- PoignardAzur 4y agoWe're still in the beginning for these tools, but already they're demonstrating some really exciting capacity. Something I haven't seen explored too much: navigation help. One of the things that takes me the most time when coding is remembering what was the next file / module / function I need to edit and jumping to it. An autocomplete engine that would suggest jump locations instead of token could help me stay in the flow much longer, with fewer worries about whether I'm introducing subtle bugs because I'm relying on the AI too much.
- Kwantuum 4y agoA lot of the comments seem to talk about the inevitable AI event horizon but unless I'm misreading this article the results are flat out bad. Even the 6 billion parameters model barely scratches a 50% success rate on a tiny problem that is trivial to fix for any human with basic knowledge of programming. Note the log scale of the graph.
- hellodanylo 4y agoYeah, I am also struggling to interpret the metrics in this post positively. The 50% success rate is also best out of 3200 completions. For best out of 1 completion, the success rate is in low single digits. I think the lesson here is that these models bring a lot more value when: 1. you have unit tests, 2. can afford compute/time to let the model try many solutions, 3. have enough isolation to run unverified code.
- kdnvk 4y ago6 billion is by no means large.
- zaidhaan 4y agoThey do note that the models "tend to do better when prompted with longer code generation tasks". But yes, the choice of scales for the graph was rather peculiar.
- tbrownaw 4y agoSounds like basically the inverse of what was on here the other day about automatically generating commit messages from a diff. Sounds kinda cool, even if trusting it would be a terrible idea.
- ilaksh 4y agoSince I am building a website https://aidev.codes https://aidev.codes to do programming based on natural language descriptions, this is extremely relevant to me. OpenAI has an 'edit' endpoint but it's 'in beta' and limited to 10-20 requests per minute. They do not acknowledge support requests about this. Azure OpenAI also has this endpoint I think but they ignore me as well. So for my edits just like everything else I have been relying on text-davinci-003 since it has much more feasible rate limits. I have just been having it output the full new file but maybe this Unified Diff thing is possible to leverage. Does anyone know, what would be the easiest way to try to run their 6B Diff Models thing against my own prompts for my service? Maybe Hugging Face?
- leo2023 4y agoThe next idea after this could be: developers draw a system diagram of the architecture, then AI writes the whole system E2E, high performance, distributed.
- Epa095 4y agoMaybe this can give a boost for languages like idris or F*, where you can specify much stronger types than in normal languages (with the price that you might have too proove the types manually). The types can help "tame" the AI generated code, and the AI can help generate the proofs. I also wonder if it could be useful in creating Coq proofs!
- ec109685 4y ago2022: engineers with 3 jobs 2023: engineers with their own AI model, typing “#fixed bugs” and spending the rest of the day by the pool.
- Jackson__ 4y agoI'm not sure if I'm just imagining it, but there seems to be a lot more negative push-back online to this than there was for copilot. It makes me wonder if it's related to recent protests in other creative fields in response to AI models, or just a weird dislike of openly released model weights?
- pklausler 4y agoHow good are these LLMs going to be at debugging code, as opposed to writing it?
- mortehu 4y agoI wrote the program between <BEF> and <MSG>, and it generated the following output: <NME> diff_model.py <BEF> import argparse import torch import transformers def main(): argparser = argparse.ArgumentParser() argparser.add_argument('--checkpoint', default='CarperAI/diff-codegen-2b-v2', choices=['CarperAI/diff-codegen-6b-v2', 'CarperAI/diff-codegen-2b-v2', 'CarperAI/diff-codegen-350m-v2'], help='Model to use') args = argparser.parse_args() model = transformers.AutoModelForCausalLM.from_pretrained(args.checkpoint) tokenizer = transformers.AutoTokenizer.from_pretrained(args.checkpoint) # Use CUDA. model = model.cuda() text = '<NME> diff_model.py\n<BEF> ' text += open("diff_model.py").read() text += '<MSG> Print a message after loading the model\n<DFF>\n' completion = model.generate(tokenizer.encode(text, return_tensors="pt").cuda(), max_new_tokens=400) completion = tokenizer.decode(completion[0], skip_special_tokens=True) print(completion) if __name__ == '__main__': main() <MSG> Print a message after loading the model <DFF> <DFF> @@ -1,4 +1,5 @@ ''' +import sys import argparse import torch @@ -8,6 +9,7 @@ def main(): tokenizer = transformers.AutoTokenizer.from_pretrained(args.checkpoint) # Use CUDA. + print("CUDA:", torch.cuda.is_available()) model = model.cuda() text = '<NME> diff_model.py\n<BEF> ' @@ -10,3 +12,4 @@ def main(): completion = tokenizer.decode(completion[0], skip_special_tokens=True) print(completion) + sys.exit(0) \ No newline at end of file [\ No newline at end of file repeated many times] It takes about 150 seconds to run on a 3090 Ti when the model is already on disk.
- shul 4y agoWhy all the hate? I for one welcome our AI overlords