34 ms·
Show HN: File-by-file AI-generated comments for your codebase
My friends and I were complaining about having to decipher incomprehensible code one day and decided to pass the code through GPT to see if it could write easily understandable comments to help us out. It turns out that GPT can but it was still a hassle to generate comments for large files.
So we decided to develop a basic web application that automatically integrates with your Github repository, generate comments, create a pull request and send you an email when it is all done.
There is definitely a lot more that can be done but we wanted to gain feedback on whether this is a problem that you face too. Do you often find it challenging to understand complex code? Do you have difficulties in writing informative comments? And if so, would you find value in a tool that can automatically generate comments for your code?
Really appreciate any feedback and suggestions! Thanks in advance!
- deleted 3y ago[deleted]
- gumballindie 3y agoCool idea, but for me the difficult part is understanding the author’s intention rather than the syntax. And thats is what needs solving. Not just at a high level but in detail.
- swiftstart 3y agoFair enough, that is a problem that we are facing too but we have no idea how to solve it at this point to be honest. Perhaps connecting one's github repository and commit history into a vectorised database for additional context? Hmmm
- gumballindie 3y agoMy advice for developers has been to write comments before writing the code. That way they describe the intention followed by an implementation. Even if there was a bug somewhere or code was poorly written at least i knew what the intention was. So maybe training a model against some really well documented codebases might help? Perhaps unit tests are a also a must since proper unit tests force granularity, thus in theory you’d have a good description of what the intention was + small chunks of code to correlate with. The issue tho is that most open source code is technical code and code that relies on it serves a business function. Either way i think the key is finding well documented well tdd’d business function code. Even if what you wish to explain is not clean and tidy but certain bits may fit in patterns that make sense when individually fit against a model. If it makes sense.
- swiftstart 3y ago"My advice for developers has been to write comments before writing the code." Haha yeah, however in our experience thus far, quite a number of developers really only explain their intentions verbally so nothing ever makes it into the codebase unfortunately. And after a few months, they themselves may not understand why the code was written in a particular manner. "The issue tho is that most open source code is technical code and code that relies on it serves a business function." Wow yeah never thought about it in that way before!
- michaelfeathers 3y agoIf they can be generated from the source they probably shouldn't be in the source. Maybe it should be an IDE plugin that displays comments for code as you hover over it.
- q7xvh97o2pDhNrh 3y agoAgreed this makes much more sense, from both a UX perspective as well as a code-cleanliness perspective. There's also the AI evolution question. When we have GPT-10 next year, do we go through and regenerate all the comments? That would introduce a lot of noise into the repo's commit history and `git blame`, which I think is another indicator that the repo is not the right place to store this sort of thing. (And it'd have to be done again every time the AI got smarter...) Having the AI perched on your shoulder and just analyzing the code as you look at it seems much simpler. Like a friendly, modern version of the pirate's parrot.
- TeMPOraL 3y ago> There's also the AI evolution question. When we have GPT-10 next year, do we go through and regenerate all the comments? That would introduce a lot of noise into the repo's commit history and `git blame`, which I think is another indicator that the repo is not the right place to store this sort of thing. Counterpoint: comments in code reflect what the writer thought at the time of their generation - whether they were a human or an AI (and, excepting automated stupidity, there would be a human reviewing and accepting AI-generated comments). "Having the AI perched on your shoulder" is like re-reading and re-interpreting what the code means. You get the benefit of experience (and improved AI models), but you'll also miss the context, long lost to time since the code in question was written. I'd say, we should to both. And code cleanliness... we won't make much more progress here than we've already made, not until we stop coding directly in the final, plaintext source form. There are too many conflicting concerns wrt. readability, and you can't have them individually optimized at the same time in a single piece of text.
- swiftstart 3y agoWow, thanks for the insightful discussion and feedback! This is definitely something that we will take into consideration and ideally, provide as an option.
- mensetmanusman 3y agoIt would be funny if comments are code in the future, as an AI just translates intentions to code that is invisible.
- rcme 3y agoThat's kind of my biggest issue with Co-pilot at the moment... it can turn comments into code, but the amount of effort required to produce the comments is greater than the effort required to write the code.
- swiftstart 3y agoIn that case, would our web application help to solve your problem? :) Specifically, you write the code, we write the comments for you.
- rcme 3y agoIt might be… I would say 95% of code doesn’t need comments. But then you encounter that one complicated function and you go wtf? So I think I would prefer something more targeted. Like a vs code extension that lets you summarize any code snippet.
- rossjudson 3y agoIt's both distressing and amusing to me that so many people think software engineers spend most of their time coding.
- swiftstart 3y agoHey, that is an interesting way to lens the possible future. There is definitely talk about the "singularity" of programming languages, similar to how you can translate from one language to another using the meaning behind the words. In fact, Copilot seems to be a glimpse into that future. Thank you for the comment!
- anotherpaulg 3y agoI really like the direct GitHub repo integration! I've thought about doing something similar as well. But keep in mind, this should be easy to do from the command line with a number of tools as long as you have a gpt-4 api key. I would probably trust gpt-3.5-turbo with this task in a pinch, but I think there would be more risk of it disrupting the original code. Here it is with aichat [1]: $ curl -s https://raw.githubusercontent.com/leachim6/hello-world/main/p/Python%203.py | aichat --model="gpt-4" -p "emit this exact code, but with helpful comments; don't put any comments before a #!shebang line if present" #!/usr/bin/env python3 # This is a simple Python script that prints "Hello World" to the console. # The first line, called the shebang, tells the operating system how to execute the script. # In this case, it specifies that the script should be run using the Python 3 interpreter. # The print function is used to output text to the console. # Here, it is used to print the string "Hello World". print("Hello World") Or with my own tool aider [2]: $ git clone https://github.com/leachim6/hello-world.git $ cd hello-world $ aider "p/Python 3.py" Added p/Python 3.py to the chat Using git repo: .git > add helpful comments p/Python 3.py <<<<<<< ORIGINAL #!/usr/bin/env python3 print("Hello World") ======= #!/usr/bin/env python3 # This is a simple Python script that prints "Hello World" to the console print("Hello World") # Print "Hello World" to the console >>>>>>> UPDATED Applied edit to p/Python 3.py Commit aad4afc aider: Added helpful comments to Python script. [1] https://github.com/sigoden/aichat https://github.com/sigoden/aichat [2] https://github.com/paul-gauthier/aider https://github.com/paul-gauthier/aider
- swiftstart 3y agoInteresting! We are personally not the most comfortable with editing things directly from the terminal, especially when GPT hallucinates, but we can definitely see how this would provide users with more flexibility. Thanks for sharing!
- anotherpaulg 3y agoI agree, sometimes you need to carefully review the changes that GPT suggests. My aider tool tries to make this easy by leveraging git. While it automatically commits the edits from GPT, it also provides in-chat commands like /diff and /undo. These commands let you quickly check exactly what edits GPT made, and undo them if they're not correct. Aider will notify GPT if you /undo its changes, and GPT will probably ask why and then try again with your concerns in mind. To manage a longer chat that includes a sequence of changes, you can also use your preferred standard git workflows like branches, PRs, etc.
- phailhaus 3y agoThis is really cool, but somewhat misses the point of comments. If the comments can be generated from the code, then they are just restating what the code is doing. That's great for docstrings! But code comments are far more useful when they explain to you why the code is a certain way. That is something that is learned by the person implementing the code, and can't be learned by looking at it directly. For example, "Uses a generator here to avoid fetching all results into memory".
- swiftstart 3y agoThanks for the feedback! We can't figure out how to hack together something like that just yet (and if it is even something that should / can be solved by a tech product) but if we do, we'll definitely share that as an update! :)
- crazygringo 3y agoI love the idea of using this to more quickly understand otherwise badly-commented or uncommented code. I'd hate the idea of this leading anyone to avoid commenting their code ("just let the AI do it!"), since comments also need to be about the why and how to use and how not to use -- not just the what.
- swyx 3y agolanding page should show example input output. youre asking for too much up front. all the best!
- swiftstart 3y agoThanks for the feedback! The web application has been updated with a link to some sample files accordingly!
- lambdaba 3y agoTrying to get GPT to generate comments at a particular level really highlights its limitations in my experience. For instance, I couldn't get it to focus on commenting on programming language aspects of the code (or only in a crude way). There's some depth it's lacking, it might be from RHLF, I don't know, but its commenting is like its writing.
- i2cmaster 3y agoHave you tried getting it to write a high level description before reproducing the code with comments? (via either FSL or instructions) Most of the reasoning ability in LLMs comes from them rambling about something and then the attention picking up on the rambling when it needs to generate the conclusion. If you skip that then the output will probably be much less coherent.
- swiftstart 3y agoWe played around with this for a bit actually. One idea we had was to generate a PlantUML diagram to show how the different components of a file or even a repository connected with one another. However, given the current limitations with GPT context, even when using GPT-4, this quickly became impractical for large files. We would need to leverage an AI with a much larger context length. That said, perhaps if the entire repository is fed into a vectorised database, a high-level overview would be possible? Just thinking aloud right now and am happy to collaborate with anyone interested in exploring this further!
- swiftstart 3y agoYeah, we had to do a lot of prompt engineering and even then, we still had to clean up the files quite a bit programmatically. Perhaps GPT-5 and beyond would make this entire process 10X easier :)
- kmoser 3y agoI entered my email address, selected a file, clicked "Upload", and a popup told me "Please select a file." Nothing I did would get it to ingest a Python file. Tried with both Firefox and Chrome with the same results. Sample Python file was only 29 lines long. What am I missing?
- swiftstart 3y agoHey, could you reach out to us at swiftstarters@gmail.com? Will be happy to help you debug this. Thanks!
- Nekobai 3y agoYep, the tool only seems to be accepting .zip files. I managed to upload a file successfully by zipping a .js file. The file upload successfully, but then when I received the results a few minutes later, the zip file containing the results was empty.
- swiftstart 3y agoHey, yeah this is an issue that we are aware of and has since been rectified. If you are still experiencing this issue, please feel free to reach out to us at swiftstarters@gmail.com Thanks!
- al_be_back 3y ago>> My friends and I were complaining about having to decipher incomprehensible code one day Your raw sample files [1] are already highly descriptive: filenames, component names, variables etc etc. the auto-generated comments add only noise (increase file size). I would use complex, hard-to-decipher incomprehensible code to show how these comments bring sense forth. Perhaps let it loose on the Linux Kernel source, see what it does - an idea Personally, I prefer the doc-as-you-code approach, especially using context-aware naming. When code's cryptic & large, i try to visualise it e.g. Sequence Diagram from code (intellij). [1] https://swiftstart.notion.site/Sample-files-acfb097bbf214210b5a21df46eaa5819 https://swiftstart.notion.site/Sample-files-acfb097bbf214210...
- davidguetta 3y ago100% I've had this problem in the past and the functinos were way more obscure than something called " get_doc_from_code ". Can it even deobfscuate code ?
- swiftstart 3y agoThanks for the feedback! That's a really good idea and we will definitely give that a shot! If it provides more clarity, we are trying to tackle this in stages. Can the AI produce helpful and descriptive code for 1. Basic + Short (less than 100 lines) files? 2. Basic + Long (more than 100 lines) files? 3. Complex / Vague + Long files? After stage 3, which I believe is what you are referring to, we then hope to explore stage 4. 4. Can AI incorporate programmers' intentions into the comments?