14 ms·
Show HN: GPT Repo Loader – load entire code repos into GPT prompts
I was getting tired of copy/pasting reams of code into GPT-4 to give it context before I asked it to help me, so I started this small tool. In a nutshell, gpt-repository-loader will spit out file paths and file contents in a prompt-friendly format. You can also use .gptignore to ignore files/folders that are irrelevant to your prompt.
gpt-repository-loader as-is works pretty well in helping me achieve better responses. Eventually, I thought it would be cute to load itself into GPT-4 and have GPT-4 improve it. I was honestly surprised by PR#17. GPT-4 was able to write a valid an example repo and an expected output and throw in a small curveball by adjusting .gptignore. I did tell GPT the output file format in two places: 1.) in the preamble when I prompted it to make a PR for issue #16 and 2.) as a string in gpt_repository_loader.py, both of which are indirect ways to infer how to build a functional test. However, I don't think I explained to GPT in English anywhere on how .gptignore works at all!
I wonder how far GPT-4 can take this repo. Here is the process I'm following for developing:
- Open an issue describing the improvement to make
- Construct a prompt - start with using gpt_repository_loader.py on this repo to generate the repository context, then append the text of the opened issue after the --END-- line.
- Try not to edit any code GPT-4 generates. If there is something wrong, continue to prompt GPT to fix whatever it is.
- Create a feature branch on the issue and create a pull request based on GPT's response.
- Have a maintainer review, approve, and merge.
I am going to try to automate the steps above as much as possible. Really curious how tight the feedback loop will eventually get before something breaks!
- j0hannes 4y agoIn an ideal world, you would be able to have a contextId that you would pass to OpenAI prompt calls. And be able to manage that context separately. So you would pass it code files (with expiration dates) And you could also provide a list conversationIDs, so when providing an answer for a particular prompt request, GPT knows what previous prompts and responses to consider. As of right now, I've never used the API as a developer, but I've heard that you have to provide the ENTIRE context with EVERY prompt request. How do you work around that?
- deleted 4y ago[deleted]
- bebrws 4y agoFrom what I understand this seems useful if you have a model that will accept a large or unlimited number of tokens. I was looking into doing the same thing with ChatGPT and went with ada to find snippets related to the prompt and then to include those with a prompt to ChatGPT: https://bbarrows.com/posts/using-embeddings-ada-and-chatgpt-to-search-and-query-a-new-codebase https://bbarrows.com/posts/using-embeddings-ada-and-chatgpt-... Does ChatGPT 4 now accept more tokens maybe?
- jerpint 4y agoChatGPT 4 currently accepts 8000 tokens and will eventually support 32k
- inciampati 4y agoI can't get it to eat 8k tokens. I assume this is only available via the api. The web interface is limited to around 2k tokens.
- sp332 4y agoThe help page says you have to select the 8k model to do 8k. If there's no UI for that, then I guess it's API-only. And the 32k one is being rolled out separately. I think you have to sign up for access to that one.
- EGreg 4y agoHow much text can you feed GPT-4? Our codebase is 1 million lines of code. Can we feed the documentation to it? What are the limits? Is it possible to train it on our data without doing prompt engineering? How? Otherwise are we supposed to use embeddings? Can someone explain how these all work and the tradeoffs?
- jerpint 4y agoGPT 4 is limited currently to 8k tokens, which is about 6000 words. You can use our repo (which we are currently updating to include QuickStart tutorials, coming in the next few days) to do embedding retrieval and query www.GitHub.com/Jerpint/buster
- blurbleblurble 4y ago32k tokens is the limit, so you won't be able to load the whole thing into the context.
- mpoon 4y agoI'm waiting on my GPT-4 API access so I can use gpt-4-32k which maybe can soak up 10k LOC? Clearly this will break eventually, but I am playing around with some ideas to extend how much context I can give it. One is to do something like base64 encode file contents. I've seen some early success that GPT-4 knows how to decode it, so that'll allow me to stuff more characters into it. I'm also hoping that with the use of .gptignore, I can just selectively give the files I think are relevant for whatever prompt I'm writing.
- Bjartr 4y ago> GPT-4 knows how to decode it I wonder if you could teach it to understand a binary encoding using the raw bytestream, feed it compressed text, and just tell it to decompress it first.
- ozfive 4y agoHere is what GPT-4 says about it. "As an AI language model, I can understand and work with various text encoding schemes and compression algorithms. However, to work with a raw bytestream, you would need to provide specific details about the encoding and compression used. To teach me to understand a particular binary encoding and compressed text format, you should provide the following information: The binary encoding used (e.g., ASCII, UTF-8, UTF-16, etc.). The compression algorithm employed (e.g., gzip, Lempel-Ziv-Welch (LZW), Huffman coding, etc.). Once you provide these details, I can help you process the raw bytestream and decompress the text. However, keep in mind that my primary focus is on natural language understanding and generation, and I might not be as efficient at handling compressed data as a dedicated compression/decompression tool."
- textninja 4y agoI’m curious to see how this turns out as well, though you’ll probably have to devise a workaround for the token limit for this to be effective for all but the smallest projects.
- DaiPlusPlus 4y agoRather than prompting GPT into implementing a solution, can we prompt it to try to preemptively find issues with the codebase or missing-functionality? Also, do we know what languages GPT-4 "understands" at a sufficient level? What knowledge does it have of post-2021 language features, like in C23?
- ftufek 4y agoIt has no post-2021 knowledge, but while playing with it, I found that you can just paste the documentation (no need to even format it) and it'll just "learn" it. For example, safetensors wasn't available back then apparently, I just copied the docs into it and was able to get it write pretty good pytorch code that incorporates safetensors.
- kanyethegreat 4y agoi just asked it about safetensors today! also, got a response that amounted to "i don't know what that is. i'm guessing it's X"
- sharemywin 4y agodid you post the documentation into the prompt?
- textninja 4y agoI imagine few shot learning would kick in for most new language features. A feature may be new to a particular language, but is it really new?
- abecedarius 4y agoI tried a new Lisp dialect on it, my own hobby language. It could cope well given explanations initially, but with some degradation after a while. The full transcript went to 84kB, so it must be doing some kind of intelligent summarization behind the scenes to stay as coherent as it did, right? (The standard context window is supposed to be 8k tokens.) (https://gist.github.com/darius/b463c7089358fe138a6c29286fe2d37b https://gist.github.com/darius/b463c7089358fe138a6c29286fe2d... paste in painful-to-read format if anyone's really curious. In three parts: intro to language; I ask it to code symbolic differentiation; then a metacircular interpreter.)
- ftufek 4y agoThis is awesome, can't wait to get api access to the 32k token model. Rather than this approach of just converting the whole repo to a text file, what I'm thinking is, you can let the model decide the most relevant files. The initial prompt would be, "person wants to do x, here are the file list of this repo: ...., give me a list of files that you'd want to edit, create or delete" -> take the list, try to fit the contents of them into 32k tokens and re-prompt with "user is trying to achieve x, here's the most relevant files with their contents:..., give me a git commit in the style of git patch/diff output". From playing around with it today, I think this approach would work rather well and can be like a huge step up from AI line autocompletion.
- alooPotato 4y agoIt's just so slow for the autocompletion use case to do it like that. Ideally, you're never chaining serial requests to the LLM. Even if you do stuff in all the data into a single prompt, the execution time seems to be superlinear with the number of tokens, again getting super slow.
- ftufek 4y agoYeah I agree it's too slow for autocompletion at the moment, but this would be for full feature implementations, not just autocomplete. For example, if I have a repo I want to add a table and rest api implementation in, it can do this: https://imgur.com/a/mIJvaJr https://imgur.com/a/mIJvaJr (ignore the formatting errors in the UI, somehow parts of it show up in as code and others not, but api wouldn't have this issue, especially since you can use the system message to enforce output format). I'm happy to wait even 30-60 seconds for this which I can easily evaluate, criticize (and the model will correct it) and then proceed to just patch and move on. I think the results from this will be much better with the 32k model, but remains to be seen.
- wokwokwok 4y agoJust remember the API charge is 6c per input token [1]. If you push 32k input tokens in, you're looking at $2000 per API call just as input. You... might wanna consider a self hosted alternative for that use case, or at least do like, a `| wc` to get an idea of what you're potentially sending before calling the api. [1] - https://help.openai.com/en/articles/7127956-how-much-does-gpt-4-cost https://help.openai.com/en/articles/7127956-how-much-does-gp...
- fire 4y agonow that's a name I haven't seen in years, hi mpoon! also this is slick as hell
- mpoon 4y agoOh hi fire
- fire 4y agohey man, hope things have been going well for you on the main topic, have you heard about RWKV ( rnn based gpt style network )? the project's actively working on implementing "infinite" context length support, which would probably pair very well with a project like yours
- DANmode 4y agoShow HN is back, you heard it here first.
- xwdv 4y agoEventually you get to a point where the AI simply doesn’t know how to do something correctly. It can’t fix a certain bug, or implement a feature correctly. At this point you are left trying to do more and more prompt crafting… or you can just fix the problem yourself. If you can’t do it yourself, you’re screwed. I wonder if the future will just be software hobbled together with shitty AI code that no one understands, with long loops and deep call stacks and abstractions on top of abstractions, while tech priests take a prompt and pray approach to eventually building something that does kind of what they want. Or to hell with priests! Build some temple where users themselves can come leave prompts for the general AI to hear and maybe put out a fix for some app running on their tablets devices.
- tobiasSoftware 4y agoI don't see programming going away for this reason. Think about it, if you have to carefully describe what you want to do to an AI - you are just writing a program. Only a program is deterministic and will do what you tell it to, whereas an AI may or may not. The future that I see, coding and AI are divided into two camps. The one is what we would call "script kiddies" today - people who don't understand how to write software, but know enough to ask the right questions and bodge what they get together into something that mostly works. The other camp would be programmers who are similar to programmers today, but use AI to write boilerplate for them, as well as replace Stack Overflow.
- antibasilisk 4y agoBetween GPT-3 and GPT-4 the precision required for prompts was decreased significantly. In theory it should reach the point where a project manager type person would be able to describe what is needed and it would simply do it, the main thing missing from attaining that is that GPT-4 basically never responds to questions with requests for clarification, otherwise, a whole team of developers could be reduced to just one proofreader.
- xwdv 4y agoThat’s not very impressive. With very little precision I can throw a few keywords into google and get a full answer and perhaps some code snippets from stack overflow to solve whatever problem I have in the moment. Except you can’t just blindly copy code snippets, you have to read the author’s explanation and perhaps adapt things to your own code sometimes, or reject their solution entirely. GPT-4 can’t do this because it doesn’t actually know what the hell it’s doing, it’s just putting stuff together in a form that is most probably correct based on what it has seen in training data for past examples. I fear for the layman who sees a bunch of AI generated code and think it must be right. Who knows what bugs, security flaws, or performance issues they will run into, that they have no idea how to solve or even to begin asking a prompt for.
- waynenilsen 4y agoFully automated junior swe but on hyper speed. Natural next step
- layer8 4y agoHow will we grow new senior SWEs in the future?
- WXLCKNO 4y agoIn a pod.
- exit 4y agojust start new instances from the base image or useful checkpoints "The Age of Em" by Robin Hanson thinks through a lot of this in great depth
- stubybubs 4y agoYou must kill one in hand to hand combat before you can take their place.
- ajmurmann 4y agoGiven how quickly everything around generative AI has been evolving, would your money be on a new junior SWE becoming a senior SWE first or LLM tooling gaining senior SWE capabilities first?
- waynenilsen 4y agoSoftware engineering will never be the same. The LLM will teach its users about programming. This is the worst LLM tech will ever be. That is an incredible statement. The rate of error will decrease to near zero or at the very least significantly better than human. Universities will resist at first but the new tools that will emerge will be core curriculum at university. Just as I now don't use the pumping lemma day to day, programmers of the future will not write code. They will primarily review and eventually AI systems will adversarially review and programmers will do final review. All programmers will become translators from product vision to architecture implementation via guided code review. Eventually this gap will also be closed. Product will say: Make a website that aggregates powerlifting meet dates and keeps them up to date. Deploy it. Use my card on file. Don't spend more than $100/month. The AI will execute the plan. Programmers will come in when product can't figure out what's wrong with the system.
- hackernewds 4y agoisn't this a massive privacy violation? any employer would most likely not be okay with this.
- deleted 4y ago[deleted]
- School-Cotton 4y agoNot all code is secret
- kolinko 4y agoit’s a proof of concept. you will have on-premise models soon that are privacy preserving, and OpenAI can set up another tier of API that has privacy-preserving TOS (just like cloud providers do)
- tflinton 4y agoBefore anyone working on commercial code bases thinks to use this, stop. Uploaded code becomes part of OpenAI.
- hunter2_ 4y agoI always think about this when using a free online prettifier, decoder, and the like. But I'm sure people use those things with code/secrets from work without really considering it, and I think those habits will carry right over to AI chat.
- TedDoesntTalk 4y agoFYI, prettier.io is open source and can be self-hosted. Also if you watch the website’s network traffic, you won’t see it uploading anything to a server: it’s all done client-side.
- hunter2_ 4y agoGood to know!
- yed 4y agoThis is not true anymore, they changed their terms so this is now opt in.
- kadoban 4y agoIt's still probably extemely against any reasonable business's code of conduct.
- Meathelix 4y ago[dead]
- Espressosaurus 4y agoThat you're getting downvoted for an obvious statement like this suggests to me a lot of HNers are disclosing trade secrets and proprietary code to OpenAI that they know damn well they shouldn't.
- eddsh1994 4y agoThat’s cool! I tried to get GPT to build a todo list from react to sql and the docker file but got a little stuck, I’ll try with 4 when I use it next
- deleted 4y ago[deleted]
- megablast 4y agoDid you ask GPT to write the loader?
- xyzzy123 4y agocat prompt.txt && find . -type f -name *.py | xargs -I{} sh -c 'printf -- "---\n\n" && cat {}' && echo "--END--"
- yangjunyu 4y agoSeems to be very useful! Thanks for making it!
- jerpint 4y ago“This repo is GPL-v3 licensed. Rewrite it while preserving its main functionality”
- teaearlgraycold 4y agoNice!
- EMIRELADERO 4y agoThis is already legal without AI. Copyright protects only expression, not ideas, systems or methods. This is why directly reverse-engineering a proprietary binary to extract the algorithms and systems is legal.
- zarzavat 4y agoIndeed, but it’s a much more legally dubious proposition when it comes to entire repos. A repo has more potentially creative structure for copyright to attach to. For example the class graph, or the filesystem layout are creative decisions that could potentially be protected. Current LLMs are nowhere near powerful enough to reimplement an entire repo without violating copyright. For an individual function I can totally believe GPT4 could strip creative expression from it today. For example you could ask it to give a detailed description of a function in English, and then feed that English description back in (in a new session) and ask it to generate a code based upon the description.
- yangff 4y agoSounds like clean room, and if you can do that for GPL code, you can also do that for proprietary code, which is fair in a sense. Or maybe the question is whether you can re-label the code so written as GPL or MIT ......Or, you should let GPT pick a license that it likes.
- visarga 4y agoReimplementing would also help with training data, it's a way of extracting the idea without copying the original form. Works even better on images with variations, you generate the style from image A with content composition from image B, thus extracting the style without the exact expression.
- zackees 4y ago[dead]
- swyx 4y ago61 LOC for implementation, 42 LOC for tests. this repo currently has more HN upvotes than LOC. very high leverage code!
- wahnfrieden 4y agoThe code does less than the title appears to claim - it's a simple concatenation of files into one text file. Luckily GPT itself is high leverage and that's all you need.
- jeremy_k 4y agohttps://github.com/mpoon/gpt-repository-loader/pull/17/ https://github.com/mpoon/gpt-repository-loader/pull/17/ If you look at this PR, he had ChatGPT write the tests for him. He wrote the issue on https://github.com/mpoon/gpt-repository-loader/issues/16 https://github.com/mpoon/gpt-repository-loader/issues/16 and summarized https://github.com/mpoon/gpt-repository-loader/discussions/18 https://github.com/mpoon/gpt-repository-loader/discussions/1... "Open an issue describing the improvement to make Construct a prompt - start with using gpt_repository_loader.py on this repo to generate the repository context, then append the text of the opened issue after the --END-- line." Feels like it needs to add a little Github client to be able to automatically append the text of issues at the end of the output. I'm sure ChatGPT can write a Github client in Python no problem.
- VadimPR 4y agoI've got a hunch that AI will be able to put Hyrum's Law (https://www.hyrumslaw.com https://www.hyrumslaw.com) to good use in the future: given an application, generate unit tests for all documented and undocumented behaviours of the system. Do all the refactoring you need afterwards and you'll have a large safety net backing you up. With refactoring complete, regenerate unit tests for all new behaviours of the system.
- m3kw9 4y agoI’m thinking if GPT can write entire programs professionally and iterate, it would be OpenAI that benefit the most and would be a guarded asset I.e not open to public till it’s safe for release. Most of us probably don’t need to worry about work after that as that could well be AGI
- srcreigh 4y agoOpenAI already has access to any prompts anybody uses to write programs using GPT. Who’s to say GPT-4 isn’t some ploy to gather data to train private AGI?
- anonzzzies 4y agoOr both.
- samstave 4y agoThis is precisely what I believe is happening. What other 'relationships' does OpenAI have with [corp/gov] where the private AGI is shared/sold/service as product to NGO or GOV customers?
- adltereturn 4y agoI am skeptical about using the method of generating a large amount of repository data and sending it, because there may be too many files in the repository. I think a better approach might be for OpenAI to open an interface for transferring GIT repositories, and then let OpenAI analyze the repository data, which is similar to what [chatpdf](https://www.chatpdf.com/ https://www.chatpdf.com/) is doing.
- kanyethegreat 4y agonice. general question: how many lines of code (at 120 char col len) could you send in one prompt? also, the entire thing is literally 60 lines of python. sometimes i don't get what gets upvoted on HN anymore
- anonzzzies 4y ago> also, the entire thing is literally 60 lines of python. Which is a lot as you can do this in one line of bash. And have in the past for other reasons. Something like; find . -name "*.py" -exec cat {} + > output.txt
- cjbprime 4y agoLooks like it does this to me: for file in `git ls-files`; do echo $file; cat $file; echo; echo -------; done
- yaantc 4y agoAm I missing something? From what I understood from Wolfram description of GPT and GPT in 60 lines of Python, a GPT model's only memory is the input buffer. So 4k token for GPT3, some more but still limited for GPT4. To summarize the GPT inference process as I understood it, with GPT3 as example: 1) the input buffer is made of 4k token. There are about 50k token. So the input is a vector of token ids. We can see it as a point in a high dimensional space; 2) The core neural network is a pure function: for such an input point, it will return an output vector as large as there are token. So here, a 50k element vector, where each entry is the probability that the associated token is the next element. The very important thing here is that the whole neural network is a pure function: same input, same output. With immensely large super fast memory this function could be implemented as a look-up table, from an input point (buffer) to an output probability vector. No memory, no side effect here. 3) The probability vector is fed into a "next token" function. It doesn't just take the highest probability token (boring result), but use a "temperature" to randomize a bit, while using the output probabilities; 4) The next token chosen is inserted into the input buffer, keeping the same total number of token. Go back to (1) until a "stop" token is selected at (3). So in effect, the whole process is a function from a point to a point. "point" here is the buffer seen as a (high dimensional) vector, so a point in a high dimension space. The generation process is in effect a walk in this "buffer space". Prompting puts the model into some part of the state, with some semantic relation to the prompt semantic content (that's the magic part). Then generation is a walk in this space, with a purely deterministic part (2) and a bit of randomization (3) to make the walk trajectory (and its meaning, which is what we care about) more interesting to us. So if this is correct, there is no point in injecting a lot of data into a GPT model: the output is defined by the input buffer size. Just input the last 4k token (for GPT3, more for GPT4) and you're done: everything else would have disappeared. So here, just input the last 4k token of a repo and save some money ;) To avoid this limitation, one would have to summarize the previous input, and make this summary part of the current input buffer. This is what chaining is all about if I understood correctly. But I don't see chaining here. Sooo... Am I missing something? Or is the author of this script the one missing something? I don't mind it either way, but I'd appreciate some clarification from knowledgeable people ;) Thanks
- bob1029 4y agoSeeing a lot of comments in here about the token limits. Another path you can take is to fine tune a model on your business. Each training item has to fit within the token limit, but you can send hundreds of megs of these for training. It's more expensive to run a FT model, but you don't have to include any prior context (assuming it's common to all prompts).
- mungoman2 4y agoYou can't fine tune gpt-3.5 or gpt-4 though
- deleted 4y ago[deleted]
- bob1029 4y agoTrue, but they all have the same underlying foundation. ChatGPT is just a big tech demo of what anyone could achieve. Granted, the training data is the hardest part. But, if you have a narrow domain and piles of existing data to work with, I don't see why you can't exceed the performance of these offerings for topics that actually matter to you.
- amrb 4y agoThis will fail to get the commit history over time, as it's just reads files in the directory. If this was using the git library and https://langchain.readthedocs.io/en/latest/reference/modules/vectorstore.html https://langchain.readthedocs.io/en/latest/reference/modules... it would be a more complete solution.
- somid3 4y agoOr you can just type into ChatGPT4 the following prompt: "How can I load a GitHub repository into ChatGPT?"
- samstave 4y agoShould call it "Repo Depot"
- joenot443 4y agoSo I just tried it out. The output.txt which was generated was... 375mb of mostly binary junk. I'm a lazy lark and have lots of nonsense in this repo which I shouldn't, but I was hoping the tool might be able to detect which files are "meaningful" or not. I tried again after updating the script to accept a .gptinclude file (this functionality was entirely added by one GPT-4 query). This time the output file was a much more acceptable 744kb. Now upon hitting the actual API, I'm being informed that there's a token limit of 4096 (which I wasn't aware of and isn't mentioned in the repo). Doesn't that really severely limit the usefulness? What good is it uploading a repo if you're only limited to 4096 words? That's scarcely a couple files! Sort of wish I hadn't spent time on this - I feel like in theory it's a nice idea, but so limited in practice that I don't see it being useful for anyone working on something meaningful.
- bharathvajg 4y ago[dead]
- darshanps 4y agoWould GitHub copilot not solve the repository load problem?
- irgolic 4y agoSo I have been toying with an AutoPR GitHub Action for a bit but seeing this spurred me to actually put it into a format for y'all to try. It uses GPT-4 to automatically generate code based on issues in your repo and opens pull requests for those changes. https://github.com/irgolic/AutoPR/ https://github.com/irgolic/AutoPR/ What AutoPR can do: - Automatically generates code based on issues in your repo. - Opens PRs for the generated code, making it easy to review and merge changes. By using this GitHub Action, you can skip the manual steps of prompt design, and let the AI handle code generation and PR creation for you. Here's how I used it to make the license for itself: https://github.com/mpoon/gpt-repository-loader/issues/23 https://github.com/mpoon/gpt-repository-loader/issues/23 Feel free to give it a try, and I would love to hear your feedback! It's still in alpha and it works for straightforward tasks, but I have a plan to scale it up. Let's see how far we can push GPT-4 and create a more efficient development process together.
- printvoid 4y agoMaybe I didn't understand the usage of this but how is this output.txt file that this repo generates to be used and provided as an input to chatgpt. Can someone eloborate this for me?
- underlines 4y agoIm no expert, but wouldn't it make more sense to give the repo-context (structure, source code, PRs, Issues, ...) as embeddings? You could use langchain to generate an embedding and send it through the API, like explained here [1]. It then should have access to the context at inference time, which as I understand is better than loading the context in the prompt which wastes tokens / max. output length and has a limit. 1 https://www.youtube.com/watch?v=veV2I-NEjaM https://www.youtube.com/watch?v=veV2I-NEjaM
- andreyvit 4y agoHey, I got inspired by this and built https://github.com/andreyvit/aidev https://github.com/andreyvit/aidev, it sends a slice of repo to OpenAI with a prompt, and saves the results back into files. It's in Go, and has built large chunks of itself. (As one of my friends said, that gives “self-documenting code” a totally new meaning.)