11 ms·
Show HN: Promptr, let GPT operate on your codebase and other useful goodies
Hi HN,
I've been working on an experimental tool that helps you use GPT to work on your codebase. I'd love to improve the tool if there's interest. New ideas welcome! I think this could also be useful for experimenting with other types of recursive prompts.
It’s a little bit Swiss Army knife and a little bit skynet:
https://github.com/ferrislucas/promptr https://github.com/ferrislucas/promptr
From the README:
Promptr is a CLI tool for operating on your codebase using GPT. Promptr dynamically includes one or more files into your GPT prompts, and it can optionally parse and apply the changes that GPT suggests to your codebase. Several prompt templates are included for various purposes, and users can create their own templates.
- ilovecaching 4y agoJust want to say that I'm not going to touch any AI tools until the legalities are ironed out, and I absolutely won't be using it at work without very clear approval. AI poses a huge risk to businesses as programmers start feeding code into other companies backends without thinking and pulling out random snippets from projects with varying licenses.
- qikInNdOutReply 4y agoSo AI products can only be developed in AI friendly legal areas? Similar to cloning and stem cell research?
- w0m 4y agos/developed/used/ I think is more the comment. OrgA will be 'all in' with the convenience of Copilot; OrgB will flee due to legality/litigious concerns.
- ilovecaching 4y agoI never said anything about AI development.
- Cloudef 4y agoIgnoring AI tools is very shortsighted IMO
- jay-barronville 4y agoBeyond shortsighted.
- roflyear 4y agoWhat's beyond shortsighted? What would you call that?
- ilovecaching 4y agoBeing cautious in order to keep one's IP and job safe during an economic downturn isn't shortsighted, it's playing the long game. AI isn't going anywhere, and waiting and seeing how the legalities shake out and how companies will want to consume this internally is the smart play. Jumping onto a new technology with unknown risks and getting burned is shortsighted.
- JonAtkinson 4y agoThe "long game" becomes short if your organisation is out-executed by a competitor who isn't as conservative in approach.
- namaria 4y agoThen I get home and think: how come we're collectively killing the biome that sustains us and degrading its ability to nurture human civilization? Then I remember I need to race to the bottom with everyone else to get that cash son!
- JonAtkinson 4y agoIt's the reality of the marketplace. I'd like to change it too.
- reisender 4y agoYou are spot on with the work comment. Posting IP you do not own to a service you may not be permitted to use is a bad idea.
- swader999 4y agoIf you've reduced your code base to vector weights and that's all you send to an api, I wonder how much you'd mitigate this concern?
- asteroidz 4y agoWise decision as far as I'm concerned. Unfettered AI experiments on production systems are an IP and (depending on your field) regulatory nightmare waiting to happen. Unless you have lawyer-vetted guidelines and oversight, use OpenAI products exclusively for experiments and personal projects.
- yosito 4y agoThe README and code examples are quite dense. To be honest, I'd rather think about my own code than try to figure out how to write Promptr commands to get GPT to do it for me. Is there any way you can simplify the syntax, or at least create some short aliases for common use cases?
- deathmonger5000 4y agoYes, I totally agree. I just wanted to get it out there. Probably a little early, but oh well.
- oslac 4y agoIt sounds like a good idea to send your codebase through a MITM to OpenAI, both of these ideas I mean.
- deathmonger5000 4y agoThe files you pass to promptr are indeed sent to OpenAI. There’s no man in the middle. Privacy is important, and I’m glad you care. The relevant code is here https://github.com/ferrislucas/promptr/blob/3ae09d1cffbb6b931399c9e7ef571317a5c274e3/gpt4Service.js#L17 https://github.com/ferrislucas/promptr/blob/3ae09d1cffbb6b93... And here: https://github.com/ferrislucas/promptr/blob/3ae09d1cffbb6b931399c9e7ef571317a5c274e3/gpt3Service.js#L24 https://github.com/ferrislucas/promptr/blob/3ae09d1cffbb6b93...
- deleted 4y ago[deleted]
- debarshri 4y agoI think, if I am not wrong, commentator is saying that you are the man in the middle.
- WrtCdEvrydy 4y agoThis is good for security overall... think about it, if we sent all of our stuff to the NSA, they would find security bugs and fix them for us.
- Frummy 4y agoI think in 5 years maybe chatGPT can revolutionize the handling of technical debt, for example with the multidisciplinary knowledge within legal and finance on top of code and architecture rewrite all of the ancient z/os systems in the financial sector into something fitting the modern age.
- rashkov 4y agoAlso it could be used to save a lot of abandonware through AI enhanced reverse engineering. This is the kind of high toil work that very few people want to do, unless it becomes a lot easier
- groby_b 4y agoI promise you that all financial and legal systems worth rewriting require a lot more context than 32K tokens. Given the utter absence of solid test coverage, a much more likely approach is massive coverage generation via ChatGPT test generation & human audits, and then a gradual rewite by humans assisted by ChatGPT. It's a really good tool, it's nowhere good enough to do automated rewrites. (And as long as the results matter, you'll continue to have humans in the loop for more than 5 years - if for no other reason than cleanly assigning legal blame)
- Frummy 4y agoYes I agree. I work with these systems, they are so complex I think all of it would have to be in memory at once. My "dream" requires many chained miracles and that something like this would be possible at all, abolishing technical debt would only be a footnote of the effects on the world. Anyway I still think it's possible and obviously in one way or another humans will always be in the loop. I'll tell you about the current state, if I write 50 lines of code in a very important place there will be 3 months of testing.
- discordance 4y agoAwesome, thanks for this! I have been thinking a lot about the time when I can start jamming with GPT inside my dev environment and code base and this is a step closer. Use case 3, where you define tests and it tries to give you a passing implementation is the dream.
- deathmonger5000 4y agoThanks for saying this. I feel the same way. There's something magic about telling the robot to "make the tests pass" and watching the implementation magically appear. It does surprisingly well sometimes. I think if I put some work into the prompts then things could perform better and more consistently.
- emaginniss 4y agoI firmly believe that there will always be a horizon effect whereby there will be a solution that matches all of the tests but fails for the general case. Computers are more likely to find that test-solving solution that doesn't handle the general case because of their iterative nature, whereas humans want to solve the general case first and then add in conditions to handle the unexpected cases.
- falcor84 4y agoAs I understand it, the proponents of TDD would argue that as long as the implementation tends towards laziness (avoiding unnecessary complexity), adding more and more tests to handle additional edge cases should cause the code to converse to solving the general case (and avoid local minima). I suppose that the crucial piece is the laziness, including refactoring whenever possible, to save time on addressing each subsequent test case.
- ornornor 4y ago`expect(true).to be true`
- qikInNdOutReply 4y agoImagine, this integrated into software as exception handling. Instead of critical error, fail and send report. Safe input state, create error report + problem description, send to chatgpt, recompile and onwards. Software might never die, but instead "self-repair"..
- photochemsyn 4y agoI'd be worried about the per-token-cost of using the OpenAI API when submitting an entire codebase. And, I'd probably only use it with open-source codebases. I've been using chatblade, it has the nice feature of having a cost estimate call: https://news.ycombinator.com/item?id=35223759 https://news.ycombinator.com/item?id=35223759 As to how to use GPT to help with coding, I wanted to be able to store the chatblade output to a file with a simple one-word call, so I wrote a wrapper that does that, and the first step was to submit this question to GPT: > "I want to use Python's subprocess module in a script (that takes command-line arguments) to manage a call to an OpenAI API (that takes a variable amount of time to complete), and which prints output to the terminal, and I also want to tee that output to a specified file for storage. What do you recommend, from the perspective of an expert Python programmer?" It laid out a nice template using several python modules, and I was able to write it in a single morning with a little bit of additional reference to pydocs and a few more questions about specifics, so now all the command-line queries get appended to a file with an up-to-date cost estimate attached to each one. And I'm pretty junior as far as programming goes.
- deathmonger5000 4y agoThanks, chatblade looks cool I'll check it out! I've found GPT to be a phenomenal tool for solving problems with code.
- braindead_in 4y agoHow can I select a block of code in vim and pass it to this? Maybe I should build a vim plugin for this with ChatGPT.
- baq 4y agoif it supports stdin and stdout, try :'<,'>!promptr -m gpt4 -t refactor -p "Cleanup the code"
- deathmonger5000 4y agopromptr doesn't support stdin in this way, but this is a great reason to make it do so.
- keeptrying 4y agoIs this javascript only at this moment? Looking for something that can handle rails
- deathmonger5000 4y agoYou can use this on any code (or text files) that you want. I've used this for rails, and it's great. I think you'd run into issues with more niche languages where the model hasn't had much exposure.
- MacsHeadroom 4y agoNo, it's any code and even things that aren't code. It will write and revise documentation for you, create jquery, nosql databases, and more. Basically anything a human can do with text files. If that's not enough for you, projects like Auto-GPT give GPT-4 full autonomy to figure out how to do complex multi-step tasks beyond modifying text all on its own with only a vague goal provided. https://github.com/Torantulino/Auto-GPT https://github.com/Torantulino/Auto-GPT
- m3kw9 4y agoMan that’d cost a pretty penny. You have input and output token costs, likely close to dollar per call. And you could need 2/3 calls that get it right, worth it?
- marpstar 4y agoI inherit small- to medium- sized projects from other developers at the agency I work with and to be able to run a project through it and get this sort of analysis without having to find the entry-points and whatnot myself... easily worth the money. It's not something I'd be using with any regularity (on the same project), but I can see where it'd be worth it, even at a dollars-per-project cost.
- redeux 4y agoDepends on your scenario, I guess. If we assume an engineer’s time is worth $75/hr and it would take an hour (much more or less depending on the size and complexity of the code base) to complete one of these tasks then this looks like a bargain, right?
- m3kw9 4y agoit likely won’t be a no brainer once you factor in testing the code and understanding the code depending on your constraints like security, performance, expandablilty etc. Not like you can run it and check in, you’d get fired doing that
- m3kw9 4y agoThis use case is why we need local models that work. In the future you will build a server to host your own models or pay.
- turnsout 4y agoNah, they'll run directly on your edge device. For 90% of tasks, there will be no market for paid or self-hosted inference.
- nico 4y agoWe already do: GPT4All https://github.com/nomic-ai/gpt4all https://github.com/nomic-ai/gpt4all
- m3kw9 4y agoThing is gpt3.5 turbo is so cheap buying hardware to run it don’t make any sense while at the same time, won’t likely able run GPT4 where you are forced to use the server because the hardware cost for that wouldn’t make sense for personal use. This is OpenAIs strategy they need to balance the cost evelope so people are always stuck between a rock and a hard place.
- nico 4y agoIt’s a loosing strategy for OpenAI. We only need so much AI power before we can make our own improved versions by leveraging that initial AI power. Individual developers are now creating models that run on phones and regular computers. OpenAI cannot possibly catch up to what millions of people can do in the open.
- Proven 4y ago[dead]
- jasonjmcghee 4y agoConsider an alternative approach: chunk up and embed your entire codebase (dramatically cheaper) and insert it into a vector store. when you go to send a query, search the store and retrieve the most relevant chunks, and those are what you send with your query. Optionally send other specific code alongside it. This should be similar in quality and dramatically cheaper. It's quite doable with langchain, and you can use Chroma + DuckDB to avoid having to pay for Pinecone. You could even maintain there vector store using git diffs to only update chunks that have changed.
- deathmonger5000 4y agoGreat ideas thank you!
- swader999 4y agoYeah this is the way!
- asteroidz 4y agoThe problem with your suggested approach is the resulting lack of holistic context. The problem of OP's approach (direct parsing) is cost and context-window-limits. There has to be a better way.
- jasonjmcghee 4y agoIt's searching by semantic meaning, so it should be able to find all relevant pieces. Using overlap during chunking should help too. Using the "give it everything" method will cause it to forget most of what you're feeding it if you have a large repo anyway, right?
- deathmonger5000 4y agoI don't think it will forget anything as long as everything fits in the context window, but I could totally be wrong. That's the big problem with the "give it everything" approach: if your codebase doesn't fit then it's game over. I've had success limiting what I give it to the relevant files.
- mzitelli 4y agoGreat to see this here. I am working on a VS Code extension that provides some nice UX to use GPT for autonomous software development. Check it out: https://github.com/MateusZitelli/PromptMate https://github.com/MateusZitelli/PromptMate
- deathmonger5000 4y agoThanks for dropping this here. promptr could sure use a better UX - something like what you've created. Happy to collaborate if there's any interest. Building great tooling for using LLM's to code is something that's really interesting to me.
- mzitelli 4y agoDefinitely, it is inspiring to see so many fantastic initiatives popping up. I am emailing you.
- anotherpaulg 4y agoThanks for sharing promptr! I will try it out. I have also been exploring a similar pattern for using GPT as a coding collaborator: - Send all the (relevant) code to GPT along with a change request - Have it reply with all the code, modified to include the requested change - Automatically replace the original files with the GPT edited versions - Use git diff, etc to review and either accept/reject the changes. GPT is significantly better at modifying code when following this "all code in, all code out" pattern. This pattern has downsides: you can quickly exhaust the context window, it's slow waiting for GPT to re-type your code (most of which it hasn't modified) and of course you're running up token costs. But the ability of GPT to understand and execute high level changes to the code is far superior with this approach. I have tried quite a large number of alternative workflows. Outside the "all code in/out" pattern, GPT gets confused, makes mistakes, implements the requested change in different ways in different sections of the code, or just plain fails. If you're asking for self contained modifications to a single function, that's all the code that needs to go in/out. On the other side of the spectrum, I had GPT build an entire small webapp using this pattern by repeatedly feeding it all the html/css/js along with a series of feature requests. Many feature requests required coordinated changes across html/css/js. https://github.com/paul-gauthier/easy-chat#created-by-chatgpt https://github.com/paul-gauthier/easy-chat#created-by-chatgp... Another HN user has also released a command line tool along these lines called gish: https://github.com/drorm/gish https://github.com/drorm/gish
- deathmonger5000 4y agoWe're thinking the same thing. You might be interested in the prompt template I used to make GPT respond in json format here: https://github.com/ferrislucas/promptr/blob/main/templates/refactor.txt https://github.com/ferrislucas/promptr/blob/main/templates/r... Having GPT's response in json was useful to be able to easily apply the changes GPT wants to the filesystem. GPT4 is significantly more consistent with only responding with json. Gish looks cool!
- nico 4y agoFor people looking for development tools with GPT. One of the best is CodeGPT, with already 350k+ downloads on the visual studio marketplace: https://marketplace.visualstudio.com/items?itemName=DanielSanMedium.dscodegpt https://marketplace.visualstudio.com/items?itemName=DanielSa...
- MacsHeadroom 4y agoI use a local LLaMA API as my code inference endpoint. No way I'm just shipping all my code to OpenAI.
- nico 4y agoYou can also use GPT4All https://github.com/nomic-ai/gpt4all https://github.com/nomic-ai/gpt4all
- MacsHeadroom 4y agoI use Vicuna[0]. It's much better than GPT4All. Vicuna is based on 13B (not 7B) and its training data includes humans chatting with GPT-4 vs GPT4All's purely synthetic dataset generated by GPT-3.5. [0] https://github.com/lm-sys/FastChat https://github.com/lm-sys/FastChat
- nico 4y agoThank you! Could Vicuña be used to further fine-tune GPT4All to make it better?
- MacsHeadroom 4y agoI think GPT4All's inferior quality dataset would make a worse combined model than strict Vicuna. Vicuna-30B will likely be better than GPT-3.5 level and approaching GPT-4 level when it's done training, but run slow on CPU.
- superb-owl 4y agoIt looks like this only works on JavaScript? It errors if you don't have a package.json
- deathmonger5000 4y agoI'd love it if you opened an issue on the github repo. I haven't seen this happen. My guess it that maybe you have an older version of node, but I might be way off. Maybe I should dockerize this, so you wouldn't need node installed to use it.
- hegem0n 4y agoI want to see the evolution of context management tools for coding with GPT. - implement a function, providing structural context (e.g. model, service, controller) with stubbed modules, but leave out utilities/libraries and database model - drop the structural context, and ask it to expand the stub implementation with some additional context about the database model - drop the database model context, and ask it to refactor the solution with some additional context about utilities/libraries I think this is doable, if we can build up some utilities for context management and iterative development, GPT should start to be usable on large code bases. It could work similarly to how one person wrote an entire novel using GPT.
- anotherpaulg 4y agoI have been experimenting with exactly this. I hit the context window limit while developing a small web app [1] by having GPT do all the coding. I've tried a few things: 1. Summarize/collapse the code. - Use GPT to summarize code blocks into a 1 line comment. - Collapse all the top level code blocks into their summary line. This turns the entire codebase into a much smaller "top level map". - Use a ReAct pattern via langchain to have the LLM itself try and determine which collapsed blocks need to be recursively "expanded" to understand the code with respect to the user requested feature/bugfix/change/modification. - Feed GPT this partially expanded code base along with the user requested change. - Have it spit back the modified version of that partially expanded code. - Apply the GPT changes back to the original source files. 2. Explicitly ask GPT "which parts of the code are relevant to this change request" or "which parts of this code would need to be modified to make the needed changes", etc. 3. Again using ReAct, give GPT tools to "grep" and "cat" the code so that it can explore the codebase to find and understand relevant chunks. I've even armed it with bash in a sandbox. It started importing and running parts of the python code I was operating over. None of these approaches have panned out fully yet. But there are promising signs. [1] This is the small webapp that GPT wrote for me. I'm working on it mainly as a forcing function to explore these sorts of "GPT as junior developer / coding collaborator" workflows. https://github.com/paul-gauthier/easy-chat#created-by-chatgpt https://github.com/paul-gauthier/easy-chat#created-by-chatgp...
- 4y ago
- mherrmann 4y agoCool stuff. I recently created a similar tool that automatically fixes errors in source code: https://github.com/mherrmann/fix https://github.com/mherrmann/fix
- deathmonger5000 4y agoThe tool you linked looks really cool nice job! "What could possibly go wrong?" - I love it!
- mherrmann 4y agoThank you! Yours looks very nice as well :)
- mov_eax_ecx 4y agoI don't think that the world needs yet another chatgpt proxy to mess with source code when copilot already fills the need, they have the engineers and lawyers to make it work. What I don't see is, unique complex applications with a great UI that don't overlap with the applications like office with copilot, or copilotX.
- deathmonger5000 4y agoThis isn't a proxy. Say you use ChatGPT to assist you while you write code... there's a lot of copy paste action in that workflow. It gets old. promptr gets rid of the copy paste. That's a much better developer ergonomic (IMO)
- tuchsen 4y agoPlus your tool is open source, and direct calls to the OpenAI API are a much cheaper than Copilot. If anything Copilot's looking like the redundant one here to me. Thanks for doing Promptr, it looks great I'm going to give it a shot
- deathmonger5000 4y agoThank you! Have fun - please share if you do anything cool!
- mov_eax_ecx 4y agotake a look at https://www.youtube.com/watch?v=s7AGkcSMiaI https://www.youtube.com/watch?v=s7AGkcSMiaI for a demo of what copilot is capable of, did you try copilot and copilot labs?, they have the best DX, look the interviews of nat fridman on the tool. The code is a prompt, a foreach file execute prompt and return whatever, it can be used as a template build something useful and not redundant, but definitely not code. Don't reinvent the wheel.