8 ms·
> Reading through these commits sparked an idea: what if we treated prompts as the actual source code? Imagine version control systems where you commit the prom
by SrslyJosh 1y ago
> Reading through these commits sparked an idea: what if we treated prompts as the actual source code? Imagine version control systems where you commit the prompts used to generate features rather than the resulting implementation.
Please god, no, never do this. For one thing, why would you not commit the generated source code when storage is essentially free? That seems insane for multiple reasons.
> When models inevitably improve, you could connect the latest version and regenerate the entire codebase with enhanced capability.
How would you know if the code was better or worse if it was never committed? How do you audit for security vulnerabilities or debug with no source code?
- Sevii 1y agoThere are lots of reasons not to do it. But if LLMs get good enough that it works consistently people will do it anyway.
- minimaxir 1y agoWhat will people call it when coders rely on vibes even more than vibe coding?
- roywiggins 1y agoHaruspicy?
- brookst 1y agoWriting specs
- auggierose 1y agoExactly my thought. This is just natural language as a specification language.
- kiitos 1y ago...as an ambiguous and inadequately-specified specification language.
- auggierose 1y agoIn the end, every specification is specified via natural language, this is just where the buck stops. All math books are written in natural language, even the ones about specification languages.
- kiitos 1y agoHuh? Is ABNF a "natural language"? Is the Go language spec a "natural language"?
- auggierose 1y agoHow is ABNF itself specified? Yes, via natural language. And the Go language spec is written in natural language, too, you can check for yourself: https://go.dev/ref/spec https://go.dev/ref/spec
- kiitos 1y agoABNF itself is specified with a well-defined grammar and syntax...
- auggierose 1y agoYes, but the spec is still done in natural language: https://www.rfc-editor.org/rfc/rfc5234 https://www.rfc-editor.org/rfc/rfc5234 Just like any other RFC.
- rectang 1y ago>> what if we treated prompts as the actual source code? You would not do this because: unlike programming languages, natural languages are ambiguous and thus inadequate to fully specify software.
- deleted 1y ago[deleted]
- a012 1y agoPrompts are like story on the board, and like engineers, depends on the understanding of the model the generated source code can vary. Saying the prompts could be the actual code is so wrong and dangerous thought
- squillion 1y agoExactly! > this assumes models can achieve strict prompt adherence What does strict adherence to an ambiguous prompt even mean? It’s like those people asking Babbage if his machine would give the right answer when given the wrong figures. I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a proposition.
- tayo42 1y agoI'm pretty sure most people aren't doing "software engineering" when they program. There's the whole world of WordPress and dream Weaver like programing out there too where the consequences of messing up aren't really important. Llms can be configured to have deterministic output too
- fastball 1y agoThe idea as stated is a poor one, but a slight reshuffling and it seems promising: You generate code with LLMs. You write tests for this code, either using LLMs or on your own. You of course commit your actual code: it is required to actually run the program, after all. However you also save the entire prompt chain somewhere. Then (as stated in the article), when a much better model comes along, you re-run that chain, presumably with prompting like "create this project, focusing on efficiency" or "create this project in Rust" or "create this project, focusing on readability of the code". Then you run the tests against the new codebase and if the suite passes you carry on, with a much improved codebase. The theoretical benefit of this over just giving your previously generated code to the LLM and saying "improve the readability" is that the newer (better) LLM is not burdened by the context of the "worse" decisions made by the previous LLM. Obviously it's not actually that simple, as tests don't catch everything (tho with fuzz testing and complete coverage and such they can catch most issues), but we programmers often treat them as if they do, so it might still be a worthwhile endeavor.
- stingraycharles 1y agoMeans the temperature should be set to 0 (which not every provider supports) so that the output becomes entirely deterministic. Right now with most models if you give the same input prompt twice it will give two different solutions.
- NitpickLawyer 1y agoEven at temp 0, you might get different answers, depending on your inference engine. There might be hardware differences, as well as software issues (e.g. vLLM documents this, if you're using batching, you might get different answers depending on where in the batch sequence your query landed).
- derwiki 1y agoTwo years ago when I was working on this at a startup, setting OAI models’ temp to 0 still didn’t make them deterministic. Has that changed?
- 1y ago
- renewiltord 1y agoIt’s been a thing people have done for at least a year https://github.com/i365dev/LetterDrop https://github.com/i365dev/LetterDrop
- gizmo686 1y agoMy work has involved a project that is almost entirely generated code for over a decade. Not AI generated, the actual work of the project is in creating the code generator. One of the things we learned very quickly was that having generated source code in the same repository as actual source code was not sustainable. The nature of reviewing changes is just too different between them. Another thing we learned very quickly was that attempting to generate code, then modify the result is not sustainable; nor is aiming for a 100% generated code base. The end result of that was that we had to significantly rearchitect the project for us to essentially inject manually crafted code into arbitrary places in the generated code. Another thing we learned is that any change in the code generator needs to have a feature flag, because someone was relying on the old behavior.
- mschild 1y ago> One of the things we learned very quickly was that having generated source code in the same repository as actual source code was not sustainable. Keeping a repository with the prompts, or other commands separate is fine, but not committing the generated code at all I find questionable at best.
- djtango 1y agoI didn't read it as that - If I understood correctly, generated code must be quarantined very tightly. And inevitably you need to edit/override generated code and the manner by which you alter it must go through some kind of process so the alteration is auditable and can again be clearly distinguished from generated code. Tbh this all sounds very familiar and like classic data management/admin systems for regular businesses. The only difference is that the data is code and the admins are the engineers themselves so the temptation to "just" change things in place is too great. But I suspect it doesn't scale and is hard to manage etc.
- diggan 1y agoIf you can 100% reproduce the same generated code from the same prompts, even 5 years later, given the same versions and everything then I'd say "Sure, go ahead and don't saved the generated code, we can always regenerate it". As someone who spent some time in frontend development, we've been doing it like that for a long time with (MB+) generated code, keeping it in scm just isn't feasible long-term. But given this is about LLMs, which people tend to run with temperature>0, this is unlikely to be true, so then I'd really urge anyone to actually store the results (somewhere, maybe not in scm specifically) as otherwise you won't have any idea about what the code was in the future.
- mellosouls 1y agoYes, it's too early to be doing that now, but if you see the move to AI-assisted code as at least the same magnitude of change as the move from assembly to high level languages, the argument makes more sense. Nobody commits the compiled code; this is the direction we are moving in, high level source code is the new assembly.
- Xelbair 1y agoWorse. Models aren't deterministic! They use temperature value to control randomness, just so they can escape local minima! Regenerated code might behave differently, have different bugs(worst case), or not work at all(best case).
- chrishare 1y agoNitpick - it's the ML system that is sampling from model predictions that has a temperature parameter, not the model itself. Temperature and even model aside, there are other sources of randomness like the underlying hardware that can cause the havoc you describe.
- never_inline 1y agoApart from obvious non-reproducibility, the other problem is lack of navigable structure. I can't command+click or "show usages" or "show definition" any more.
- saagarjha 1y agoJust ask the AI for those obviously
- visarga 1y agoThe idea is good, but we should commit both documentation and tests. They allow regenerating the code at will.
- pollinations 1y agoI'd say commit a comprehensive testing system with the prompts. Prompts are in a sense what higher level programming languages were to assembly. Sure there is a crucial difference which is reproducibility. I could try and write down my thoughts why I think in the long run it won't be so problematic. I could be wrong of course. I run https://pollinations.ai https://pollinations.ai which servers over 4 million monthly active users quite reliably. It is mostly coded with AI. Since about a year there was no significant human commit. You can check the codebase. It's messy but not more messy than my codebases were pre-LLMs. I think prompts + tests in code will be the medium-term solution. Humans will be spending more time testing different architecture ideas and be involved in reviewing and larger changes that involve significant changes to the tests.
- maxemitchell 1y agoAgreed with the medium-term solution. I wish I put some more detail into that part of the post, I have more thoughts on it but didn't want to stray too far off topic.
- 7speter 1y agoI think the author is saying you commit the prompt with the resulting code. You said it yourself, storage is free, so comment the prompt along with the output (don’t comment that out that if I’m not being clear); it would show the developers(?) intent, and to some degree, almost always contribute to the documentation process.
- maxemitchell 1y agoAuthor here :). Right now, I think the pragmatic thing to do is to include all prompts used in either the PR description and/or in the commit description. This wouldn't make my longshot idea of "regenerating a repo from the ground up" possible, but it still adds very helpful context to code reviewers and can help others on your team learn prompting techniques.
- torben-friis 1y agoPlus, commits depend on the current state of the system. What sense does “getting rid of vulnerabilities by phasing out {dependency}” make, if the next generation of the code might not rely on the mentioned library at all? What does “improve performance of {method}” mean if the next generation used a fully different implementation? It makes no sense whatsoever except for a vibecoders script that’s being extrapolated into a codebase.
- croes 1y agoYou couldn’t even tell in advance if the prompt produces code at all.
- lowsong 1y agoI'm the first to admit that I'm an AI skeptic, but this goes way beyond my views about AI and is a fundamentally unsound idea. Let's assume that a hypothetical future AI is perfect. It will produce correct output 100% of the time, with no bugs, errors, omissions, security flaws, or other failings. It will also generate output instantly and cost nothing to run. Even with such perfection this idea is doomed to failure because it can only write code based on information in the prompt, which is written by a human. Any ambiguity, unstated assumption, or omission would result in a program that didn't work quite right. Even a perfect AI is not telepathic. So you'd need to explain and describe your intended solution extremely precisely without ambiguity. Especially considering in this "offline generation" case there is no opportunity for our presumed perfect AI to ask clarifying questions. But, by definition, any language which is precise and clear enough to not produce ambiguity is effectively a programming language, so you've not gained anything over just writing code.
- handoflixue 1y agoWe already have AI agents that can ask a human for help / clarification in those cases. It could also analyze the company website, marketing materials, and so forth, and use that to infer the missing pieces. (Again, something that exists today)
- layer8 1y agoIf the AI has to ask for clarification, you can’t run it as a reproducible build step as envisaged. It’s as if your compiler would pause to ask clarifying questions on each CI run. If the company website, marketing materials, and so forth become part of the input, you’ll have to put those in version control as well, as any change is likely to result in a different application being generated (which may or may not be what you want).
- gitgud 1y agoThis is so eloquently put and really describes the absurdity of the notion that code itself will become redundant to building a software system
- paxys 1y agoForget different model versions. The exact same model with the exact same prompt will generate vastly different code each subsequent time you invoke it.
- dragonwriter 1y agoAlso, while it is in principle possible to have a deterministic LLM, the ones used by coding assistants aren't deterministic, so the prompts would not reliably reproduce the same software. There is definitely an argument, for also committing prompts, but it makes no sense to only commit prompts.
- TechDebtDevin 1y agoSome code is generated on the fly, like llm ui/ux that writes python code to do math. Idk kinda different tho.