5 ms·
I'd guess these model's understand works more closely to people so encoding in text is more token efficient and things like comments help. Also syntax seems a
by eldenring 3y ago
I'd guess these model's understand works more closely to people so encoding in text is more token efficient and things like comments help.
Also syntax seems a lot easier to understand for them than semantics/logic. If you've used GPT-4 it almost never makes syntax errors. Logical errors on the other hand...
- kevinlu1248 3y agoFrom my experience, GPT-4 never makes syntax errors directly but when making edits to existing code it's harder to prevent these syntax errors from appearing. We used to add a second pass to check for these syntax errors. It also frequently makes undefined variables and the like, however.
- reitzensteinm 3y agoDid you get rid of the second pass? I'm working on something quite similar and find a pass that inspects and rejects erroneous code to be a big boost to correctness.
- kevinlu1248 3y agoWe got rid of it. Our new edit framework works around search-and-replace pairs with an example at https://github.com/sweepai/sweep/blob/d37dda3a626f09dea3b3221dd0254671407ccc1b/sweepai/core/prompts.py#L325-L329 https://github.com/sweepai/sweep/blob/d37dda3a626f09dea3b322...
- _boffin_ 3y agoI’ve built out an end to end automated fix pipeline, it’s getting the bug fixes right, but been having trouble with line number errors. Looking forward to reading through your docs and repo later tonight to see how you’re addressing issues like this.
- kevinlu1248 3y agoWe used to use line numbers but it became problematic so we switched over to search-and-replace pairs, which works significantly better. The only potential problems are with setting up a fuzzy search system since sometimes the search doesn't match exactly with the code (missing comments, etc.). We're going to write about our core algo and diff managing system soon.
- reitzensteinm 3y agoAh, interesting. I completely abandoned diff style updates in favour of AST substitution, but that was only possible because my tradeoffs are different to yours. I'm building a bot that's building itself, so it doesn't have to support large legacy code bases with different languages.
- kevinlu1248 3y agoThis is interesting, I'm wondering what you mean by AST substitution. Is this like an agent that traverses the tree and picks what to edit? Is this language model based? Also, thankfully we don't support too many uncommon languages. The most recent ones we added support for are embedded templates (ERB/EJS for flask and Ruby) and mustache. Fortunately many uncommon languages are subsets of other languages.
- reitzensteinm 3y agoThe agent specifies what function, class, method etc to replace, along with its full source. It's more costly, but I believe it leads to fewer hallucinations as it is generating a coherent piece of code. But it requires parsing AST and language specific instructions. And things like metaprogramming or macros could cause some hairy confusion. All of these factors don't hurt my use case.
- kevinlu1248 3y agoWe have a similar method under the hood, except it's purely text-based search-and-replace. The model decides what to replace. It seems to be consistent and is easy to implement.
- reitzensteinm 3y agoMy gut feeling based on my experience over the last couple of months is that substitution of an entire function is more reliable than some lines of a function. The surrounding context reduces the chance of hallucinations. Gut feeling doesn't account for much though - I'm working on an evals system to be able to quantify system performance. It won't be cheap to run. It could easily be that your method is superior.
- anotherpaulg 3y agoThat edit format looks familiar! https://github.com/paul-gauthier/aider/blob/d5a7aac560d4584dc66da39d55af734a74febada/aider/coders/editblock_prompts.py#L24 https://github.com/paul-gauthier/aider/blob/d5a7aac560d4584d... https://aider.chat/docs/benchmarks.html https://aider.chat/docs/benchmarks.html
- kevinlu1248 3y agoYup it's based on the aider blogs. They're perfect for our use case and are very reliable compared to our old attempts.