3 ms·
I've been working on some basics in this direction... taking the expected type at the cursor and using this to inform exactly what type & function definitions t
by disconcision 3y ago
I've been working on some basics in this direction... taking the expected type at the cursor and using this to inform exactly what type & function definitions to splice into the prompt: https://andrewblinn.com/papers/2023-MWPLS-Type-directed-Prompt-Construction-for-LLM-powered-Programming-Assistants.pdf https://andrewblinn.com/papers/2023-MWPLS-Type-directed-Prom...
We're currently working on two forks off this... one, using the expected type information in conjunction with existing grammar constrained generation to enforce per-token type-correct generations, and two, using some program synthesis unevaluation techniques to also provided runtime trace data relevant to the current program hole we're trying to generate a completion for.
But yeah in general there is so much existing work on static and dynamic program analysis which can be applied here... I think a lot of the interesting challenges are going to be UI/UX ones... interactive processes to help more precisely specify intent and iteratively valid generations as more and more code is written autonomously.
- pedrovhb 3y agoThat's interesting. In my experience, although synctatic errors do happen, they're not nearly as common as, say, hallucinating methods or tripping over the exact API of a library. One thought I've had but haven't experimented with yet is that we could leverage a lot of the existing tooling that was made for humans - e.g. tab-complete providers like Jedi, which do some type inference behind the scenes. It's able to provide suggestions of valid members for a given cursor position, and so logits could be warped to prefer tokens which match the suggestions (so if the output at a given time is `math.sq`, `math.sqrt` would be much preferred over `math.square_root` which doesn't exist). You'd have to be a bit smart around this though because in since situations such as when using an identifier for the first time it's not yet in scope and you don't want the LLM to never create variables. Maybe some beam search shenanigans and heuristics could be enough to get useful output, but at that point it no longer seems like a quick thing to just try out :)
- disconcision 3y agoit's been done, in a limited way: https://arxiv.org/abs/2306.10763 https://arxiv.org/abs/2306.10763; we're working on extended this, but yeah, there are a lot of details to work out