3 ms·
Code generation should be relatively straightforward with language model. Start with a problem description and let the LLM generate the code. It should be fairl
by choeger 2y ago
Code generation should be relatively straightforward with language model. Start with a problem description and let the LLM generate the code. It should be fairly simple to ask for a proof of correctness for the generated code, then verify that proof (for instance by creating property-based tests), and finally ask for a human-readable summary of the proof. Pass this to human code review and you have a pretty good code generation machine. Failure at any step can simply lead to another iteration.
The thing is: SE doesn't work like this anymore. People aren't used to formalized requirements and take 99% just for granted. (Which web developer really takes network errors into account nowadays?).
With "AI" we might have the tools we need to really solve the software crisis.
- intended 2y agoThis is likely why Coding assistants are talked about in glowing terms. “35% increased coder productivity with co-pilot.” This still hits the “expert required” problem. Use an LLM to code in a language or domain you have 0 experience in. Contrast your experience with that of any explainer video where an expert in a domain uses an LLM to make an MVP. Experts will get something up in the span of a video. They will avoid innumerable mis-steps and rabbit holes because they know code that gets to the result effectively on sight. Non experts wont even know what part of the output is wrong. ——- Generalize this problem away from code. Think of financial models and another version of the verification problem. Your LLM works, it feeds signal into a financial model that predicts the future value of a firm. The model output is dramatically out of sync with market estimates. If an analyst showed this result, they would be told to fix their excel or tweak the assumptions. Thats the typical “verification” loop. Let’s say hyou tune the model to get to 99.9%. I still have to run sampling /EVALS / QC all the time to make sure my output is within tolerances. I still need to know that I can expect X% accuracy, that we have a plan for the 1-X% cases, and that this isn’t going to overwhelm the people with domain expertise in my team.