4 ms·
I am as enthusiastic about formal methods as the next guy, but I very much doubt any LLM-based technique will make it economical to write a substantial fraction
by bwestergard 7mo ago
I am as enthusiastic about formal methods as the next guy, but I very much doubt any LLM-based technique will make it economical to write a substantial fraction of application software in Lean. The LLM can play a powerful heuristic role in searching for proof-bearing code in areas where there is good training data. Unfortunately those areas are few and far between.
Moreover, humans will still need to read even rigorously proved code if only to suss out performance issues. And training people to read Lean will continue to be costly.
Though, as the OP says, this is a very exciting time for developing provably correct systems programming.
- zozbot234 7mo agoLLMs are writing non-trivial math proofs in Lean, and software proofs tend to be individually easier than proofs in math, just more tedious because there's so much more of them in any non-trivial development. Some performance issues (asymptotics) can be addressed via proof, others are routinely verified by benchmarking.
- madrox 7mo agoThis assumes everything about current capabilities stay static, and it wasn't long ago before LLMs couldn't do math. Many were predicting the genAI hype had peaked this time last year. If you want it to be a question of economics, I think the answer is in whether this approach is more economical than the alternative, which is having people run this substrate. There's a lot of enthusiasm here and you can't deny there has been progress. I wouldn't be so quick to doubt. It costs nothing to be optimistic.
- candiddevmike 7mo ago> and it wasn't long ago before LLMs couldn't do math They still can't do math.
- Hammershaft 7mo agoPro models won gold at the international math olympiads?
- bandrami 7mo agoThey have trouble adding two numbers accurately though
- evrimoztamur 7mo agoWhy are they expected to?
- otabdeveloper4 7mo ago[*] According to cloud LLM provider benchmarks.