5 ms·
There are LLM SQL benchmarks. [1] And state of the art solution is still only at 77% accuracy. Would you trust that? [1] https://bird-bench.github.io/ https://
by moltar 2y ago
There are LLM SQL benchmarks. [1] And state of the art solution is still only at 77% accuracy. Would you trust that?
[1] https://bird-bench.github.io/ https://bird-bench.github.io/
- flappyeagle 2y agoYes. Ask it to do it 10 times and pick the right answer
- pclmulqdq 2y agoThat only works if you assume the fail cases are uncorrected. Spoiler alert: they are not.
- flappyeagle 2y agoAsk 10 different models then
- pclmulqdq 2y agoSame problem: The models are also correlated on what they can and can't solve. To give you an extreme example, I can ask 1000000 different models for a counterexample to the 3n + 1 problem, and all will get it wrong.
- deleted 2y ago[deleted]
- flappyeagle 2y agoNo. What a bizarre example to choose. This is so easy to demonstrate. They will all come back with the exact same correct answer
- pclmulqdq 2y agoIf it's so easy, go do it. You can publish the result in any math journal you like with just a title and a number, because this is one of the hardest problems in mathematics. For reference: https://en.wikipedia.org/wiki/Collatz_conjecture https://en.wikipedia.org/wiki/Collatz_conjecture
- flappyeagle 2y agoMy guy, every LLM has read Wikipedia
- pclmulqdq 2y agoI don't know if you're purposely being dense. The first sentence of Wikipedia is that this is a famous unsolved problem. So no, sampling 1000000 LLMs will not get you a solution to it. I guarantee you that.
- flappyeagle 2y agoIt will get you the correct answer, not a solution. Once again it’s a terrible example, I don’t know why you used it. It’s certainly not a gotcha
- pclmulqdq 2y agoThe reason I used it is that the correct answer to the actual problem is unknown and nobody has any idea how to solve it. No amount of sampling an LLM will give you a correct answer. It will give you the best known answer today, but it won't give you a correct answer. This is an example where LLMs all give correlated answers that do not solve the problem. If you want to scale back, many programming problems are going to be like this, too. Failure points of different models are correlated as much as failure points during sampling are correlated. You only gain information from repeated trials when those trials are uncorrelated, and sampling multiple LLMs is still correlated.