3 ms·
The problem with all of these trick questions is that you don't know what was already in the training data - readily and easily available for completion. I.e.
by beders 2y ago
The problem with all of these trick questions is that you don't know what was already in the training data - readily and easily available for completion.
I.e. you can't tell if a result was produced by "reasoning" or by a simple lookup.
- Hugsun 2y agoThat's very true. I thought about speaking more to that issue but the post was already longer than I wanted. You might find this short analysis interesting. https://www.arnaldur.be/experimenting/with/large-language-models-1 https://www.arnaldur.be/experimenting/with/large-language-mo... Here I do an exhaustive analysis of a range of simple math questions. I then visualize the output so you can see the failures. I do contend though that the tikz question mentioned in the post can't possibly be wholly represented in the training data. There is of course tikz code on the internet but it find it highly likely that the model extrapolated based on having seen my tikz code, and a bunch of other tikz code. I go further to contend that the extrapolation requires reasoning to achieve.
- imtringued 2y agoWell, arithmetic problems are good simple benchmarks for multi step problem solving, but then you get the weirdos who claim you shouldn't let an LLM do the dirty work. When someone uses addition as a benchmark, they are not necessarily in need of a calculator, they are in need of an LLM that demonstrates the skills that it shares in common with arithmetic and the actual task. For example, if you are given the task to calculate ten additions, you must know how to split the problem into subproblems until you can compose a series of skills to actually solve the problem. If the model knows a different operation like subtracting, it will know how to do so even though it hasn't seen an example involving ten subtractions. Apply this to logical operators, set operators, modus ponens, etc and you will get very far even with very rudimentary skills. The point here is that it should be easy for a machine to verify that the LLM does indeed have to think for itself.