3 ms·
ARB: Advanced Reasoning Benchmark For Large Language Models (2023)
- nabla9 3y agoThese broad context reasoning benchmarks miss the point. Researchers in child cognitive psychology make better tests for reasoning ability. Children have very limited knowledge, so the tests are based on core concepts like numbers, physical object properties like that object can't be in two places simultaneously. Also some objects have agency and some don't, some things are platonic and some concrete. For example, can LLM can determine reliably what is a "moral subject", or "has agency" using simple language that is suitable for children (no specific terminology that works as a hint). Surely LLM can't do legal reasoning if it can't do that. In my own experiments, LLM's are too much syntax driven to be reliable. When you frame the exact same question in 10 different ways, you get 4 different answers. As long as I can make LLM to think that Number 4 pencil is responsible for stabbing someone, it really can't do common sense reasoning.