4 ms·
This wasn't that hard to see. > Our systematic review of 445 benchmarks reveals prevalent gaps that undermine the construct validity needed to accurately measu
by proc0 11mo ago
This wasn't that hard to see.
> Our systematic review of 445 benchmarks reveals prevalent gaps that undermine the construct validity needed to accurately measure
targeted phenomena
Intelligence has an element of creativity, and as such the true measurement would be on metrics related to novelty, meaning tasks that have very little resemblance to any other existing task. Otherwise it's hard to parse out whether it's solving problems based on pattern recognition instead of actual reasoning and understanding. In other words, "memorizing" 1000 of the same type of problem, and solving #1001 of that type is not as impressive as solving a novel problem that has never been seen before.
Of course this presents challenges to creating the tests because you have to avoid however many petabytes of training data these systems are trained with. That's where some of the illusion of intelligence arises from (illusion not because it's artificial, since there's no reason to think the brain algorithms cannot be recreated in software).
- peacebeard 11mo agoIn my opinion a major weakness in how people reason about this issue is that they describe solving problems as EITHER recall of an existing solution OR creative problem solving. Sure it is possible for a specific solution to be recalled, but it's not possible for a problem to be absolutely unrelated to anything the system has ever seen before and still be solvable. There are many shades of gray in the similarity a problem may have to previously seen problems. In fact, I expect that there are as many shades of gray as there are problems.
- proc0 11mo agoThe difference is that humans don't memorize petabytes of problems, so from a relative perspective people are constantly solving novel problems they never saw before. I'm thinking this is a requirement for dynamic, few-shot learning. We can clearly see LLMs fail when you throw even a small wrench in the prompt.
- peacebeard 11mo agoHumans encounter massive numbers of problems in their experience that informs their problem solving. The same is true of LLM. LLMs do not actually have all their training data memorized. I’m not sure what your basis is for saying “LLMs fail if there is a small wrench in the prompt.” They also succeed despite wrenches in the prompt with great regularity.
- proc0 11mo agoLet's clarify, this isn't about whether the models are capable. They are very capable and impressive. This is more about whether we can use the same type of metric we use for humans to compare and conclude if they are "intelligent". It's not just semantics, the metrics are supposed to tell us the potential of the model. If they can solve extremely hard PhD problems, it should be the case that we're already in the singularity, and they should be solving absolutely everything in whatever field they were trained in, because it's not just PhD level, it's a machine that has a ton of memory, compute and never sleeps. However, once you use these models extensively, it becomes apparent they are just synthesizing data, and not as much understanding it in a way that would allow them to extrapolate into anything else as humans do. I think this point is a little hard to explain. I'll just emphasize, these are smart systems, and they can do a lot, but there is still a disconnect between, let's say, a PhD level model and a human with a PhD, in the "quality" of what we would call "intelligence" of both entities (human and machine).
- Kostchei 11mo agoHuman metrics of intelligence have always felt like rubbish. We never did this well. I would describe intelligence as effective adaption leading to survival and growth or prospering. Memorization, comprehension, speed of response etc. those are magnifying factors that are valued, we view them as components of intelligence, but llms are proving this is not the whole, without effective application, they are not intelligence. Perhaps learning is the difference? How to measure that? Someone describing string theory is the literary equivalent of fractal structures in snowflakes. Lovely, complex, possibly unique, but not proof of a level of intelligence- for the string theorist maybe it is intelligent, perhaps persuading someone to fund their grant, which enables them to eat, shelter etc. Might be a bit harsh on string theory. Saying it is proof of an amount of intelligence leads us to falsifiable statements.