3 ms·
> say "better than 95% of humans on 95% of intellectual tasks" but if we used that definition we already have AGI and almost no one thinks we have achieved AGI
by computably 16d ago
> say "better than 95% of humans on 95% of intellectual tasks" but if we used that definition we already have AGI and almost no one thinks we have achieved AGI
What matters isn't 95% of humans, it's 95% of actual professionals. Benchmarking an AI accountant against people with zero accounting experience is worse than worthless.
- EthanHeilman 15d agoThat's what 95% is intended to capture. That some level of expertise in an area is should be captured by 95% sample of the population, you could push it to 99% of 99.9%, but you want it to be quantified by a number to avoid arguments that something is not AGI because obscure field or expert exists that AI can not do. 95% better at 95% of the population is already approaching ASI. One could even argue that AGI is 50% better than 50% of the population. All that said, this over rotation on benchmarking misses something critical. What is general intelligence? We assume that humans have it and we assume it is captured by benchmarks on "intellectual tasks", but it is probably the case that general intelligence is based displayed by judgement on uncertain outcomes. Benchmarks by their very nature have certain outcomes, they have a wrong and right answer. Test AIs on questions we don't have the answers to and there is no clear right answer, but there will be at some point in the future. What will the economy do? Which US senators will be be re-elected that polling correctly suggests will not be re-elected. Which US senators, currently not in office, will actually pass bills representing the wishes of their voting base? What published papers will be seen are groundbreaking in 5, 10 15 years? What approach to unifying physics should be taken?