3 ms·
https://arxiv.org/abs/2303.12712 https://arxiv.org/abs/2303.12712 > As far as I'm aware they were still hallucinating pretty hard and incredibly biased and ove
by mrshadowgoose 3y ago
https://arxiv.org/abs/2303.12712 https://arxiv.org/abs/2303.12712
> As far as I'm aware they were still hallucinating pretty hard and incredibly biased and overly agreeable.
Certainly. But those aren't blockers to displaying the ability to reason within limited contexts. There are hundreds of thousands of human beings that fit that description.
- dmbche 3y agoThanks for the paper! I'll be looking into it. The reason I'm bringing this up is that these issues, in humans, are solved by redundency (i.e. hiring someone else to look over their work). As long as errors persist and, more importantly, are impossible/hard to be warned of, competent humans will need to oversee and validate every word they type, making their "advantage" much more nuanced. It kind of turns into having a friend that can google pretty well answer your questions. I'm not an expert, feel free to show me wrong - I'm reading the paper now and expecting it to be at least a touch troubling :)
- MacsHeadroom 3y agoThere's another paper which showed GPT-4 (or was it 3.5?) going from something like 78% to 99.9% success on difficult multi-part reasoning tasks by running two copies of itself where one is an editor/reviewer. I'm on my phone at the moment. But maybe someone else can link it.
- p-e-w 3y ago> As long as errors persist and, more importantly, are impossible/hard to be warned of, competent humans will need to oversee and validate "Competent" humans make errors too. This is a non-argument.