4 ms·
> At the same time, SAFE is more than 20 times cheaper than human annotators.> LLMs have achieved superhuman performance on reason- ing benchmarks (Bubeck et al
by Take8435 3y ago
> At the same time, SAFE is more than 20 times cheaper than human annotators.> LLMs have achieved superhuman performance on reason- ing benchmarks (Bubeck et al., 2023; Gemini Team, 2023; Li et al., 2022) and higher factuality than humans on sum- marization tasks (Pu et al., 2023). The use of human annota- tors, however, is still prevalent in language-model research, which often creates a bottleneck due to human annotators’ high cost and variance in both evaluation speed and qual- ity. Because of this mismatch, we investigate how SAFE compares to human annotations and whether it can replace human raters in evaluating long-form factuality.
I am not an expert in this research but this seems like this is just a slippery slope all in the pursuit of cost, first and foremost.
Given that the summary focuses on cost and this paragraph mentions cost as the first point, it sure seems like these folks only goal goal is to just take humans out of the mix entirely when it comes to facts.
Is this a good idea? I am not sure.
- skeledrew 3y agoIf they can guarantee human level or better performance, no problem at all (from a technical perspective; social is another matter).
- cl42 3y agoI'd argue the issue is data set drift. This works with Google _today_ but will it work with Google tomorrow? Will it work when Google changes its search algorithms/rankings? Will it work when AI overtakes human content on Google? You can't trust that this will work on your knowledge domain or that it'll work in the future.