6 ms·
>> on a corpus of ... abstracts Wait a minute. Isn't the abstract the summary? I've read accounts of people digging through a paper and seeing impossible, obvi
by swsieber 3y ago
>> on a corpus of ... abstracts
Wait a minute. Isn't the abstract the summary? I've read accounts of people digging through a paper and seeing impossible, obviously fabricated numbers. I don't think those types of issues would surface in an abstract.
- visarga 3y agoWhat they are doing is just conditioning on the topic, author, institution and media attention to predict replicability. But they predict averages, can't make reliable predictions about specific papers. And the study is focusing only on Psychology, even though they trained their embeddings on 2m paper abstracts.
- mike_hearn 3y agoWell it's worse than that. It's surprisingly common to read a paper and discover the abstract doesn't match the claims in the body of the paper itself, or that the claims don't match the data in the tables. And of course press releases about a paper sometimes don't match the paper abstract, etc. The amount of slippage in what's being claimed as papers get summarized is far too large in some fields and a major contributor to distrust in science, as it causes authors to try and have cake+eat it. Get challenged on an untrue statement you made to the public? Refer to the body of the paper where the claim isn't made, or is made equivocally, and then blame journalists for "mischaracterizing" your work. It probably is possible to pick up signs of pseudo-science automatically with ML, but you'd want to use really big GPT-4 style LLMs with large context windows and detailed instructions. We're not quite there yet, but maybe next year.