4 ms·
Arxiv accidentally created a crackpot filter from apparently a simple semantic classifier along with a count of stop words used in the article, which acts as a
by mcbits 8y ago
Arxiv accidentally created a crackpot filter from apparently a simple semantic classifier along with a count of stop words used in the article, which acts as a very coarse "style" metric.[1] Reasonable-looking crackpot work is usually classified as "general physics" instead of being rejected entirely. Unfortunately it also lumps some legitimate but unconventional research in with the crackpots and makes Arxiv itself somewhat of a gatekeeper.
What I'm imagining is a sort of layered approach with raw "article" (or some other unit) data at the bottom and indexing, tagging, clustering, filtering, reviewing, commenting, linking, etc. layered on top with possibly many implementations to choose from.
The analysis/filtering would also apply to people augmenting the data, so if some users are really good at tagging certain types of junk as junk, you could easily filter out that junk. If there emerges a cluster of users who keep tagging certain interesting material as "woo" then you could filter them out or even use them to discover interesting material.
I doubt there's a silver bullet (at least today) that could reliably distinguish between unconventional-bad and unconventional-good work, but keeping the baby and the bathwater together opens the door to such an algorithm in the future.
[1] https://arxiv.org/abs/1603.03824 https://arxiv.org/abs/1603.03824