5 ms·
I'd favor something like clustering and client-side filtering over gatekeeping. Mainstream academic research can coexist just fine with fringe theories and indu
by mcbits 8y ago
I'd favor something like clustering and client-side filtering over gatekeeping. Mainstream academic research can coexist just fine with fringe theories and industry research in the same database as long as one can efficiently distinguish between them.
And keeping everything together in one (de)central database would make it possible for people, so inclined, to annotate other people's work with new references to support or debunk the work long after it was published, or to clarify ambiguous language, etc. Those annotations, too, could be subject to filtering as needed.
People could build reputations and whole careers around tying up loose ends instead of the "publish or perish" grind.
- nl 8y agoMainstream academic research can coexist just fine with fringe theories and industry research in the same database as long as one can efficiently distinguish between them How do you propose that is done? For my job I build neural networks for text processing. I spend a lot of time reading papers in the field. And yet if I look at something in an adjacent field (even something as close as something like open information extraction) I have trouble telling which papers are important. How on earth am I supposed to tell if something in a further removed field which attracts more crackpots (say probability theory or something) is a fringe theory or a breakthrough from a new author? I'd note the example of the Gaussian correlation inequality[1] where even people in the field weren't aware it had been proven for 3 years after publication[2]. [1] https://www.quantamagazine.org/statistician-proves-gaussian-correlation-inequality-20170328/ https://www.quantamagazine.org/statistician-proves-gaussian-... [2] https://en.wikipedia.org/wiki/Gaussian_correlation_inequality https://en.wikipedia.org/wiki/Gaussian_correlation_inequalit...
- mcbits 8y agoArxiv accidentally created a crackpot filter from apparently a simple semantic classifier along with a count of stop words used in the article, which acts as a very coarse "style" metric.[1] Reasonable-looking crackpot work is usually classified as "general physics" instead of being rejected entirely. Unfortunately it also lumps some legitimate but unconventional research in with the crackpots and makes Arxiv itself somewhat of a gatekeeper. What I'm imagining is a sort of layered approach with raw "article" (or some other unit) data at the bottom and indexing, tagging, clustering, filtering, reviewing, commenting, linking, etc. layered on top with possibly many implementations to choose from. The analysis/filtering would also apply to people augmenting the data, so if some users are really good at tagging certain types of junk as junk, you could easily filter out that junk. If there emerges a cluster of users who keep tagging certain interesting material as "woo" then you could filter them out or even use them to discover interesting material. I doubt there's a silver bullet (at least today) that could reliably distinguish between unconventional-bad and unconventional-good work, but keeping the baby and the bathwater together opens the door to such an algorithm in the future. [1] https://arxiv.org/abs/1603.03824 https://arxiv.org/abs/1603.03824