3 ms·
I'm a PhD student, and one of the things you quickly become aware of is the fire-hose of new science that is published daily. The discovery/prioritisation probl
by tompccs 9y ago
I'm a PhD student, and one of the things you quickly become aware of is the fire-hose of new science that is published daily. The discovery/prioritisation problem is being "solved" by publishers recommending other articles (published by them) on their website, or sites like ResearchGate (basically Facebook for academics and researchers) doing the same with similarly obscure algorithms.
I started a side project that pulls in RSS feeds from various academic publishers and uses a simple regression analysis based on tiles and abstracts to recommend you papers. I've become busy with my PhD but if anyone is interested, it's on github:
https://github.com/tompccs/scoopy https://github.com/tompccs/scoopy
- tekkk 9y agoInteresting. I just happened to think of a new side-project for myself using RSS and Scrapy of which I'm familiar with Scrapy. Cool if you can parse RSS-feeds so easily with feedparser-library. Makes it probably quite easy compared to the Scrapy hacks I've done trying to navigate through a messy HTML-structure that has no unique identifiers whatsoever. By the way you can easily put that thing into a Lambda without or with Serverless. I don't know do you need any data to be stored in S3 right now but it shouldn't take long to do it. =) Using machine learning would be for my own project only in the far future but still, data science yo. Single layer neural network with one activation function or what was the great description someone gave here for linear regression.
- twic 9y agoWhen i was a PhD student, i early on set up alerts on Zetoc [1] for tables of contents on the top journals in my area. A couple of weeks later, i wasn't even reading all of the tables of contents, let alone the papers. What might have worked is getting together with a few like-minded students and postdocs and doing it collectively; it would have lowered the individual workload, and created a social pressure to actually do it. It could be a sort of superficial version of a journal club. Machine learning is cool, but in life sciences, it's traditional to solve problems by throwing more postdocs at them. You could perhaps do this with human effort, but mediated (and perhaps assisted) by machine. Imagine a tiny HN instance with a handful of users, where the new page is fed from RSS. Or, at a larger scale, Reddit, with subs for various disciplines, and some process for routing RSS items to the appropriate subs. [1] http://zetoc.jisc.ac.uk/ http://zetoc.jisc.ac.uk/
- Matumio 9y agoFor machine learning there is the Arxiv Sanity Preserver, which can recommend new papers based on your library. http://www.arxiv-sanity.com/ http://www.arxiv-sanity.com/ It didn't really preserve my sanity, though. I'm reading those papers in my spare time, and it got me to read many more papers than I ever did before.
- tompccs 9y agoNot seen that before (my field isn't comp sci or ML), but it looks great! Obviously web interface is the way to go for a serious project, but I actually found it fun learning how to implement an interactive command line app. It's incredibly responsive and has a four-key interface, so if you are good at skim reading you can train it with a lot of papers in very little time.
- jarvist 9y agoA bespoke solution (learns by your 'up votes') for any area of the ArXiv has been developed by Anton Lukyanenko: http://arxivist.com/ http://arxivist.com/
- deleted 9y ago[deleted]