3 ms·
I used to, I now use zotero to save whole pages onto webdav, from there bunch of scripts peel the ads off, scrape the text, convert to PDF, store in cms and ind
by kusmi 10y ago
I used to, I now use zotero to save whole pages onto webdav, from there bunch of scripts peel the ads off, scrape the text, convert to PDF, store in cms and index for full text search on solr. Also hooked up Dropbox to do the same for one click archiving from mobile. Since Dropbox and the webdav are shared between my partners and I, it's a convenient way to build knowledge base. Experimenting hooking up Telegram and slack as well to integrate everything for no hassle user-end. The real pain in the ass is passing the URL itself, consistently, without insisting users use another third party app.
*Forgot to mention the best part: Backend pools these full-text documents, cleans and parses for NLP, then generates meaningful tags, and organizes documents in an auto generated folder hierarchy which is based on word2vec/doc2vec and content clusters. Whole thing runs on a dedicated server with two 1070 GTX video cards for the NLP work which is training and re-evaluating constantly as new content pours in.
Altogether it was 2-3 years of work.
- x0x0 10y agoThat's a crazy (awesome?) level of effort. May I ask why you went to it? What type of knowledge base are you building?
- kusmi 10y agoI work in molecular genetics research, and I've noticed how disconnected subfields within the life sciences can be. For example, a classical geneticist working with yeast can spend their entire career oblivous to the fact that the problem they've been working on has been indirectly solved by a biochemical engineer working in the same building. This happens for a number of reasons, they publish in different journals, use different terminology, the relationship may not be obvious, etc. Originally, I only had metadata extractors and various NLP parsers specific to the bio/life sciences, but I felt that was too limiting and began to expand it. The backend which ties all the services is almost entirely written in Lua/Torch, and Redis. And everything is built around the Alfresco CMS which comes with Solr, and Mattermost as a locally hosted slack alternative. Mattermost bots report on new content (http://imgur.com/a/P3YK1 http://imgur.com/a/P3YK1, http://imgur.com/a/GTVEX http://imgur.com/a/GTVEX) wherever it comes from. There is too much information to stay on the bleeding-edge of things without serious commitment of resources, which start-ups don't have. My intention was to track the content a group of people go through in a day and visualize connections in the data that may not have been obvious before. Essentially, it's meant for harvesting IP in biotech field.
- narak 9y agoThis is really cool. Do you mind putting your contact info in your "about" box if someone wants to get in touch to learn more?
- ravendug 10y agoSounds awesome! Have you ever considered turning it into a product/service?
- kusmi 10y agoI was considering piecemealing it out as Saas. The CMS component is heavily dependent on Alfresco, which is a bit of a nightmare to work with, to the extent I code around it instead of directly integrating the components. If I find funding, I wouldn't mind splitting this off as its own project (apart from our more wet-lab oriented work which has nothing to do with software). I'd need to payroll a Java developer.
- banku_brougham 10y agoYes. this sounds so awesome i thought it was made up. but i can imagine how this would evolve over two years
- kusmi 10y agoI even had to build a server specifically for this, running it on AWS for example would bankrupt me.
- haffi112 10y agoAre you planning on sharing your setup? This sounds like a great way to organize research in a group at a university.
- kusmi 10y agoOriginally no, it was supposed to be a tool to help us generate IP and analyze patents. I'm considering peddling it to other biotechs on a word-of-mouth basis and see where that goes. To make this into a production ready service would require more effort than just me alone, so I'm still exploring options.