6 ms·
I did skim, but I got the impression it was tested as a NO because of all the irrelevant articles (whole city pages, death of Epstein etc. ) This is the scrapi
by quickthrower2 3y ago
I did skim, but I got the impression it was tested as a NO because of all the irrelevant articles (whole city pages, death of Epstein etc. )
This is the scraping tarpit! And why I avoid scraping projects. Even if there is a nice api it is still like a scraping project as there is tonnes of raw crap to deal with and categorise.
Machine learning might help. Not necessarily a LLM but could just use attention architecture to predict yes/no relevant or irrelevant, maybe.
- benreesman 3y agoIf anyone wants to try something like this I recommend looking at: https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/45573.pdf https://static.googleusercontent.com/media/research.google.c... I ran across it trying to solve a (regime of) a problem in proteomics, but it’s got old-school P(click | impression) style problems written all over it.
- margalabargala 3y ago> there is tonnes of raw crap to deal with and categorise This is a fair use case for LLMs. "Read this article and decide if it meets this criteria". There will be some false positives and some false negatives but it will probably have a 95%+ success rate out of the box, which is more than good enough for a project like this.
- bemmu 3y agoThis works. I recently used this approach to find baby names on Reddit. I would show GPT a list of articles and ask "list the article IDs among these which appear to be about baby names". Then further show each of those articles and ask "what baby names were mentioned in this post and its replies?".
- dangus 3y agoCould it be done with “scroll around on a map” performance? I would think you’d have to process the whole Wikipedia data set ahead of time. I don’t know if that’s difficult or not.
- margalabargala 3y agoI would bet it could be made to be that performant. You wouldn't need full GPT-4 performance to do this sort of sentiment analysis, a cheaper, faster model would work fine. Furthermore, you wouldn't need to ingest the whole article; just the intro paragraph would be fine. Overwhelmingly, the information needed to determine if this is the sort of article suitable for inclusion is in the intro section at the very top of the page; just one or a couple paragraphs. Add to that that every page could be run in parallel, and you're looking at probably less than two seconds of processing time for loading all articles in an arbitrary location. Bonus points that an article's suitability is unlikely to change, so this only needs to be done once, and then the article's id and suitability bool can be stored in a table somewhere.