4 ms·
Afaik there is no meta data scrapped (technically I think some HTML meta tags are scrapped but not sure if used). There is machine-learning in the annotations
by lrei 9y ago
Afaik there is no meta data scrapped (technically I think some HTML meta tags are scrapped but not sure if used).
There is machine-learning in the annotations - categories rely on a (cross lingual) text classifier, entities rely on matching to Wikipedia articles, maybe a bunch of small other things to - I don't know all the details - there are some papers published about it.
- deleted 9y ago[deleted]
- oceanbreeze83 9y agointerested in knowing about these papers. can you point me to some?
- pbadenski 9y agoA query to Google Scholar and quick skimming through the papers yields this as interesting start: Event Registry – Learning About World Events From News http://wwwconference.org/proceedings/www2014/companion/p107.pdf http://wwwconference.org/proceedings/www2014/companion/p107.... Using news articles for real-time cross-lingual event detection and filtering http://ai2-s2-pdfs.s3.amazonaws.com/f917/c0cff24fed1af45f94c53b74ca0229874966.pdf http://ai2-s2-pdfs.s3.amazonaws.com/f917/c0cff24fed1af45f94c...
- lrei 9y agoCorrect. Main author of the project is Gregor Leban: https://scholar.google.co.uk/citations?user=5pAxBWsAAAAJ&hl=en https://scholar.google.co.uk/citations?user=5pAxBWsAAAAJ&hl=... The original crawler (newsfeed.ijs.si) paper is from Trampus, Mitja and Novak, Blaz: The Internals Of An Aggregated Web News Feed. Proceedings of 15th Multiconference on Information Society 2012 (IS-2012).