3 ms·
1. Is the index public ? 2. Any chance for a rss feed search ?
by arromatic 2y ago
1. Is the index public ?
2. Any chance for a rss feed search ?
- marginalia_nu 2y ago1. I'm not sure what you mean. The code is open source[3], but the data is, for logistical reasons, not available. Common Crawl is far more comprehensive though. 2. I've got such plans in the pipe. Not sure when I'll have time to implement it, as I'm in the middle of moving in with my girlfriend this month. Soon-ish. [3] at https://git.marginalia.nu/ https://git.marginalia.nu/ , though still some rough edges to sand down before it's easy to self-host (as easy as hosting a full blown internet search engine gets).
- arromatic 2y agoThanks . What you answered at 1. is what I meant. I was looking for a small web dataset but cc is too big for me process . 1. Do you know any dataset of rss feeds that are not 100s of gbs ? 2. How does your crawler handle malicious site when crawling ?
- marginalia_nu 2y ago1. Here are all RSS feeds known to the search engine as of some point in 2023: https://downloads.marginalia.nu/exports/feeds.csv https://downloads.marginalia.nu/exports/feeds.csv -- it's quite noisy though, a fair number of them are anything but small web. You should be able to fetch them all in a few hours I'd reckon, and have a sample dataset to play with. There's also more data at https://downloads.marginalia.nu/exports/ https://downloads.marginalia.nu/exports/ , e.g. a domain level link graph, if you want to experiment more in this space. 2. It's a constant whac-a-mole to reverse-engineer and prevent search engine spam. Luckily I kinda like the game. It's also helpful that it's a search engine so it's quite possible to use the search engine itself to find the malicious results, by searching for the sorts of topics where they tend to crop up, e.g. e-pharama, prostitution, etc.
- arromatic 2y agoOn 2. I meant malware that could affect your crawling server not spams. And thanks for the data .
- marginalia_nu 2y agoMalware authors typically focus on more common targets, like web browsers. I'm quite possibly the only person doing crawling with the stack I'm on, which means it's not a very appealing target. It also helps that the crawler is written in Java, which is a relatively robust language.
- arromatic 2y agoApologies for too many questions but resources on search engines are scarce . How do I visualize the link graphs or process them ? is there any tool preferably foss . Majestic seem to have one but it's their own .
- marginalia_nu 2y agoI don't know if there's any real good answers. It's hard to visualize a graph of this size, but most graph libraries will at least consume it assuming you have a decent amount of RAM.