5 ms·
I've always been curious about how search engines seed their scanning and index programs. Like how do you know what domains, ips, etc.. to start scanning and wh
by cloudyporpoise 3y ago
I've always been curious about how search engines seed their scanning and index programs. Like how do you know what domains, ips, etc.. to start scanning and where is the origin?
- ddorian43 3y agoStart with Common Crawl and go from there.
- djoldman 3y agoThat may be a much easier question to answer than discovery. How do you discover relevant new domains?
- marginalia_nu 3y agoI've actually sort of solved this recently. Marginalia's ranking algorithm is a modified PageRank that instead of links uses website adjacencies[1]. It can rank websites even if they aren't indexed, based on who is linking to them. Vanilla PageRank can't do this very well. Domains that aren't indexed don't have (known) outgoing links, in the periphery of the rank. There's a some tricks to get these to not mess up the algorithm completely, but they basically all rank poorly. That's even without considering all the well known tricks for manipulating vanilla pagerank. The modified version seems very robust with regards to both problems. [1] https://memex.marginalia.nu/log/73-new-approach-to-ranking.gmi https://memex.marginalia.nu/log/73-new-approach-to-ranking.g...
- marginalia_nu 3y agoIt's basically seeded with my personal bookmark list. Like a few dozen links. Not exactly this, but close enough: https://memex.marginalia.nu/links/bookmarks.gmi https://memex.marginalia.nu/links/bookmarks.gmi I've changed the crawler design a couple of times, but the principle for growing the set of sites to be crawled is to look for sites that are (in some sense) adjacent to domains that were found to be good.
- cloudyporpoise 3y agoSo if there was a new domain, unlinked by anything - this wouldn't find it?
- marginalia_nu 3y agoIt wouldn't. But such islands are typically not very interesting either. The context of who links to a domain is very important for a search engine for many tasks, not just discovery.
- cloudyporpoise 3y agoVery cool. Reason I ask is at first glance the header "Search the Internet" to me, implies you are searching the entire internet. It sounds like a more appropriate header would be "Search the obsecure Internet"
- marginalia_nu 3y agoTo be fair, no search engine lets you search the entire Internet, not even Google does this. Internet arguably doesn't even have a size. You can construct a website that's like n.example.com/m which links to '(n+1).example.com/m' and 'n.example.com/(m+1)', for each m and n between 0 and 1e308.
- Lex-2008 3y agoI did it! For every two numbers, calc.shpakovsky.ru has a static(-looking) webpage showing their sum (or difference, etc). Together with links to several other pages. The only limitation I know of is 4k URL length. Interestingly enough, major search engines are rather smart about it and cooled down their indexing efforts after some time. Guess, I'm not the first one to make such a website.
- marginalia_nu 3y ago
- gertgoeman 3y agoI remember reading somewhere that Google used dmoz (https://en.wikipedia.org/wiki/DMOZ https://en.wikipedia.org/wiki/DMOZ) as seed page for their crawler. Not sure if it's true though...