3 ms·
This list is used to determine what domains and sub-domains "belong" together, in the sense that they are controlled and/or owned by the same entity. For insta
by randomstring 10y ago
This list is used to determine what domains and sub-domains "belong" together, in the sense that they are controlled and/or owned by the same entity. For instance *.google.com is all google, so x.google.com and y.google.com can be trusted to share the same SSL key, safely share javascript (XSS), etc. However x.blogger.com and y.blogger.com are probably two completely separate blogs, people, domains, SSL keys, javascript domains, etc. And you wouldn't want to see x.blogger.com's web pages showing up in search results for y.blogger.com.
Secondly, maintaining this list is a pain. You have hundreds of two letter top level domains, one for each country. Each country with its own NIC in charge of sub-domains. Each NIC with the power to add or delete subdomains (check out https://en.wikipedia.org/wiki/.uk https://en.wikipedia.org/wiki/.uk for just one of hundreds of examples). Some "countries" even sell off their top level domain like .nu . Then you have .us (https://en.wikipedia.org/wiki/.us https://en.wikipedia.org/wiki/.us) that has wild card domains like http://vil.stockbridge.mi.us/ http://vil.stockbridge.mi.us/ where the vil. is fixed and part of the domain and stockbridge.mi is the important domain information. Of course they're always adding more top level domains: .ninja, .wtf, etc (maybe there is a .etc now?). Then you have all the blogging and hosting platforms that use personalized domains for hosting content. Many are listed in the publicsuffix list, but I'm guessing not all!
I ended up writing my own publicsuffix parser in PERL a few years back for the blekko search engine. The main purpose being to be able to group web pages together by site/owner. There is nothing quite like feeding every URL on the internet through your parser to find bugs and corner cases.