4 ms·
It turns out to be a complete pain for working out scope in Web crawlers because you need to separate the domain from the path part and deal with them separatel
by junklight 17y ago
It turns out to be a complete pain for working out scope in Web crawlers because you need to separate the domain from the path part and deal with them separately. If it was in the opposite order it would be much simpler to process.
The Heritrix crawler (primarily worked on by the Internet Archive) introduces a "surt" form which is basically the domain in the same order as the path so that Reading from left to right it goes from least specific to most specific.
- jacquesm 17y agoCrawling is an inconvenience, sure it is problematic, but phishing is much more problematic. You can train a machine to parse that thing right-to-left no problem, to tell users to start middle-to-left and then middle-to-right is too much of a burden.
- junklight 17y agoAgreed. Not so much "a burden" as something that you will never be able to teach a large number of people.