3 ms·
Interesting aspect! The biggest problem in this would probably be how to recognize if "content" changed. A site can change the full design, navigation, footer
by Sujan 8y ago
Interesting aspect!
The biggest problem in this would probably be how to recognize if "content" changed. A site can change the full design, navigation, footer and header and everything and still have the exact same "content". For a human being this will be simple enough to understand, but a tool might have its problems with that.
- Cogito 8y agoYes, this is a fundamental issue if you wanted to do this at scale. There are a few solutions to this already, using solutions like outline.com to pull the content out of the cruft, but I don't know how many of these are general purpose and how many are purpose built for each site (and maintained for the current version of the site, perhaps?) As seen in the article, most links are to a small number of sites, so perhaps hard coding the content extraction would be feasible, especially for an initial study. It would be interesting I think to see just how many links have identical content, but you're right in that the number will be skewed greatly if there are any ads or similar included.