4 ms·
Well that sure is an arrogant reply, and not even close to true. In the US everything published (online or not) is copyrighted. Since you're clearly not a cont
by brwnll 11y ago
Well that sure is an arrogant reply, and not even close to true. In the US everything published (online or not) is copyrighted.
Since you're clearly not a content creator, and have decided since you didn't work on it it has no value, but putting something online does not make content public domain.
- jumperjake 11y agoYou're confusing redistribution with scraping. Scraping publicly-accessible content is legal. Redistributing it is not (at least not in the US). Google scraps webpages as a core competency.
- chinathrow 11y agoGoogle also adheres to the robots.txt standard. Most of the scrapers I block don't.
- jumperjake 11y agoNot correct. Google will completely ignore the rules in robots.txt if it deems it acceptable. I think there's a link to this somewhere in this comment page.
- chinathrow 11y agoThey do not index the content but might add the URL, correct. You can have a meta noindex present and they won't index even the URL.
- kuschku 11y agoPutting it online allows me to scrape it, but not to republish it. I might just want to train a neural network on extracting data out of websites, and, for that usecase, wget -m every website I can find. How I use the data you present to me, as long as I don’t republish it, is my decision. If you give me a license to read it, you also give me the license to copy it onto up to 7 different media at the same time, and to show it to up to 7 friends at the same time.