3 ms·
Did you see our reply? https://commoncrawl.org/blog/setting-the-record-straight-common-crawls-commitment-to-transparency-fair-use-and-the-public-good https://co
by ccgreg 8mo ago
Did you see our reply? https://commoncrawl.org/blog/setting-the-record-straight-common-crawls-commitment-to-transparency-fair-use-and-the-public-good https://commoncrawl.org/blog/setting-the-record-straight-com...
Also, if your site has CC-BY-NC-SA markings, we have preserved them.
- username223 8mo agoHopefully my site is no longer part of Common Crawl. I'm not interested in participating in your project, block CCBot in robots.txt, and have requested deletion of my data via your form.
- ccgreg 8mo agoDid you see our reply? Edit: by which I mean, we sent you an email that explains what we did and how to verify it. Did you not receive an email reply? If not, please contact us again. Also, if your site has CC-BY-NC-SA markings, we have preserved them.
- username223 8mo agoI don't care. Is blocking your bot and requesting removal sufficient? If not, what is?
- ccgreg 8mo agoPlease read our email reply. I have no idea if we received your request —- your HN username doesn’t match any request we have received.
- username223 8mo ago"We have initiated the process to remove your content from the Common Crawl Dataset. This is a multi-step process, involving first a nocrawl directive, followed by removal of the URLs from the primary index files, and finally removal of the content from the deep archive. We will advise when the process is complete." Received April 2024. I have not been advised. Please advise.