4 ms·
I feel like the solution is a better common crawl. As nice as it would be to block the frontier AI labs from getting access to information, we should reset the
by mips_avatar 3mo ago
I feel like the solution is a better common crawl. As nice as it would be to block the frontier AI labs from getting access to information, we should reset the baseline of information accessibility so there's less marginal advantage on these labs.
I worry a lot of the anti scraping rhetoric will just injure the open web and put somebody like cloudflare in charge.
- jay_kyburz 3mo agoI agree, if up-to-data data was available somewhere else and free, there would be no reason to pay hackers and scrape. You could perhaps even get website operators to "push" new data to a common crawl database. The scrapers would learn there is no value on scraping X domain because the data is available elsewhere more easily.
- jay_kyburz 3mo agoHow about a website header with a link to a static zip that contains the whole website in one hit. The Zip could be hosted on some big public sever. Perhaps even mirrored locally for each nation.
- mips_avatar 3mo agothat's hard to do with rendered content, oftentimes the result depends on a backend service. Maybe you should make the service it's running public but that might be a line most aren't willing to cross.
- jay_kyburz 3mo agoI was thinking you scrape your own website every day in the middle of the night when traffic is low, and make that available. They can come and collect it every day if they want to.
- mips_avatar 3mo agoYeah. Though I guess the point I thought of was like a deals site. That would have infinite pages and content
- Symbiote 3mo agoI have essentially this at work, but the scrapers ignore it. (Or at least many, many scrapers ignore it.)
- jay_kyburz 3mo agoIt wont work unless everybody does it because otherwise it's more work for the scrapers not less. They need to implement two systems instead of one. And we'll never get everybody to do it.
- flaburgan 3mo agoWell this is not what is happening in practice, Wikipedia / Wikidata, OpenStreetMap, OpenFoodFacts... All provide APIs and even a full dump of their database available to download for free, but no, the stupid bots still DDoS them 24h/24.
- pocksuppet 3mo agoWhy don't they take legal action?
- jay_kyburz 3mo agoThere is nobody to sue. The traffic is coming from millions of residential IPs.
- inigyou 3mo agoWhy would that stop them?
- nobodywillobsrv 3mo agoFeels like it would be a good time for freenet and the like to catch on.
- andai 3mo agoWhat really confuses me is ... people always say, it's because companies are gathering data for AI training. Then why would they need to scrape the same page thousands of times per day? Edit: the article says millions of times per hour? (!?) The article is also astonished by this, and speculates it might be some kind of underground AI labs but... millions of them? Or does it only take one with too much money and a badly configured scraping setup?
- mips_avatar 3mo agoI kind of wish the recent Google monopoly court ruling had forced Google to open up their index to anyone, not just Perplexity/other big players.
- catdog 3mo agoThat's really a huge issue right now (to some extent even before the AI hype) that almost everywhere google is effectively the only entity explicitly allowed to scrape.
- cyberax 3mo agoHah. I have a homelab with a couple of sites, including a personal Forgejo installation. Last night my server turned off because it went into thermal protection shutdown. Turns out, my all-in-one cooler has inoperative fans, which I normally never really notice. The passive heat dissipation from the water cooler is more than enough. However, this time they hammered my computer for 12 hours with about 200 requests per _second_ to my Forgejo.
- ronsor 3mo agoIronically in early 2023 a lot of websites went out of their way to block Common Crawl. Unsurprisingly that shifted scraping toward individual actors whereas the previous solution in research was to download CC dumps and process them.
- ccgreg 3mo agoWe aren't sure if that really made a significant difference in Common Crawl's data quality. It does hurt our dataset from a humanities point of view, alas.
- II2II 3mo agoI'm sure there are those who would participate, either because they want their data to be captured by AI labs or as a form of compromise. That said, the approach is flawed. It looks like the people doing the scraping want everything. There are some people who do not want their data to be captured by LLMs. A common crawl would make it easier to those people to opt out, limit what is captured, or to poison the data. (I'm assuming the only way to avoid fragmentation is for the crawl to be done in the open and by consent.) Then there is the question of who would pay for the crawling and hosting. You could try charging for access to the dataset, but that would only encourage others to develop and sell their own dataset (especially since there are likely many who would want their interest in such a dataset to be confidential).
- ccgreg 3mo agoIf you're referring to Common Crawl, which has existed since 2008, indeed your predictions are somewhat accurate. It's easy to opt out or limit what is collected. The crawling itself is inexpensive to us and the hosting is from the AWS Open Dataset Sponsorship Program. And there's no charge for downloading it.
- mips_avatar 3mo agoThanks for making common crawl as good as it is. It’s a really important part of making the Internet better
- ccgreg 3mo agoAppreciate your kind words! Many people have worked at Common Crawl over the years, and it's been a labor of love fueled by positive comments like yours and the large list of PhD theses helped by our public web dataset.