8 ms·
So instead of scraping IA once, the AI companies will use residential proxies and each scrape the site themselves, costing the news sites even more money. The o
by f33d5173 8mo ago
So instead of scraping IA once, the AI companies will use residential proxies and each scrape the site themselves, costing the news sites even more money. The only real loser is the common man who doesn't have the resources to scrape the entire web himself.
I've sometimes dreamed of a web where every resource is tied to a hash, which can be rehosted by third parties, making archival transparent. This would also make it trivial to stand up a small website without worrying about it get hug-of-deathed, since others would rehost your content for you. Shame IPFS never went anywhere.
- terminalshort 8mo agoBut don't you have to sign a license agreement that prohibits scraping in order to purchase a subscription that allows you to bypass the paywall?
- fartfeatures 8mo agoIPFS was an attempt at this: https://en.wikipedia.org/wiki/InterPlanetary_File_System https://en.wikipedia.org/wiki/InterPlanetary_File_System
- lukeasch21 8mo agoCoincidentally most of the funding towards IPFS development dried up because the VC money moved onto the very technology enabling these problems...
- Seattle3503 8mo agoIs there a good post-mortem of IPFS out there?
- iririririr 8mo agoWhat do you mean? It is alive and "well". Just extremely slow now that interest waned.
- __MatrixMan__ 8mo agoIt's been several years, but in my experiments it felt plenty fast if I prefetched links at page load time so that they're already local by the time the user actually tries to follow them (sometimes I'd do this out to two hops). I think it "failed" because people expected it to be a replacement transport layer for the existing web, minus all of the problems the existing web had, and what they got was a radically different kind of web that would have to be built more or less from scratch. I always figured it was a matter of the existing web getting bad enough, and then we'd see adoption improve. Maybe that time is near.
- iririririr 8mo agooh I mean slow in terms of adoption and public interest. my bad. i expressed awfully. But you are right on the reason it "failed". People expected web++, with a "killer app", whatever that means. Imagination is dead.
- __MatrixMan__ 8mo agoI'm still working on what I think could be a killer app for it, but progress happens on holidays and vacations and weekends only if I'm lucky, so as you say... it's slow :)
- giantrobot 8mo agoI see the primary issue with IPFS is a significant majority of all web users are on mobile. They can't act as content hosts or routers. In P2P parlance they can only ever act as leeches. Even people with full fledged computers the market is dominated by laptops. These have similar availability issues as phones even if they don't have the same storage or connectivity limitations. Compared to the total number of users on the Internet relatively few have stable always-on machines ready to host P2P content. ISPs do not make it easy or at times possible to poke holes in firewalls to allow for easy hosting on residential connections. This necessitates hole punching which adds non-trivial delays on connections and overall poorer network performance. It's less about imagination being dead but instead limitations of the modern Internet retards momentum of P2P anything.
- Operyl 8mo agoThey already are, I've been dealing with Vietnam and Korea residential proxies destroying my systems for weeks, I'm growing tired. I cannot survive 3500 RPS 24/7.
- raincole 8mo agoEven if the site is archived on IA, AI companies will still do the same.
- CqtGLRGcukpy 8mo agoThe AI companies won't just scrape IA once, they're keeping come back to the same pages and scraping them over and over. Even if nothing has changed. This is from my experience having a personal website. AI companies keep coming back even if everything is the same.
- giancarlostoro 8mo agoWeird, considering IA has most of its content in a way you could rehost it all idk why nobody’s just hosting a IA carbon copy that AI companies can hit endlessly, and then cutting IA a nice little check in the process, but I guess some of the wealthiest AI startups are very frugal about training data? This also goes back to something I said long ago, AI companies are relearning software engineering poorly. I can think of so many ways to speed up AI crawlers, im surprised someone being paid 5x my salary cannot.
- mlnj 8mo agoUnless regulated, there is no incentive for the giants to fund anything.
- cm2187 8mo agoThere is no problem that cannot be solved with creating a bureaucracy and paperwork!
- jniles 8mo agoI understand this is tongue-in-cheek, but do you have an alternative/better proposal?
- cm2187 8mo agoLet the market do. If good data is so critical to the success of AI, AI companies will pay for it. I don't know how someone can still entertain the idea that a bureaucrat, or worse, a politician, is remotely competent at designing an efficient economy.
- demetris 8mo agoI don’t believe resips will be with us for long, at least not to the extent they are now. There is pressure and there are strong commercial interests against the whole thing. I think the problem will solve itself in some part. Also, I always wonder about Common Crawl: Is there is something wrong with it? Is it badly designed? What is it that all the trainers cannot find there so they need to crawl our sites over and over again for the exact same stuff, each on its own?
- ccgreg 8mo agoMany AI projects in academia or research get all of their web data from Common Crawl -- in addition to many not-AI usages of our dataset. The folks who crawl more appear to mostly be folks who are doing grounding or RAG, and also AI companies who think that they can build a better foundational model by going big. We recommend that all of these folks respect robots.txt and rate limits.
- demetris 8mo agoThank you! > The folks who crawl more appear to mostly be folks who are doing grounding or RAG, and also AI companies who think that they can build a better foundational model by going big. But how can they aspire to do any of that if they cannot build a basic bot? My case, which I know is the same for many people: My content is updated infrequently. Common Crawl must have all of it. I do not block Common Crawl, and I see it (the genuine one from the published ranges; not the fakes) visiting frequently. Yet the LLM bots hit the same URLs all the time, multiple times a day. I plan to start blocking more of them, even the User and Search variants. The situation is becoming absurd.
- ccgreg 8mo agoWell, yes, it is a bit distressing that ill behaved crawlers are causing a lot of damage -- and collateral damage, too, when well-behaved bots get blocked.
- toomuchtodo 8mo agoAI browsers will be the scrapers, shipping content back to the mothership for processing and storage as users co browse with the agentic browser.
- pigggg 8mo agoAI companies are _already_ funding and using residential proxies. Guess how much of those proxies are acquired through being compromised or tricking people into installing apps?
- golem14 8mo agoDoes anyone know if Teslas do this? I noticed Tesla cars want to have access to local WiFi and eat up oodles of bandwidth …
- Aurornis 8mo ago> So instead of scraping IA once, the AI companies will use residential proxies and each scrape the site themselves, costing the news sites even more money. News websites aren’t like those labyrinthian cgit hosted websites that get crushed under scrapers. If 1,000 different AI scrapers hit a news website every hour it wouldn’t even make a blip on the traffic logs. Also, AI companies are already scraping these websites directly in their own architecture. It’s how they try to stay relevant and fresh.
- dawnerd 8mo agoHello hi, I work on a news site and we absolutely notice and it does mess up traffic logs.
- shark_laser 8mo ago> I've sometimes dreamed of a web where every resource is tied to a hash, which can be rehosted by third parties, making archival transparent. This would also make it trivial to stand up a small website without worrying about it get hug-of-deathed, since others would rehost your content for you. Shame IPFS never went anywhere. You've just described Nostr: Content that is tied to a hash (so its origin and authenticity can be verified) that is hosted by third parties (or yourself if you want)
- Hendrikto 8mo agoWith Nostr you can host your content anywhere, but for it to actually be discoverable, you need to declare that host. Third parties therefore cannot really solve the problem for you, without your help.
- zaphirplane 8mo agoWe don’t lack the technology to limit scrapers, sure it’s an arms race with AI companies with more money than most. Why can’t this be a legal block through TOS
- lxgr 8mo agoBut hey, paywalled sites might be getting 2-3 additional subscriptions out of it!
- nerdponx 8mo agoIt's almost as if this isn't about scraping and more about shutting down a "free article sharing" channel that gets abused all the time.
- j45 8mo agoBlocking the internet archive sounds like non-tech leadership making decisions without understanding how ubiquitous and moot it is to simply get it another way. Kind of sucks because the news are an important part of that kind of an archive.
- WalterBright 8mo ago> I've sometimes dreamed of a web where every resource is tied to a hash, which can be rehosted by third parties, making archival transparent. I wrote a short paper on that 25 years ago, but it went nowhere. I still think it is a great idea!
- jeron 8mo ago>The only real loser is the common man who doesn't have the resources to scrape the entire web himself. definitely, this is going to hurt those over at /r/datahoarder
- Denatonium 8mo agoIt would be nice if IA could create a browser extension or TLS-intercepting proxy that end users can run over their own computers and connections, allowing crowd-sourced scraping. It would need an allow/deny-listing feature for sites to passively crawl, and I'm not sure how you could prevent data poisoning, but it would at least get around the issues of blocking.