4 ms·
Respect to them for not naming names, that's the classy move, but I wouldn't blame them if they did.
by stusmall 2y ago
Respect to them for not naming names, that's the classy move, but I wouldn't blame them if they did.
- ryandrake 2y agoWhy is it "classy" to not name names when a business (which likely holds it self out there as reputable) behaves badly, especially when it behaves in a way that costs you money? Everyone is so vague and coy. These companies are being abusive and reckless. Name and shame!
- WhyCause 2y agoIn the article, they mention that they are working with the crawling company to be reimbursed for the download costs. Naming and shaming the company while you're trying to work with them is a real good way to not get what you want.
- deleted 2y ago[deleted]
- OtherShrezzing 2y agoNaming names isn't really required. The hosts have a $5,000 bandwidth fee, but so do the consumers. There's maybe 10 companies with the financial & compute resources to let a $5,000-per-month-per-website bug run rampant before taking the harvesting service offline. Meta/Google/Whoever may benefit from economies of scale, so they're not seeing the full $5,000 their side, but they're hitting tens of thousands of sites with that crawler.
- energy123 2y agoI thought data egress is much more expensive than downloading it?
- echoangle 2y agoYou know you can hit the data rate they were complaining about by using a residential fiber connection, right? 10 TB per day is about 1 Gigabit continuous if I’m not mistaken. There are probably millions of people that could to this if they wanted to.
- OtherShrezzing 2y agoThere's millions of people who could do that to an individual website. There are remarkably few organisations who could do that simultaneously across the top 100,000 or so sites on the internet, which is how readthedocs has encountered this issue.
- Joel_Mckay 2y agoIt is 100% likely a cloud provider IP range. They are a persistent source of spam email servers, scrapers, and bot probes. The simple reason is the operators quickly dump a host, and the next user is left wondering why their legitimate site is instantly spelunking spam ban lists. It is the degenerative nature of cloud services... and unsurprisingly we end up often banning most parts of Digital Ocean, Amazon, Azure, Google, and Baidu. Have a wonderful day, =3
- immibis 2y agoIt's the degenerative nature of assuming an IP corresponds to a user. They have not corresponded to users for over a decade. I once discovered I'm banned on my mobile phone connection from at least one app which doesn't know that CGNAT exists (a very poor assumption for mobile phone apps in particular). If you must block IPs, do it as a last resort, make it based on some observable behavior, quickly instated when that behavior occurs, and quickly uninstated when it does not.
- Joel_Mckay 2y agoReally depends on the use-case, but yeah the response happens in a proportional manner. We also follow the tit-for-tat forgiveness policy to ensure old bans are given a second chance. Mostly, we want the nuisance to sometimes randomly work, as it wastes more of their time fixing bugs. And note, if a server is compromised and persistently causing a problem... we won't hesitate to black hole an entire country along with the active Tor exit nodes and known proxies lists (the hidden feature in context cookies). Have a great day friend, =3
- croemer 2y agoWhy are you ending all your messages with =3 ?
- Joel_Mckay 2y agohttps://www.jpl.nasa.gov/images/pia22092-arp-142-the-penguin-and-the-egg https://www.jpl.nasa.gov/images/pia22092-arp-142-the-penguin... Don't worry about it friend =3