4 ms·
Given that the US courts have definitively ruled web scraping legal, doesn't that present a risk that third-parties could create their own APIs to interface wit
by nugget 6y ago
Given that the US courts have definitively ruled web scraping legal, doesn't that present a risk that third-parties could create their own APIs to interface with Reddit, especially as it relates to fetching content? I think if Reddit tries to lock down its content, that becomes a losing battle, when there are probably much more creative ways to monetize.
- comex 6y agoI think the legal situation is more complicated than you’re suggesting. It’s one thing if you’re scraping without logging in, which is possible on Reddit but of course limits you to read-only access. If you want to participate in discussions, on the other hand, you need an account, which means you’re agreeing to a EULA, which means unauthorized clients can potentially be sued over either breach of contract or contractual interference. Maybe. I believe those legal aspects are still largely untested in this context.
- nugget 6y agoI believe that whatever the user wants to do on their own device is allowed; this is why adblock is legal, regardless of what a website's EULA says, for example. If I had a local software client that automated posting of content to Reddit, it would seem difficult to detect and stop. If detected, would Reddit ban my user account? Doesn't seem like the kind of relationship I would want with my power users. Why not out-compete the third-party providers (who are bubbling up CX research for free) which should be easy if the platform owner itself takes it seriously.
- White_Wolf 6y agoThere is a huge difference between adblock and a scraper. The scraper is automated browsing and data harvesting. This can be quantified and creates extra load. Adblock works with local data after the server did it's part and sent all data to the device. It doesn't create additional load on the server.
- toomuchtodo 6y agoA fair compromise would be a scraper that runs as a browser extension, so as to not create extra load (for example, RECAP [1], which scrapes PACER and ships the viewed case PDFs off for archiving and indexing). Shipping post, link, and comment data somewhere globally accessible using a browser extension while logged in sidesteps Reddit’s walled garden efforts without any load beyond that which they would’ve expended serving typical user requests. Maybe IPFS as a target? [1] https://free.law/recap https://free.law/recap
- White_Wolf 6y agoYes. That would be a great middle ground. As a side note:I am pro-scraping as long as common sense has a say in it and it's not abused. I do scrape one in a blue moon when I want to research a topic in depth.It's great for gathering a lot of information to help learn about a topic fast(er).