5 ms·
As somebody who did plenty of scraping for his own little projects, lets just say that reddit their security concerns are just PR for "we do not want to keep su
by benjiro29 3mo ago
As somebody who did plenty of scraping for his own little projects, lets just say that reddit their security concerns are just PR for "we do not want to keep supporting old.reddit".
While yes, you can not simply scrap new reddit as easily as a pure html scraper is cheaper to run. And while the new reddit its js "slow down" scraping as it need to run over a headless browser. Do you slow scraping down a lot?
No. Because you can simply spin up more instances and route over more proxies. Given the fact that you can scrap millions of pages per day, on a single cheap 1 CPU node...
The limiting factor for scraping a website is not html vs js / client vs headless browser. Its avoiding detection by running from different ips, faking browser information and masking your traffic to not give off the smell of automatization (and avoiding tls fingerprinting).
I wrote this type of stuff before we had AI to help write it, in a few days time. Now with a LLM at your fingertips in a few hours, you will have a complete working version with cloaking build in.
And the larger size difference does not matter, you can block irrelevant files via faked cache responds, to give the server the impression that your a repeat client, and no need to flood your browser instance with images, fonts etc. Some think that this is a defense, aka throttling you with increased bandwidth usage lol
So in my personal opinion, its not about dealing with scrapers, but trying to push people to new reddit, so they can finally phase out old reddit.
PS: So much fun reading new reddit on your phone browser, to have it pop-up a "use our app" banner, that prevents you from interacting with the website anymore. And even blocks firefox its scroll ability (and thus accessing to the browser its nav bar). Tip for people: delete the sites cookies and cache regularly, to prevent this usage tracking, so the banner newer shows.
- red-iron-pine 3mo ago> "we do not want to keep supporting old.reddit" the second old.reddit goes is the second I stop using that site it's already very obviously chock full of slop and AI bullshit, with some subs having long unavoidable lists of "your post was removed because you did not verify" trails that imply that automated posters found a thread and went hard but add in privacy crushing bs -- no reason to be there anymore. Back to old school forums
- avadodin 3mo agoOld–school forums never left, still inhabited by the few boomers and genXers who could figure out the Internets back in the day and aren't dead yet. The gamification of Reddit and the ephemeral nature of the main page along with subreddits with varying amounts of overlap contributed to more active interactions, though. Some of the volume was lower quality but not all. The boomer posts about his petunia breeding program on a BBS and he gets feedback half a year later after the other active user gathers the energy to travel to the computer room again. On Reddit, maybe a zoomer would have asked if you can smoke petunias in Brazilian Portuguese, but that may still count as welcome human interaction to the old man. None of that applies to new deaddit, though.
- calvinmorrison 3mo agothe big click of reddit like yahoo forums and other earlier things is that each 'sub' reddit can have its own rules culture, etc. its like facebook groups in a way. Unfortunately with how bad reddit is, I see mostly my interest niche hobbies moving to facebook.
- nothercastle 3mo agoForms are sadly dead. They cost a lot to operate and maintain security updates for so without a critical mass of users they are dead
- samplatt 3mo agoJust like with VR, you can keep on claiming forums are "dead" all you like. Me and all my friends will keep on using them daily. therearedozensofus.gif
- soperj 3mo ago"Good thing you're the majority of people, then." - samplatt [0] [0]: https://news.ycombinator.com/item?id=49002685 https://news.ycombinator.com/item?id=49002685
- kmarc 3mo ago> So much fun reading new reddit on your phone browser, to have it pop-up a "use our app" banner, that prevents you from interacting with the website anymore Hmmm, I remember this from the past, but one of my adblock rules probably catches this one, I haven't experienced it for years.
- apples_oranges 3mo agoI always surf in anon windows that forget all cookies. Works since 2010 without any plugins. :)
- rnd0 3mo agoWorks a treat, I reckon -until the moment comes when you actually want to interact and/or post.
- slifin 3mo agoI can't use new reddit - old reddit lets me disable infinite scroll I think all major platforms should have an option to disable infinite scroll Otherwise I look up and its 3 hours later
- 0123456789ABCDE 3mo ago> the purpose of the system is what is does
- wongarsu 3mo agoNew reddit is more annoying (to me as human, but also to scrapers) in two very important ways: - old reddit shows more or less the raw comment tree, sorted, like HN. New reddit does aggressive folding. You get a couple comments for each branch and have to click "show more" to see more - new reddit forces you to log in to view anything tagged NSFW, including entire subreddits. Old reddit does not enforce this, along with not enforcing some other restrictions reddit has introduced over time Of course none of this stops a scraper. And there is still the API. But these things make it easier to detect scrapers because they have to make a lot more requests and have to handle user sessions that make them easier to track
- tlonny 3mo agoAs I understand it, anything publicly viewable is free to be scraped without issue. If you scrape stuff behind a sign-up, you potentially open yourself up to being sued for breach of ToS. Now I'm sure the Chinese labs don't give a rats arse about that threat, but I imagine it stops the likes of Google, OAI, Anthropic from scraping and instead forces them to purchase the data.
- benjiro29 3mo agoNot how it works ... Anything public viewable is still subject to a TOS. And the whole copyright or whatever still applies. If your argument was valid, we can scrape news websites and show their content freely. No ... You get sued and lose the case. See Google News that got their behinds in court and lost. Short summaries are allowed / transformative, but you can not just take content (even without a login). Not without getting into civil court if somebody wants to press the matter. Now, if you scrap and never make that data public or transform it (LLMs). Then it becomes a harder matter to deal with.