5 ms·
Reddit's Robots.txt Changed
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- skilled 2y agoEvery search engine other than Google has stopped indexing pages from Reddit. Google has not commented on whether they plan to respect it. Rich Results[0] say they're using a version from June 25. The new version was last modified July 1. [0]: https://search.google.com/test/rich-results https://search.google.com/test/rich-results
- fragmede 2y agoGoogle is paying Reddit $60 million to not have to respect it. The media reports it's for training data for AI, but none of us have read that contract.
- skilled 2y agoYes, it makes sense from a 600 million monthly visitors point of view, but that would still be a pretty interesting development if they don't respect it.
- fragmede 2y agoif they're getting the data directly from Reddit then they can respect it while still having the content because they're not crawling it via http.
- darreninthenet 2y agoIsn't this monopoly abuse?
- fragmede 2y agoany other company, specifically Microsoft/Bing and OpenAI, are free to pay reddit for the same privilege, so I don't see how it could be considered that.
- darreninthenet 2y agoI actually meant from Reddit as they hold a monopoly on online forums imo, but I'm no expert in this, was just wondering
- altdataseller 2y agoBing is still indexing it. As is DDG. https://www.bing.com/search?q=reddit+travel&qs=n&form=QBRE&sp=-1&ghc=1&lq=0&pq=reddit+travel&sc=9-13&sk=&cvid=C9E48332A62E44E29BCA9E364B91B635&ghsh=0&ghacc=0&ghpl= https://www.bing.com/search?q=reddit+travel&qs=n&form=QBRE&s...
- skilled 2y agoThey are not indexing new pages. I should have made that more clear.
- altdataseller 2y agoThey can and will probably respect it and still be able to have reddit results. “User-agent: * Disallow: /“ in the `robots.txt` simply instructs all web crawlers not to crawl any pages on the website, not whether to show pages in search results. Of course new pages wont be shown if they cant be crawled if theres no other way you can get it.. But that line simply prevents crawlers that abide by robots.txt from retrieving the content of any pages. But Google doesnt need to crawl Reddit anymore. Reddit is directly giving the data to them to serve to users! Bing, DDG and any other search engines now are basically forced to start paying Reddit. And presumably, the Reddit execs have calculated that future revenue will be more because of this. Don’t be fooled: This isnt a “good for internet, morality” decision. Reddit is a public company now and has shareholders. They are doing this because it will equate to more $$$$
- dageshi 2y agoYup, this is what I thought would happen. It wouldn't surprise me if reddit goes a step further and requires login to view pages. AI is the death-knell of the web as we've known it for the past three decades. Once freely available information will retreat behind login walls and charge bigco's for access to train their models. I wonder if some standardised data API will be settled upon, perhaps it already exists?
- deleted 2y ago[deleted]
- fragmede 2y agothat's not ai but the result of the linked in scraper lawsuit. if it's not behind a login, it's fair game to be scraped. Behind a login, the site operator doesn't have to let you scrape them.
- alecco 2y agoThe rugpull is the community giving content for free to these corporations. Perhaps for the next big public forum we should require better licensing. Or just avoid these VC-backed sites altogether and go back to something like Usenet.
- CobrastanJorji 2y agoInteresting. There have been all sorts of "Reddit's content is key to search results, adding 'reddit' to search results makes them good" stories. And there's been a lot of talk about how some of the big ML makers, notably Google, depend on Reddit's content to train their AI. And Google has that recent $60 million deal for content. So clearly Reddit's execs have been talking about how their content is valuable and they shouldn't give it away for free. But at the same time, blocking search engines from indexing your social media site is a dangerous game. Any search engine that respects this is gonna effectively de-list Reddit. That's no good for views, and views is what makes Reddit money. Presumably they have negotiated private deals with Google and probably Microsoft for this and are trying to sell their data to ML companies, because otherwise this would seem suicidal. Kind of a shame. The information is still going to get shared around to all the giant corporations, but Reddit will presumably make it harder to access for all the little guys. And the more they tie the content to dollars, the more managers on the inside will start doing stupid things to try and generate more of whatever the most valuable kinds of content are.
- deleted 2y ago[deleted]
- swarnie 2y ago> So clearly Reddit's execs have been talking about how their content is valuable and they shouldn't give it away for free. Yeah see the API changes from last year and their pre-IPO investor roadshow. This ground is well trodden.
- elaus 2y agoIn my personal experience, Reddit has become mainstream in the past few months/years. Even my completely non-tech friends use it, having not even heard about it 1-2 years ago. So maybe they can afford to be delisted from search engines, because the site is addictive enough to bring back its users – and new ones because it's mentioned everywhere. Like Discord or Facebook, where most content is also unavailable to search engines (although they probably rely much more on the social aspects).
- 2y ago
- benreesman 2y agoIf anyone is curious how deeply destructive, or how deeply approved by the NSA, or how deeply self-sabotaging for society the modern AI training data pipeline is becoming I’d refer them to SB 1047. OpenAI is openly collaborating with the NSA, Google is manipulating the definition of a web crawl, Anthropic has installed a bunch of humanitarians from Jump Trading as the leading mech interp group that makes strident claims about how all this stuff works based on weights you do not and never will have access to. They’re telling you: “And you will do nothing, because you can do nothing.” I invite you to join me in proving that we can in fact do something.
- aydyn 2y agoDo what? Keep in mind youre outnumbered by normies who just do not care by a million to one.
- benreesman 2y agoThe public is distressingly indifferent to both the handling of their data and their pricing power relative to the modern cartel. The enterprise is keenly aware that they will be decades recovering from how badly they got put over a barrel on cloud. The enterprise cares deeply about where their privileged and in many cases legally regulated data goes and is stored. Big business needs an ally in the tech game, and for once that is looking likely to trickle down to John Q Taxpayer.
- alecco 2y ago"Never underestimate a horde of basement-dwelling autists."
- benreesman 2y agoAgreed. None of the modern AI shops has anything like the level of elite technology expertise that the serious Unix companies did in the 80s or 90s: the Sun Microsystems people would eat OpenAI and still be hungry. And they got their ass kicked by Linux. Not even a close, hard-fought match. Brutally humiliated and driven out of business and sold off for parts. The reckless contempt that the modern AI cartel holds open models in, and GGUF in, and even Apple in on this stuff is staggering and makes me wonder if any of the leadership even knows the history. No one has ever tried to tell the technology community what they can and can’t do on their own machine and lived to tell the story. The people pushing DRM over the years need a restraining order against their own employees. Bill Gates is like, a thousand times smarter and more ruthless than any of these people and he backed off in his prime.
- deleted 2y ago[deleted]
- Seattle3503 2y agoTechnically this title violates HNs title policy as it should just be "Reddits robot.txt" or something, but "Reddits robot.txt changed" is more useful. I'm curious to see if mods change it.
- deleted 2y ago[deleted]
- Lorin 2y agoOh great now we can rely on Reddit's own renowned search functionality /s What are they thinking?
- Lutzb 2y agoAlso my first thought. Reddit search is woefully inadequate for its own content.
- Ekaros 2y agoWhy would you even search? You are supposed to browse exactly what is shown to you. Or in worst case ask the same question again.
- alecco 2y agoWith their new Google contract they'll get a search upgrade, for sure. I wouldn't be surprised if they get acquired altogether at some point.
- sunaookami 2y agoRelated thread for the official blog post: https://news.ycombinator.com/item?id=40799275 https://news.ycombinator.com/item?id=40799275 Side note: They seem to serve other robots.txt for different User-Agents & IPs: https://merj.com/blog/investigating-reddits-robots-txt-cloaking-strategy https://merj.com/blog/investigating-reddits-robots-txt-cloak...
- hamilyon2 2y agoIs it a good time to start competitor? Given that reddit might take Quora's path to oblivion.
- daghamm 2y agoThere are multiple alternatives, but very few user seem to want to switch https://www.reddit.com/r/RedditAlternatives/wiki/index/ https://www.reddit.com/r/RedditAlternatives/wiki/index/ I think a moderated mirror of certain subs would be a better solution. Because people like freedom but love their cats and memes much more.
- shitlord 2y agoThat wiki page is a little outdated. The pinned thread is a little more recent but still missing a lot of offshoots like TheMotte.org. Ironically the only off-site that ever really took off is the snuff film site which has 2.3M users now.
- eps 2y agoHave you figured out how you are foing to moderate communities and the content? Because that is Reddit's magic ingredient and you will have to do it 24/7 from day one.
- alecco 2y agoIt's even worse. Many subreddits were taken over by bad actors. When they saw a popular community, they found at what times the mods were inactive (e.g. late night or weekends) and posted evil stuff. Then went to cry to the admins about the evil stuff staying up for hours and got their friends to become mods of the community. In the case of reddit, many people suspect the admins and 3 letter agencies were in the know, or at least played along to their advantage. And if you self-host to avoid the risk of handing over your community to those sites under control, you get flooded with DDoS and hacking attacks (including state-sponsored actors). Sad times.
- eps 2y agoThis is not good. 80% of my Google searches for other people's opinions now end with "site:reddit.com", and there is surprisingly quite a few of them. The alternative is Reddit's own search and it tends to produce less relevant results.
- daft_pink 2y agoKagi, you can thank me later
- aaron695 2y ago[dead]
- nubinetwork 2y agoMost robots don't honour robots.txt anyways...
- rany_ 2y agohttps://www.reddit.com/robots.txt https://www.reddit.com/robots.txt has additional comments: # Welcome to Reddit's robots.txt # Reddit believes in an open internet, but not the misuse of public content. # See https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy Reddit's Public Content Policy for access and use restrictions to Reddit content. # See https://www.reddit.com/r/reddit4researchers/ for details on how Reddit continues to support research and non-commercial use. # policy: https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy User-agent: * Disallow: /
- aydyn 2y ago"Public content" aka _our_ content that we sell for millions and steal from our userbase.
- jkhanlar 2y agolol what? Just 20-30 minutes ago I saw this at #52, and tried to find it again now and see it #468 https://archive.ph/ReTR5 https://archive.ph/ReTR5 but I wonder if that is algorithmically natural or whatnot, lol
- jkhanlar 2y agoAlso it seems that since 2018 it has not actually changed, lol http://web.archive.org/web/20180501000000*/https://old.reddit.com/robots.txt http://web.archive.org/web/20180501000000*/https://old.reddi...
- deleted 2y ago[deleted]