30 ms·
Google is the only search engine that works on Reddit now, thanks to AI deal
- VoidWhisperer 2y agoWow, reddit found a way to make themselves even less useful somehow. After the API fiasco, that seemed like it'd be pretty hard to do.
- wvenable 2y agoBut, apparently, they did finally find a way to make money.
- LunaSea 2y agoBarely enough to pay the CEO
- jasode 2y ago>But, apparently, they did finally find a way to make money. The most recent 10-K financial results 2024-03-31 (filed 2024-05-08) shows they actually lost money: https://www.sec.gov/edgar/browse/?CIK=1713445 https://www.sec.gov/edgar/browse/?CIK=1713445 (For 2024-Q1, Reddit lost -$575 million on revenue of $242 M.) If the quoted "$60 million deal"[1] from Feb 2024 is accurate, that small amount from Google may not be enough for Reddit to turn a profit. It remains to be seen what the Q2 or Q3 financials will show. [1] https://www.google.com/search?q=google+ai+deal+reddit https://www.google.com/search?q=google+ai+deal+reddit
- wuiheerfoj 2y agoWow, perhaps I’m naive but what the hell are they spending over $800M a year on? That seems an obscene amount for a glorified message board. I just read they have 2000 employees which is also puzzling to me
- toomuchtodo 2y agoThey were a public good currently larping as a for profit concern now run by a vanity and wealth driven executive driving it into the ground while it flails to monetize when that is likely incompatible with the entity. Compare and contrast to say, HN, run on two servers in a colo with less than a handful of mods.
- splwjs 2y agoit's not just a message board, it's an influence machine. They need to make sure the stuff they want people to think is posted often and has a big number next to it, they need to make sure things that people like are associated with the stuff they want people to like/think/do and things that people don't like are never associated with the stuff they want people to like/think/do. They need to make sure that people who say the wrong things are silenced or persuaded to leave, etc etc. Man they probably have at least one contact in at least one intelligence agency and they have to make sure not to run afoul of that contact. Like the news isn't just a list of what happened recently, political debates aren't just two guys talking, and reddit/twitter aren't just message boards.
- alephxyz 2y agoThey spent 400M on R&D this quarter, which means more "personalisation"/ad targeting and probably cooking up some DOA chatbot/assistant product that's costing them a ton in compute
- some_random 2y agoAlmost 200 million is in CEO compensation https://www.statista.com/statistics/1453196/reddit-top-executives-compensations/ https://www.statista.com/statistics/1453196/reddit-top-execu...
- Hikikomori 2y agoThe only things it does for me is forcing me to use Google as a large amount of the answers I need is on reddit.
- immibis 2y agoThat's what Google is paying them for :)
- brewdad 2y agoSo then this gambit worked. It sucks and I hate it. I will continue to use DDG/Bing first but it looks like I'll be hitting up Google more often too.
- WarOnPrivacy 2y ago> The only things it does for me is forcing me to use Google Startpage, Kagi and Lukol are 3 that source from Google. I imagine there are others.
- stainablesteel 2y agowhich is ironic because pre-AI every solid piece of obscure information and non-programming question usually had an answer on reddit, its an extremely valuable dataset looking back. but moving forward i think its only going to become less valuable and people will probably manually/custom-scrape all the questions out of worthwhile subreddits and open up their data for free
- splwjs 2y agoWhen I was young, my brother knew a guy who was really into movies. If you wanted to know about a movie you couldn't remember, you would go talk to that guy. For a while, the internet had an end-run play that made that guy less useful. You can just go on the internet for obscure movie information, buddy. But now it seems like knowing a movie guy is going to be the only way to get a real person's opinion on movies. The internet is about to forget everything without a profit motive and just start telling you that the latest product from a monolith corp like disney is the only movie worth watching. If someone scrapes all the useful movie opinions off of reddit and spends their time crafting it into a usable format, that guy's probably got a company. But not Bill. Bill's just a guy you can know or not know. You can't monetize knowing Bill. Sidenote that's probably why it irked me so bad when some bozo coined the phrase "social capital".
- splwjs 2y agoIf they kept their API open then by now the entirety of the site would be ai slop that was built with chatgpt and launched with the api. Then again most of what that site does is just blend and regurgitate the information that's currently on it anyway.
- miohtama 2y agoThose AI bots would likely to be more intelligent commentors than Redditors
- abdullahkhalids 2y agoThe API changes and these robots.txt were part of the same strategy - preventing third parties from scrapping their data and reducing the AI generated content that makes it into their data. So they can sell that data and make money.
- kjkjadksj 2y agoTheir dataset is already polluted with misinformation campaigns and shilling
- AlexandrB 2y ago> their data Love how it's their data when it might make them money but not their data if they get sued.
- abdullahkhalids 2y agoThat's fair. I agree with you that in some sense it is user data. And that Reddit is operating unethically.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- popcalc 2y ago# Welcome to Reddit's robots.txt # Reddit believes in an open internet, but not the misuse of public content. # See https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy Reddit's Public Content Policy for access and use restrictions to Reddit content. # See https://www.reddit.com/r/reddit4researchers/ for details on how Reddit continues to support research and non-commercial use. # policy: https://support.reddithelp.com/hc/en-us/articles/26410290525844-Public-Content-Policy User-agent: * Disallow: / Source: https://www.reddit.com/robots.txt https://www.reddit.com/robots.txt
- will0 2y agoLooks like it changed a month ago: https://old.reddit.com/r/redditdev/comments/1doc3pt/updating_our_robotstxt_file_and_upholding_our/ https://old.reddit.com/r/redditdev/comments/1doc3pt/updating...
- immibis 2y agoNobody who wants to be successful obeys robots.txt. And I do mean nobody.
- chippiewill 2y agoThey changed it to disallow so that scrapers can't just claim the robots.txt gave them permission.
- toomuchtodo 2y agoIndependent scrapers can launder the data between Reddit and AI consumers. The only folks this hurts is users seeking info via search engines and folks willing to kowtow to rules that are potentially low impact to evade. Next steps would be (from an adversarial perspective) browser extensions that stream back data for ingestion similar to Recap for Pacer [1]. [1] https://free.law/recap/faq https://free.law/recap/faq (full disclosure: assisting someone pursuing regulatory action against reddit in the EU for a separate issue from scraping, it's a valuable resource, but the folks who own and control it are meh)
- Elfener 2y agoI mean, the reddit company did go public, so things like this were inevitable. Also things like the API fiasco, and also small annoyances like the fact that when you click on an image on reddit, it now goes to a wrapper html page instead of just the actual image (this was one of the reasons reddit was better than most social media...).
- mrec 2y agoMaybe it's just me or something temporary (I use Old Reddit, like all right-thinking folk) but for the past couple of days the image wrapper page seems to have been sent to the glue factory. I'm just getting the image now, unadorned.
- account42 2y agoIt used to be that Reddit didn't host images and you'd have to link (often shitty) external image hosts. The someone created imgur to host images for reddit. And slowly but surely imgur became just another shitty image host (and social media site for some reason). Then reddit wanted some of the dough imgur was pulling in (probably just making losses) and added their own image hosting. At first it worked just like you are saying with you getting direct links to the image file. Now they also turned into yet another shitty image host. Part of the blame for the redirect-to-wrapper page lies with browsers. If browsers didn't let servers reliably differentiated between a direct request and an <img> embed then this practice would not be as widespread.
- onetokeoverthe 2y ago[dead]
- yjkk 2y ago[flagged]
- lifestyleguru 2y agoI deeply regret every minute spent on and kilobyte of text contributed to reddit.
- Ylpertnodi 2y agoI don't. There's nothing around that is similar...with the same traction. The various 'verses are variations on cat pics. I'm still looking, though.
- wccrawford 2y agoWhile it's still not Reddit, but I've been enjoying Lemmy. I have a similar range of communities on each, and other than some annoying groupthink, the content is often similar. And to me, forgetting to log in to each of them feels similar, too. For what that's worth. (I hate both of them when not logged in.)
- trallnag 2y agoI can confidently state that I'm a net negative for Reddit, looking at the dozens of banned accounts in the trash bin of my KeePass vault
- card_zero 2y agoI mostly contributed to r/nonsense and I'm pleased by the thought of that sub's content being used to train future AI, with information about the architectural uses of super-tall chef's hats, the prehistoric invasion of Europe by Beak People, and so forth.
- nerfbatplz 2y agoI propose we change the term enshitification to engoogleification in regards to the internet.
- crazygringo 2y agoThis is about Reddit disallowing other search engines. Blame Reddit, not Google.
- dvngnt_ 2y agoplenty of blame to go around
- crazygringo 2y agoYou'll have to demonstrate that. Is Google's contract with Reddit exclusive, so that other search engines aren't given the opportunity to also pay? I highly doubt that, especially since the DOJ would go after that immediately because of antitrust. So no, pretty sure the blame here is 100% on Reddit unless you have evidence otherwise.
- dvngnt_ 2y agoI don't think the DOJ acts immediately. > so that other search engines aren't given the opportunity to also pay? this makes it harder for new engines if google has exclusive deals with some of the most popular sites
- crazygringo 2y ago> if google has exclusive deals My comment said, show me that the Google deal with Reddit is exclusive. You haven't done that. And there's no reason to think it would be, because of antitrust. The DOJ doesn't have to act "immediately", the point is that obvious antitrust violations come with fines that make it unprofitable to attempt in the first place. And this would be black-and-white obvious antitrust violation, given Google's monopoly status in search. This isn't a gray area where it might be worth it for Google to roll the dice.
- debacle 2y agoReddit has been ripe for disruption for years. It's just waiting on an inflection point and someone to take it behind the barn.
- onlyrealcuzzo 2y agoOr for Google to buy it. They could monetize it much better while being less annoying. Ultimately - Google is getting everything they want from Reddit with this deal without having to buy it outright. Short of Reddit transforming to an entirely different product (difficult) - I'm not sure where the major growth opportunity is for it.
- rob74 2y agoIt wouldn't be the first time they have done something like this either. Remember https://en.wikipedia.org/wiki/Google_Groups https://en.wikipedia.org/wiki/Google_Groups ?
- Suppafly 2y ago>Remember https://en.wikipedia.org/wiki/Google_Groups https://en.wikipedia.org/wiki/Google_Groups It'd be somewhat hilarious if google bought reddit just to archive it and shut it down.
- jessriedel 2y agoVery few of the reddit users who are providing the content for free are motivated by which search engines are allowed to index the content, so I don't see how this would make it more ripe for competition. (If you just mean society would now be even better off if reddit were disrupted, ok, maybe, but that's a different thing.)
- crazygringo 2y agoThe network effects are too strong. Remember, the only reason Reddit "won" was because Digg destroyed itself with a radical upgrade that everyone hated. Reddit would have to do something similarly self-inflicted, and I can't even guess where people would go. Reddit was already an alternative to Digg -- what's the alternative to Reddit? I mean, it's certainly not Quora.
- causal 2y agoIt feels like Reddit is approaching an inflection point anyway where bot-made content is concentrated enough to spoil the whole experience. Closed servers like Discord and Slack may be the last haven of online human interaction.
- deleted 2y ago[deleted]
- onlyrealcuzzo 2y agoThis is an interesting development. How many other sites might have leverage to charge to be indexed? I don't want to live in a world where you have to use X search engine to get answers from Y site - but this seems like the beginning of that world. From an efficiency perspective - it's obviously better for websites to just lease their data to search engines then both sides paying tons of bandwidth and compute to get that data onto search engines. Realistically, there are only 2 search engines now. This seems very bad for Kagi - but possibly could lead the old, cool, hobbiest & un-monetized web being reinvented?
- ColinHayhurst 2y agoKagi uses at least Google and Mojeek edit: > Realistically, there are only 2 search engines now. https://seirdy.one/posts/2021/03/10/search-engines-with-own-indexes/ https://seirdy.one/posts/2021/03/10/search-engines-with-own-...
- WarOnPrivacy 2y ago> Realistically, there are only 2 search engines now. From the article: Many alternatives to GBY [Google, Bing, and Yandex] exist, but almost none of them have their own results; This seems to assert that ~0 other search providers do any crawling at all. Ever. Are we sure that's the case? (they could crawl but never ever return those results == more odd).
- ColinHayhurst 2y agoIt's a very long article so understandable that you did not read on and learn about other search engines crawling beyond GBY. Still there are indeed very few that are crawling at web scale, and internationally. We are at 8 billion pages and totally independent [0], hence expressing our concerns to 404 media after being blanked by Reddit. [0] https://www.mojeek.com/about/why-mojeek https://www.mojeek.com/about/why-mojeek
- WarOnPrivacy 2y ago
- dvngnt_ 2y agosite:reddit.com works for kagi for new posts this week?
- AndroidKitKat 2y agoKagi gets part of their index from Google, per the article, so perhaps that's the reason Kagi still works. Wonder if Vlad and Kagi will do (or have done) the calculus to see if buying crawlability from Reddit itself is cheaper than buying results from Google for Reddit search.
- hugh_kagi 2y agoNot yet but it's something we want to look into.
- rozab 2y agoBasically all 'independent' search engines piggyback off Google or Bing https://help.kagi.com/kagi/search-details/search-sources.html https://help.kagi.com/kagi/search-details/search-sources.htm... >Our search results also include anonymized API calls to all major search result providers worldwide
- ColinHayhurst 2y ago>Basically all 'independent' search engines piggyback off Google or Bing Incorrect: https://www.mojeek.com/about/why-mojeek https://www.mojeek.com/about/why-mojeek
- lpod 2y ago[flagged]
- jedberg 2y agoThey changed robots.txt a month or so ago. For the first 19 years of life, reddit had a very permissive robots.txt. We allowed all by default and then only restricted certain poorly behaved agents (and Bender's Shiny Metal Ass(tm)) But I can understand why they made the change they did. The data was being abused. My guess is that this was an oversight -- that they will do an audit and reopen it for search engines after those engines agree not to use the data for training, because let's face it, reddit is a for profit business and they have to protect their income streams.
- JohnMakin 2y agoOne (in this case, 2) company's incentive for profit should not take priority over the usability/well being of the internet as a whole, ever, and is exactly why we are where we are now. This is an absolutely terrible precedent.
- jedberg 2y agoI agree with you in theory, but in practice someone has to pay for all this magic.
- JohnMakin 2y agoThis is a false dichotomy. You can have services, and not have them devolve into complete unusability in the name of profit. This isn’t sustainable either. The myopic pursuit of short term gains at the expense of the product will collapse at some point in the future, no matter how much you believe in this weird frog-boil internet we’ve inherited now.
- talldayo 2y ago> The myopic pursuit of short term gains at the expense of the product will collapse at some point in the future, The myopic pursuit of short-term gains is the only playbook that works. Long-term business strategy is a gamble, and today's businesses have all learned that they'd rather make hay when the sun is shining than be remembered as a good business. Twitter tried a long-term playbook to reverse their unprofitable sinkhole of a website. That ended up with them being undervalued and sold to the highest bidder.
- PaulRobinson 2y agoThis is great. It means I won't see Reddit content popping up all over search results in other engines. Can Medium do the same? And perhaps Quora?
- lfkdev 2y agoYeah awesome, reddit was one of the last useful results beside the spam blogs and ai generated articles.
- bdjsiqoocwk 2y agoWhat a weird thing to say. Reddit has for a long time been a place where real people hang out and have real conversations, unlike quora and medium.
- MattPalmer1086 2y agoIts not strange to me. Every single time I've followed a Reddit link from search results, I've got a short and fairly useless conversation that doesn't help me at all. So I have never understood why people like it. Obviously, people do see value in it, or they wouldn't keep saying so! I would happily exclude Reddit links from search results though.
- candiddevmike 2y agoI think Reddit lost that kind of authenticity a while ago. Advertisers know the "search:reddit.com <product>" trick, and when you look at the number of upvotes, it costs _pennies_ to get your product trending in the comments.
- Suppafly 2y agoI don't search reddit for <product> though I search it for <highly technical issue with product> because reddit is the only place where real people discuss such issues and the solutions to them.
- VancouverMan 2y ago> where real people hang out and have real conversations I don't consider the discussions there to be "real" in any meaningful way, thanks to the extensive moderation. From what I've seen, there typically ends up being a small handful of moderator-enforced narratives that are deemed "acceptable" for a given subreddit, and any commenters deviating from those narratives get banned, or their comments end up as "[removed]" by "[deleted]", or the comments get obscured with the "comment score below threshold" notice. It's generally some of the most one-sided and blandest discussion around. Given that there's often no meaningful back-and-forth involving differing perspectives of any sort, I'm not even sure if it should be considered "discussion". It's more like regurgitation and repetition. I've found the situation to be particularly bad on the Canadian locale-specific subreddits, for example, but a enough of the tech-oriented ones I've seen seem to end up like that, too.
- nomilk 2y agoSuppose a crawler or rival search engine doesn’t respect robots.txt, reddit can’t stop them. Make it a bit trickier, yes, but not stop them.
- eschneider 2y agoIt is evidence that they didn't have permission if you sue them.
- kingnothing 2y agoThere's no grounds on which to file suit. The 9th circuit court found web scraping is legal. https://techcrunch.com/2022/04/18/web-scraping-legal-court/ https://techcrunch.com/2022/04/18/web-scraping-legal-court/
- tagawa 2y agoThis is not even scraping - it’s just crawling and indexing.
- miyuru 2y agoreddit blocked datacenter IPs even before this change.
- nomilk 2y agoCould a motivated scraper not buy IPs/proxies that aren’t in those ranges, i.e. to blend in with general users?
- xeromal 2y agoJust like every security feature in the physical and digital worlds, security just inconveniences honest people and the cost to bypass reduces the amount of people who try. Eventually it becomes expensive to scrape reddit's data and most people will stop.
- 2y ago
- tempfile 2y agoHopefully this paves the way for antitrust action, but I won't hold my breath. Reddit's justification for this is profoundly wrong. Their "public content policy" is absurd doublespeak, and counter to everything the open internet is and hopes to be. You cannot simultaneously call yourself "open" and "public" while refusing access to automated clients. Every client is automated. They even go so far as to say that "crawling" (also known as "downloading") is an "abuse" and violates user privacy. This is absurd, and not justified. I would love to see legislation that restricted server operators' ability to prohibit automated access in this way, but I suppose it will never happen. Some people in this thread have attempted to justify the policy by saying "they have to protect their income streams". No they don't. You don't have a right to an income stream, and you certainly don't have a right to lie in order to get all the benefits of an open internet with none of the downsides. Noting of course that the "downsides" are in this case actually just "competitors".
- semiquaver 2y agoSorry, what is the antitrust concern about Reddit blocking crawlers that aren’t paying them? Surely you don’t think Reddit has a monopoly on anything? Or are you somehow suggesting that it’s google’s fault that Reddit took this step? I don’t see any indication that’s the case.
- em-bee 2y agonot that reddit has a monopoly, but that google has. google is using their power to prevent others from competing. the problem here is of course that if reddit would be in financial trouble (i don't know if they are but let's imagine they need this money), they'd be between a rock and a hard place. google should not be allowed to make exclusive deals, and reddit could not survive without the deal, then what would be left? google buys reddit, or the relevant authority approves of the deal? i thought about the same problem with firefox. let's assume firefox is forced to allow people to make a choice of the default search engine (just like microsoft was forced to allow a choice of default browser on windows) then google might stop paying mozilla, and they could end up in financial trouble. ideally no company ever depends on a single other company that much, but that only works if we don't allow companies to grow that much in the first place.
- r_singh 2y agoI wonder how Aaron Swartz would react to this
- geodel 2y agoMy guess is he'd freak out once he'd hear that lawyers, law enforcement may get involved on this issue.
- deleted 2y ago[deleted]
- ykonstant 2y agoIt's ironic, because Reddit is the only search engine that works on Google now thanks to shittening.
- voisin 2y agoMakes sense that Google did this deal since their search quality tanked and they became an de facto front end UI for Reddit.
- NoMoreNicksLeft 2y agoUp until 2016 (I think, +/- 1 year), if you could remember 3 uncommon words in a comment, you could find any reddit post instantly on Google. I'd want to follow up on a thread from weeks ago, and it was magic. Number one result. Then one day that just stopped working, and even adding site:*.reddit.com didn't fix it. At the time, I think, I didn't realize that it was mostly Google's fault, I thought maybe Reddit had changed their infrastructure so that it couldn't be crawled properly. Google hasn't been a search engine in a long while, it's just an advertisement engine now.
- dev1ycan 2y agoit's so bad it's crazy, you can legit not find stuff on the internet anymore, it's the same with youtube, I search something and get like 20 or so results and then everything else is hidden. it started when youtube removed the ability to search for videos older than 5 years, if I had to guess? cost saving, have every old video in cheaper storage... but it sort of fragments youtube, every couple of years you only get newer content.
- stuffoverflow 2y agoOne day I was searching live videos of a local band from youtube and when sorting by upload date the oldest video was from 2010. I knew for a fact there had to be older videos so I got a youtube API key and searched via the API, ended up finding multiple videos starting from 2006. Learned that youtube is full of videos that are basically impossible to find with the regular search.
- dev1ycan 2y agoOne of my past times is looking up anime opening reactions for fun/to hear people listen to bands I enjoy that sometimes do anime ops, searching that is so scary, you see the same 4 people or every month or so you get 1 or get 1 removed, you can't tell me there's not more than 4 people on the internet that upload opening reactions... when anime is a billion+ dollar industry, if you know particular channels you can see daily uploads on plenty of channels, that are simply search banned for whatever reason by the algorithm. But yeah, the most outrageous one is older videos, I do believe the reason is that they are using some long term cloud storage that is cheaper for older videos so they removed the ability to search by date. Additionally, I don't believe the API fully fixes it, because Bing has a wrapper for youtube and searches do not really vary
- lowbloodsugar 2y agoFunny that source of TFA blocked me from reading the whole thing.
- endejerv 2y ago[flagged]
- endejerv 2y ago[flagged]
- roughly 2y agoBoy, the LLMs have really been an apocalypse moment for the web, haven’t they? Between hoovering up and monetizing every bit of content they can without any attribution or compensation and the absolute flood of mediocre generated content, they’ve really done in the last straggling remains of the open internet. It’s not like everyone wasn’t already pulling the same grift, but quantity really does have a quality all its own.
- imglorp 2y agoOf course, we have to be careful not to villainize a neutral tech. Instead let's call it what it is: unchecked capitalism and monopolistic behaviors. Capitalism seems to work ok for the common good until you remove all the protections. LLMs provide a defacto monopoly for the owner which must already be a near monopoly: they take vast resources to train; only a giant corp can afford to buy all the content and provision enough resources to train one. LLM did not enshittify what's left of the internet, greed did it.
- synicalx 2y ago> Of course, we have to be careful not to villainize a neutral tech This is a very good point IMO. If we're going to chastise LLM's we may as well give servers, switches, routers, fiber-optic cable, and silicone a bollocking as well since that's ultimately what's facilitating all this.
- latexr 2y agoNo, those are not comparable. If someone criticises the electric chair, it’s not reasonable to defend it with “if we’re going to chastise the electric chair we may as well give wood, metal, chairs, and electricity a bollocking since that's ultimately what's facilitating all this”. Things are more than the sum of their parts. If you have a ton of beneficial things which can be cobbled together into one bad detrimental thing, the existence of the latter does not remove the benefits of the former.
- latexr 2y agoOn the one hand, you’re absolutely right. But on the other hand it’s not like it matters in practice. Isn’t most technology technically neutral? But it’s also made to be used by people, who can do so beneficially or detrimentally. Criticising a technology is a shorthand for criticising how it’s used.
- mediumsmart 2y agothat is awesome but I can't open old.reddit.com in my browser so its a non-issue.
- daft_pink 2y agoI don’t understand how this isn’t anti-competitive behavior. It seems like reddit has to offer this deal with similar terms to google’s competitors.
- talldayo 2y agoThey do offer that deal to others; a big news story was when OpenAI bought Reddit's data they were selling: https://openai.com/index/openai-and-reddit-partnership/ https://openai.com/index/openai-and-reddit-partnership/
- dathinab 2y agoyep, but for things which are "only" search engines it's not a viable offer. Only if you expect "big AI business value" from it does it make sense, maybe.
- Suppafly 2y agoMost business deals are anti-competitive in some way. What makes you think this specifically rises to the level where they'd legally have to offer similar terms to competitors?
- daft_pink 2y agoI’m not sure. Maybe the angle is that Google is anti-competitive by signing an agreement that limits information to it’s rivals. Being forced into using google services, because they are paying information companies to deal only with them seems like a disaster for the web.
- deleted 2y ago[deleted]
- carlosjobim 2y agoWhy in the world would they have to do that? There are thousands of exclusive business-to-business deals being signed into action every second of the day.
- 2y ago
- mutatio 2y agoIt's funny in the context of Google's past motto of "don't be evil". I feel the right thing for Google here would have been to decline any deal regarding exclusivity, then Reddit wouldn't have pulled the trigger with its robots.txt update. The entire manoeuvre required both parties.
- peddling-brink 2y agoGoogle should abandon its mission to “organize the world’s information” because doing so requires spending money for valuable data, and others might not want to spend that money?
- tbeseda 2y agohttps://archive.li/GS2I0 https://archive.li/GS2I0
- dathinab 2y agoWorse it doesn't even really "work" anymore, giving how most search are flooded with garbage SEO results and payed advertisements "basically" looking like search results (most times more garbage not what you are looking for results, int he cases where it isn't it quite often times is on the line of "googles algorithm blackmailing companies to buy ads for users which want to find them through google but wouldn't without ads".) I wonder if this might affect redis, as in slowly kill it's user base especially when it comes to user providing (and often also looking for) high quality content, because who of such users would want to use google search?
- john-radio 2y ago> Worse it doesn't even really "work" anymore, giving how most search are flooded with garbage SEO results and payed advertisements "basically" looking like search results ... I don't understand what you're saying. That's exactly why people append `site:reddit.com` to their searches in the first place, because those search results typically aren't like that.
- wwweston 2y agoOr at least, reddit posts and comments that are content messaging / marketing (human or AI) fit in better with earnest and natural posts, so that they're more effective.
- ppcland 2y ago[dead]
- wtf242 2y agoThis problem is only going to get worse. for my thegreatestbooks.org site i used to just get indexed/scraped by google and bing. now it's like 50+ AI bots scraping my entire site just so they can train a LLM to answer questions my site answers without having a user ever visit my site. I just checked cloudflare and in the past 24 hours I've had 1.2 million bot/automated requests
- sct202 2y agoThere's a new setting in Cloudflare to block AI/scraper bots. https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click https://blog.cloudflare.com/declaring-your-aindependence-blo...
- venkat223 2y agogoogle is selfish
- venkat223 2y agoGoogle is selfish
- StrauXX 2y agoIANAL but as far as I understand the current legal status (in the US) a change in robots.txt or terms and conditions is not binding for web scrapers since the data is publicly accessible. Neither does displaying a banner "By using this site you accept our terms and conditions" change anything about that. The only thing that can make these kinds of terms binding is if the data is only accessible after proactively accepting terms. For instance by restricting the website until one has created an account. Linkedin lost a case against a startup scraping and indexing their data because of that a few years ago.
- jpalomaki 2y agoQuite sure they are also enforcing these with some technical measures to limit scraping.
- renlo 2y agoAs was LinkedIn, who was forced to rate stop limiting / IP-banning scrapers for public pages.
- therealdrag0 2y agoReally? That seems strange.
- altdataseller 2y agoMany other websites enforce rate limiting or IP banning for public pages (using products like Cloudflare). Why is this not legal?
- qingcharles 2y agoAt the federal level; but states have their own laws. For instance, it can get you 5 years in prison in Illinois to violate a web site ToS. https://www.ilga.gov/legislation/ilcs/ilcs4.asp?DocName=072000050HArt%2E+17%2C+Subdiv%2E+30&ActID=1876&ChapterID=53&SeqStart=59300000&SeqEnd=60000000 https://www.ilga.gov/legislation/ilcs/ilcs4.asp?DocName=0720...
- numbers 2y ago"Information is power. But like all power, there are those who want to keep it for themselves. The world’s entire scientific and cultural heritage, published over centuries in books and journals, is increasingly being digitized and locked up by a handful of private corporations." - Aaron Swartz (2008)
- Khelavaster 2y agorobots.txt isn't legally binding. Can Reddit really force Bing not to crawl it..?
- melodyogonna 2y agoWait that's actually terrible.
- bitpush 2y agoWhen Microsoft strikes an exclusive deal with OpenAI to use their models, it is a smart, brilliant, clever move. When Apple strikes an exclusive deal with suppliers for parts, it is sound business practice. When Google strikes an exclusive deal with Reddit, it is .. Some of you have no idea how businesses work, and it shows.
- riku_iki 2y ago> When Google strikes an exclusive deal with Reddit, it is .. It's because reddit is selling content created by users, base on promises that reddit supports open internet, open data, etc, without their consent and sharing revenue, which maybe legal but likely not ethical.
- bitpush 2y agoLet's get specific. You're confusing with copyright and licensing. The users hold the copyright (reddit claim that they made the meme) but reddit has the non-exclusive right to redistribute and license the content. Two different things.
- riku_iki 2y ago> but reddit has the non-exclusive right to redistribute and license the content. that's what I said: it is legal.
- arnaudsm 2y agoI understand the AI context, but this is dangerously anticompetitive for other search engines. This is a dangerous precedent for the internet. Business conglomerates have been controlling most of the web, but refusing basic interoperability is even worse.
- zooq_ai 2y agoThere is nothing preventing search companies paying the same $60 Million to license content. If reddit had exclusive agreement, it would be anti-competive. This is classic HN anti-Google tirade (and downvoting facts, logic and concepts of free market)
- pluc 2y agoPaying 60 million to every site you want to index is also a bad precedent to set. Why can Reddit get paid and XYZ can't?
- zooq_ai 2y agoAnyone can ask for licensing deal. I'm sure NY Times, Conde Nest all have licensing deals. Mr. Beast signed a deal with Amazon. Joe Rogan with Spotify. Why is it hard to understand? Even HN can get a licensing deal if they want to. If you are producing content, you have every right to do what you want to with the content.
- SlackingOff123 2y agoReddit is not producing any content; its users are.
- zooq_ai 2y agoNot the point. If users don't like it they can go somewhere else to post. For practical purposes, reddit can do whatever they want with users post. It's right there in TOS
- thih9 2y agoStory / rant warning. I remember seeing an unhelpful hyperlink for the first time. It was a random word in the body of a random tech site that redirected to a list of articles from that site tagged with that term. I remember being stunned, my expectation was that the link would lead me to another website, one that would be an authoritative source on that term and freely accessible. 20 years later we get a paywalled article about fragmented web – and we’re not slowing down.
- lmeyerov 2y agoFWIW, we inquired to the reddit sales team about paying for data sometime last year, as we do similar elsewhere for use cases like helping emergency responders, and even though they were launching the program and asking for customers... no email back. Nor on our second and I think third attempt. I'm not sure what to make of that.
- morkalork 2y agoHow much were you willing to pay? Still, rude of them not to even discuss the issue. Every time I've gone to buy data, if I'm too small of a fish, vendors have always been happy refer me to a reseller.
- lmeyerov 2y agoWe do 4-6 figures/yr for providers which is normal in our world An enterprise sales team with only 1 customer happens (eg, Mozilla 's search bar), but... That's surprising here, and scary as a sustainable & scalable business. Ignoring 5-6 figure/yr inquiries says a lot to me. In contrast, we did that same-day with Twitter without talking to anyone.
- heisenbit 2y agoCertainly rude but also possibly legally problematic. If they were judged to be in a dominant position in a market and were found making deals with exclusivity then it can get expensive. It all depends of course what the market is. If one looks as reddit not as a whole but as a collection of niches then one could imho find niches where reddit has a dominant knowledge position.
- jumploops 2y agoIIRC, GPT-2 was primarily trained on Reddit[0] [0]https://www.reddit.com/r/ChatGPT/comments/133xgb5/gpt2_was_primarily_trained_through_reddit/ https://www.reddit.com/r/ChatGPT/comments/133xgb5/gpt2_was_p...
- neilv 2y agoI'm concerned multiple ways by this, but I also could see some positive fallout from this, if it sets precedents that help protect 'content' owners from AI goldrush companies just taking everything.
- gtirloni 2y agoAI companies are the least of our worries in the Reddit situation. The fact that Reddit has full control of user-generated data to do as they please gives them freedom to do as they please. I think this is the crux of today's issue. AI companies like Google, Microsoft and OpenAI have deep pockets to 'unprotect' themselves from anything. The barrier to entry is for small AI companies and those aren't really making an impact currently.
- r_singh 2y agoThinking from reddits perspective they have nothing to lose really. It’s not like other search engines are going to pay any attention to the robots txt and Google’s AI would have still scraped data from Reddit regardless of the deal. Now they will just feel less bad about not citing sources possibly, depending on the user experience they want to deliver.
- dbg31415 2y agoEvery time I think, “How scummy…” Reddit always finds another way to go lower.
- earthboundkid 2y agoThey literally think the scissor statement is a real thing that will really work, fml.
- 1vuio0pswjnm7 2y agoWorks where archive.li is blocked: https://cc.bingj.com/cache.aspx?d=5070227914243&w=ljIRk8yx42zrUAY3rWIxncz2TzI8fz6F https://cc.bingj.com/cache.aspx?d=5070227914243&w=ljIRk8yx42...
- TesterVetter 2y ago[dead]
- 1vuio0pswjnm7 2y ago"If you use Bing, DuckDuckGo, Mojeek, Qwant or any other alternative search engine that doesn't rely on Google's indexing and search Reddit by using "site:reddit.com," you will not see any results from the last week." The veracity of this statement is questionable. I found at least four web search engines not using Google's index that produced results from the last week. Example: Recent eruption at Yellowstone Black Diamond Pool https://www.ecosia.org/search?method=index&q=site:reddit.com+black+diamond+pool&p=1 https://www.ecosia.org/search?method=index&q=site:reddit.com... https://search.brave.com/search?q=reddit.com+black+diamond+pool&source=web&offset=0 https://search.brave.com/search?q=reddit.com+black+diamond+p... https://api.yep.com/fs/2/search?client=web&gl=all&no_correct=false&q=site:reddit.com+black+diamond+pool&safeSearch=off&type=web&limit=0 https://api.yep.com/fs/2/search?client=web&gl=all&no_correct... POST /sp/search HTTP/1.0 host: www.startpage.com content-length: 74 content-type: application/x-www-form-urlencoded query=site:reddit.com black diamond pool&abp=-1&t=&lui=english&sc=&cat=web At least for this example, I got the same desired result using Reddit site search. https://old.reddit.com/search/?q=black+diamond+pool https://old.reddit.com/search/?q=black+diamond+pool If anyone has some good examples of search queries that I can test showing why a search engine must be used, please share.
- lopis 2y agoEcosia does use Google or Bing, which you can select in the settings.
- niutech 2y agoBut Brave Search has their own independent index.
- 1vuio0pswjnm7 2y ago"Its search engine uses Microsoft's Bing's technology, with whom it has a long-term arrangement." https://www.bbc.com/news/business-53922786 https://www.bbc.com/news/business-53922786
- myrandomcomment 2y agoSo I went Slashdot, Digg, Reddit. I stopped spending any time on Reddit 5 years ago. Not worth it.
- nullc 2y agoIt's weird to say that reddit "works" with google. Every page they serve to google is stuffed full of hidden unrelated content, so any reddit result in google is unlikely to actually contain what you were searching for. Google really should blacklist reddit entirely for this practice, but sadly as bad as reddit is it's still a much higher quality result than average for google.
- account42 2y agoUgh it's absurd at how incompetent Google is at filtering out "related" content or similar volatile "sidebar" feeds in the sitesthey index that has nothing to do with the main content and won't be there when the user actually opens that link.
- blackeyeblitzar 2y agoWe need laws that make it so that giant platforms like Reddit have no exclusive rights to content submitted by users. It would be ridiculous for only Google to be able to train AI on YouTube or Reddit content for example.
- ChrisArchitect 2y agoFine with this. This is the world OpenAI created. And all the people that started searching with +Reddit tacked on weirdly like 5 years ago. Reddit's covering themselves from internal user-concern and their general exposure to AI training and Google was smart enough to get on that quickly. We'll see what Bing's take is and what changes if anything now that 404medias's outrage farming is at play. This isn't a recent change afterall, month ago?
- ein0p 2y agoGood for other search engines, I suppose. Reddit is a giant toxic pile of bovine manure.
- manishsharan 2y agoFor my use cases , Google is pretty much useless without Reddit For example, when I search for product reviews, I always specify reddit. Otherwise the search results are inundated with SEO spam.
- cyanydeez 2y agoWork is such a flimsy word for qhat google currently does with search As soon as someone shows me a search engine that restores quality of searxh, im getting a subscription for work. It really cany be hard to whitelist sources and index appropiately. Get goimg nerds , google has fallen.
- ozgrakkurt 2y agoStopped using reddit after they hindered login-less viewing and blocked vpns. Everyone who respect themselves should start moving away from it imho. Same thing with google
- dakial1 2y agoAnd now lets watch white, grey and black hat SEO destroy reddit even more.
- bkjshki 2y ago[dead]
- Slava_Propanei 2y ago[dead]