10 ms·
The shady world of Brave selling copyrighted data for AI training
- ricardo81 3y agoMy entirely biased opinion is https://www.mojeek.com/ https://www.mojeek.com/ - a traditional search engine crawler (as in, follow links on the web) that identifies its user agent. Dead Simple. The open web, you can search on it.
- verisimi 3y agoHow long until IP works its way onto ai training data or ais themselves? Ie that for some specific instance, the training is intentionally wrong, so as to check and prove that there has been a breach of IP.
- homeless_engi 3y agoDo you mean like trap streets? Seems like a good idea for model makers https://en.wikipedia.org/wiki/Trap_street https://en.wikipedia.org/wiki/Trap_street
- lacrimacida 3y agoYeah, something like that may be already happening and various actors building their cases as we speak.
- DesktopMonitor 3y agoWhile not intentionally wrong, Van Halen's brown M&M's rider comes to mind as an example of a similar measure.
- the8472 3y agoDepends, how do you distinguish humans acquiring knowledge by ingesting copyrighted content vs. a human using an AI that ingested copyrighted content?
- JumpCrisscross 3y ago> how do you distinguish humans acquiring knowledge by ingesting copyrighted content vs. a human using an AI that ingested copyrighted content Doesn’t matter when the content is reproduced verbatim, as Brave is doing. If I memorise your content and then repeat it as my own, I’m not somehow off the hook for copyright violation and plagiarism.
- throwaway72762 3y agoI think this title is overstated. It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt. (Also, crawling as a service has been a thing for a while.)
- jsnell 3y ago> It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt I feel like I'm missing something. What the article claims they're doing is: 1. Misrepresenting what rights they have, and selling access to those rights. 2. Stealth-crawling the web, hiding from the webmasters just how much Brave is crawling their site, and making it impossible to block just their crawler. How is either of these the right thing? I mean, for somebody besides Brave. What "attempt" are they making that other companies aren't?
- woah 3y agoIs there something wrong with accessing information that someone has posted for public access?
- JumpCrisscross 3y ago> Is there something wrong with accessing information that someone has posted for public access? The Wikipedia example is glaring. They’re scraping content, stripping attribution and reselling it with a right to lock it down in a way that is not allowed by the original license. Brave is laundering copyleft content while lying to their customers by selling a license they can’t give. If you’d like, you can sidestep the morality of copyright entirely and focus on the plagiarism and fraud.
- theamk 3y agoYes. Legalities aside, stripping attribution (author names) from contents which specifically requires keeping it, it a really shitty thing to do. (The fact that they include original URL does not change much, given that they explicitly market it as "Data for AI" and those systems never have attribution)
- isodev 3y agoI firmly believe that corps like these don't deserve the benefit of the doubt. Google, Brave and really anyone big enough to allow themselves to do bad things and get away with it must adhere to a standard where they proactively show their stuff doesn't have malicious intents.
- sourcecodeplz 3y agoAs always, if the product is free, you are the product...
- isykt 3y ago“Although the saying tells us “If it’s free, then you are the product,” that is also incorrect. We are the sources of surveillance capitalism’s crucial surplus: the objects of a technologically advanced and increasingly inescapable raw-material-extraction operation. Surveillance capitalism’s actual customers are the enterprises that trade in its markets for future behavior.” Excerpt From The Age of Surveillance Capitalism Shoshana Zuboff
- woah 3y agoDid you read the article? This is about a paid web crawling api that the author thinks is too good or something. Nothing about a free product
- xp84 3y agoFrom article: > without any worry for copyright infringement because Brave acts as a middleman. This isn’t how law works. Unless Brave is explicitly indemnifying all their customers (which their lawyers would have to be insane to let them do), any trouble you could get in, is going to be 100% your problem. Pointing the finger at Brave could theoretically get them in trouble too, but would in no way let you off the hook.
- lopatin 3y agoWhy use brave if my info is already being leaked by third parties? E.g. experian. Is it worth the inconvenience and their repeated tricky attempts at monetizing their security conscious niche? Not being facetious, just a real question from a non security conscious person.
- soundnote 3y agoYou get degoogled Chromium with e2ee bookmarks etc. sync and a lot of nice convenience features like vertical tabs and mobile background video playback. And if it's your cup of tea, they let you straight up pay money for the search engine.
- asynchronous 3y agoIt’s built in Ad blocker and other features are heads and tails above anything else I’ve used before, personally.
- hartator 3y ago> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about your browsing? Other search engines usage, but also crawl pages on your computer to help build their search index? Ref: https://github.com/brave/web-discovery-project/blob/main/modules/web-discovery-project/sources/README.md https://github.com/brave/web-discovery-project/blob/main/mod...
- drusepth 3y agoThis specific feature is already opt-in, but historically the answer has always been "yes" for dozens of 'features' like this that fly under the radar until users start complaining, and then eventually get converted to opt-in or removed in order to save face.
- choppaface 3y agoAnd Google gets the same data joining your cookies ever since Google Plus unified auth across their properties a decade ago. Wait you mean you thought G+ was supposed to compete against Facebook-the-product and not just Facebook-the-ad-network? Oops Brave is perfectly OK with having oopsies too
- jrmg 3y agoIf you don’t trust Brave then, yeah, they could be doing anything in the browser or on their servers - but that snippet you quoted is a slightly out of context statement from a big document about how they collect data like this, but _don’t_ collect or store it in a way that they could associate it with a user. If you don’t trust that they’re doing what they say they are, then the document doesn’t mean anything. Although that would also mean the quote is kind of meaningless…
- hartator 3y agoThe rest of the document is worst. They say they are using your computer to crawl pages you visit and report back to their server. Even Google doesn't do that.
- 411111111111111 3y agoIt's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting on you vs Google or MS
- siquick 3y agoThere are multiple ways you can pay Brave. https://brave.com/firewall-vpn/ https://brave.com/firewall-vpn/ https://account.brave.com/?intent=checkout&product=search https://account.brave.com/?intent=checkout&product=search https://brave.com/search/api/ https://brave.com/search/api/
- nicce 3y agoPeople still like to defend Brave when it gets caught on shady things over and over again. I guess there are no too many other options. For some people it is already too difficult to install uBlock or know its existence.
- lalaland1125 3y agoIt's because a lot of people are bought into BAT (Brave's cryptocurrency) and have a strong financial incentive to shill Brave.
- DaSHacka 3y agoBAT gives like no money, especially after the crypto crash. Its far more likely its just the browser wars of old, but with even less options to choose from people are going to be more adamant their choice is the best.
- boondoggle16 3y agoI do not use BAT or any crypto. Brave just works, and it blocks ads automatically when I tell friends to install it on their computers. I used to recommend Firefox, but Mozilla has totally jumped the shark (privacy violations [multiple], wastes too much money, blocks APIs that are useful with no real security risks while approving APIs with little use that do have security risks, etc, very user hostile). Chromium is obviously not trustworthy at this point, let alone Chrome. So that leaves like, Safari and Opera? Brendan Eich is the CEO of Brave, and I trust him. Mozilla was good until he was ousted for political reasons.
- 6gvONxR4sf7o 3y ago> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes > 2) The nature of the copyrighted work > 3) The amount and substantiality of the portion used in relation to the copyrighted work as a whole > 4) The effect of the use upon the potential market for or value of the copyrighted work [emphasis from TFA] HN always talks about derivative work and transformativeness, but never about these. The fourth one especially seems clear in its implications for models. Regardless, it makes it seem much less clear cut than people here often say.
- flangola7 3y agoMicrosoft is gambling on the hope that model training will be ruled fair use. This makes it seem that outcome is unlikely.
- brookst 3y agoDo you think a human learning something from reading is fair use? Or are we all copyright violators because reading that article altered our connectomes, and we may recall parts of it later?
- snickerbockers 3y agoYes it is considered fair use but it's also completely irrelevant because we're talking about a computer program not a person.
- ethanbond 3y agoThe point being raised is quite specific. Not sure if you’re willingly ignoring it or what? The answer is no, because you reading the article didn’t dramatically degrade its market value. An AI ingesting all content on the internet and then being ultra-effective at frontrunning that content for a large number of future readers does degrade its market value (and subsumes it into the model’s value).
- kodah 3y agoUnpopular opinion: the next iteration of privacy laws needs to factor in AI. If AI is allowed to slurp up PII or derogative works and the people defending it defend it with the zeal of cryptobros then we're in for a decade of real pain in terms of both copyright law, PII, and IP exposure.
- asynchronous 3y agoAI is going to do that irregardless- the debate is essentially going to revolve around how and what people can make new commercial works from that data.
- _fbpp 3y agoThe fun part is that the GDPR already does. The answer is you're not allowed to use personal data for AI. (And "personal data" here covers things like all public social media posts) Facebook recently got told by the CJEU that, no, they can't use people's posts to target advertisements. Even if those ads are what's paying for the platform. That you can't claim such processing as "part of the contract" unless it is absolutely necessary in the same way the post office needs an address to send a parcel. If Facebook can't even do that, there is no way LLMs will be allowed. (And remember. The GDPR does not care if your system doesn't distribute personal data. Any kind of processing at all falls under the GDPR's requirements) OpenAI is already being chased by the EU's privacy agencies. Right now they're in the process of asking pointed questions, things will heat up after that.
- BeFlatXIII 3y agoEnd result: EU AI enjoyers use a VPN plus a US-based credit card borrowed from a friend.
- lern_too_spel 3y agoBrave continues to be shady. They claim to respect robots.txt but don't identify their crawler if you want to block it. > They don't mention their crawler anywhere in their docs, either. So, if you wanted to block Brave from crawling and indexing and ultimately selling your content to third parties, your only option for the time being would be to block all crawlers, which is how Brave would be able to "respect robots.txt".
- k__ 3y agoThe websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!
- jonathansampson 3y agoI'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anonymity as a prerequisite, but it is also open-source for public scrutiny and evaluation: https://github.com/brave/web-discovery-project https://github.com/brave/web-discovery-project.
- deleted 3y ago[deleted]
- k__ 3y ago.es?
- yamsamnow 3y agoIt's not a clever approach, it's basically scraping Google results because that's where your users are searching. You follow the bread crumbs from Google searches. Cliqz entire history was based on this kind of thing, milking off other search engines by just deducting their ranking methods, it's parasitic. There's no cleverness about it.
- throwaway675309 3y agoI don't know a lot about this particular approach but your comment that it's just using Google results is blatantly false. It all depends on the search engine that the brave user is leveraging, or no search engine if they type in the URL directly into the header.
- niemandhier 3y agoThis discussion on fair use are always quite anglocentric. Atricle 3 and 4 of the EU 'Copyright in the Digital Single Market' give data miners quite extensive rights. Move operation to the EU, train a foundational model, than train a constitutional model based on that. As much as I hate the upcoming AI regulation, the CDSM is solid. https://academic.oup.com/grurint/article/71/8/685/6650009 https://academic.oup.com/grurint/article/71/8/685/6650009 https://eur-lex.europa.eu/eli/dir/2019/790/oj https://eur-lex.europa.eu/eli/dir/2019/790/oj Update: Fixed wrong link
- pedrocr 3y agoIt's not clear that "data mining" covers this use. These models are huge, big enough that they can just contain direct copies of copyrighted works. They've been shown to reproduce them relatively easily. The argument is that they've actually generalized enough or learned enough that they're now no longer the sum of the dataset. I can definitely see that being possible but the way the technology works it's really hard to know if that has happened or if what's happening instead is a bunch of copyright washing. There are some things that would make for good faith displays by the players in the space. For example, Microsoft has been investing a lot and yet their code offering is not trained on their internal code base. Same for Google. Start by doing that and I'll entertain the argument that your tools are fair use or data mining.
- niemandhier 3y agoMy reading of the relevant laws would actually lead me to believe that this is not a problem, as long as those reproductions are not returned and the eights holder did not opt out. But courts might decide differently. Regarding the copyright of returned material here is a good discussion: https://copyrightblog.kluweriplaw.com/2023/05/09/generative-ai-copyright-and-the-ai-act/ https://copyrightblog.kluweriplaw.com/2023/05/09/generative-...
- JumpCrisscross 3y ago> as long as those reproductions are not returned That’s the author’s entire gripe. Brave reproduced a Wikipedia entry without attribution and then slapped a copyright on it to boot.