8 ms·
Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT
by Zaheer 3y ago
Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)
- supriyo-biswas 3y agoOn that note, I also wonder if they end up getting this information anyway through another source like Common Crawl.
- ramraj07 3y agoHoping this is what they’ll use to train future models and deprecate the older ones before the legal cases proceed any further.
- p-e-w 3y agoThe legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.
- DSingularity 3y agoI think a key idea is that with the amount of jurisdictions and number of courts the odds that a clean and sympathetic judge can be found approach one. I would argue that European jurisdictions are inherently less likely to be in pockets of American corporate interest and they are more likely to hear cases where fundamental human freedoms are at stake because both of these are existential threats to European independence. In the US similar arguments can be made in states vs federal or the various federal circuits. Courts are more deliberate than you would like — no denying that. But this is a feature not a flaw. It may be that damage will be done by then. Perhaps irreversible. But I would like to think if there is a will there is a way and that if things are terrible enough the governments will be bold in their responses.
- p-e-w 3y agoThe corporations that provide AI hold all the power because people (and businesses!) want to use their products. Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users. Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent? Any individual government (except, perhaps, the combined US and EU governments) is powerless against today's technology megacorporations, because they can take much more away from a country than that country can take from them. If push ever comes to shove, it will become obvious where the true power lies. So far, the corporations have barely even tried to throw their weight around.
- oli-g 3y ago> Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users. That's one possible outcome. (ETA: You DO have a point here, but...) The other is, you know, something like every website explicitly telling me, via an annoying popup, how much they value my privacy. Also, me not being able to access half of US news sites to this day. The last time EU raised their finger, every technology company (FAANG included) shat their pants. And that was simpler times, times when a cookie stored in your temp folder without websites shouting they're about to do so, was somehow the biggest concern of an EU netizen. It almost seems ridiculous, compared to the damage AI could do (the extent of which which nobody really knows).
- ben_w 3y ago> Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent? Bof, les alternatives à ChatGPT ne sont pas si mal. And even if the open source alternatives were far behind rather than just a bit — all this talk about corporate moats and their absence may be blind to the strengths of OpenAI's offerings, but even so it can be replaced if it must — the storms of protest in France are normally by the people, not by the corporations.
- moonchrome 3y agoIf you think copyright lawyers and the entertainment industry is going to let some AI upstarts launder their IP without a fight you aren't paying attention.
- p-e-w 3y ago> AI upstarts You mean corporations that wield more power than most governments, and have revenues equivalent to the GDP of entire countries? If Universal or 20th Century Fox were to ever become a serious obstacle, Google and Microsoft are simply going to buy them. This isn't the early 2000s anymore. The power balance has shifted dramatically.
- astrange 3y agoFAANG already haven't bought or started competitors to the record labels they resell in their music stores. Don't see why they'll start now.
- ben_w 3y agoI just looked it up because I have no idea how big the music industry is, and… US$26.2 billion globally in 2022 according to IFPI, and US$31.2 billion according to Statista. Other than Netflix, I think FAANG just doesn't care that much about such a small market (the market being "actually producing it", given they're already part of the previous numbers for selling and streaming it). And of course, both A's and the N of FAANG have their own commissioned TV/film content.
- deleted 3y ago[deleted]
- fakedang 3y agoI thought the Hollywood strike was about the entertainment industry planning on using AI to substitute extras? Sorry but they're all in bed together.
- astrange 3y agoThe legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.
- nextaccountic 3y agoLossy compression of a 1MB original image into a 20kb compressed image doesn't make copyright go away But that's essentially what LLMs are doing, lossy compression of the entire web
- bayindirh 3y agoYet gleefully emits its training data when one asks the right questions. It can be code, prose or images. Yeah, doesn't remember. Mhm... Oh, it just can't remember the license terms of the code it "reads", so it can't comply with these licenses or help people to comply with these licenses. Convenient.
- ben_w 3y ago> If anything is transformative, an AI that doesn't memorize its input is. I suspect the answer to the question "is it, though?" is one for the lawyers and lawmakers rather than for the software developers, and it may well vary wildly by jurisdiction.
- __loam 3y agoFair use specifically has a clause about disrupting the market for the original work lol. Being transformative isn't the only aspect of fair use, and even if training is legal, you're still a douche for training on art without permission.
- nologic01 3y agoIt doesnt memorize anything. It just needs gazillion parameters that approach the size of the training set to finesse its conversational accent.
- 3y ago
- raincole 3y ago(fortunately)
- brianjking 3y agoYeah, here in the USA we haven't figured out Section 230 yet. There is no hope for sensible (or illogical) AI regulation.
- babl-yc 3y agoGPT-4 finished training in August 2022, before the release of ChatGPT. If they had announced this sooner hardly anyone on the internet would have noticed. Props to them for adding it now.
- selcuka 3y agoMaybe some people weren't aware, but GPT-3 (and GPT-2, before that) APIs had been around for some time when ChatGPT was launched. I joined the private beta in early 2021.
- kolinko 3y agoPreviously they used OpenCrawl afaik, so they didn’t have a dedicated crawler
- gmerc 3y agoIt’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.
- zarzavat 3y agoThat would be a hilariously bad idea for them. Their business is based on fair use. The only way to enforce restrictions against scraping is through copyright law because obviously you can run the spidering code from any jurisdiction you want, so any law that says “thou shall not scrape” is toothless unless it acts through copyright. Any workable restrictions against using scraped data would also make ChatGPT illegal too.
- gmerc 3y agoNonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.
- staticman2 3y agoCopyright laws do in fact (or have in fact) acted retroactively.
- mmmmmmtoes 3y agoWhen? Not doubting, just curious about scope and type of scenarios where it's happened.
- MWil 3y agoI cannot for the life of me find the links but I feel like this happened with Monopoly or some other board game.
- staticman2 3y agoI'm going largely by memory but when the U.S. expanded copyright at one point they actually took some stuff out of the public domain. You can look it up but the current formula is authors life plus 70 and a different formula for corporate works, and when they expanded it most recently there were actually some public domain works that become not public domain retroactively. (A quick google search reveals the 1976 Act added 19 years to the terms of existing copyrights, this might be what I'm thinking of-- in other words some works that had copyright expired then had them renewed and removed from the public domain.) There's also copyright reversion, which is a related new provision that applied to older copyrighted works. Quoting from an article I just pulled up "...the 1976 Act created a new right allowing authors and their heirs to terminate a prior grant of copyright, the Act also set forth specific steps concerning the timing and contents of the termination notice that must be served in order to effectuate termination. The termination of a grant may be effective “at any time during a period of five years beginning of the end of 56 years from the date the copyright was originally secured”..." But this is a red herring because the fact a model has been trained in the past doesn't mean a copyright lawsuit is "retroactive". The infringement would presumably be occuring anew every day you make it available on your web site.
- zitterbewegung 3y agoAt least now you can see if your website is being crawled by them. It also exposes them to be easily targeted to send them invalid data or even misinformation. Before people will already doing that before by putting information that people wouldn’t see like white text on a white background.
- gwern 3y agoTheir papers say they were using Common Crawl for crawling. If you didn't want your pages in Common Crawl (eg. Twitter didn't) for use in many downstream analyses or uses beyond just OA, you could already have said so in your robots.txt.
- flangola7 3y agoThat's not consent though. Consent is not granted until explicit stated in the affirmative. Try applying "assume yes initially, until told otherwise" to entering someone's house or touching someone's body and let me know how that works out for you.
- 6gvONxR4sf7o 3y agoOpt out != opt in. This reminds me of the beginning of the hitchhikers guide to the galaxy where Dent’s house is being demolished but the notice had been on display in a locked basement below city hall or something. He could have objected, technically!
- gwern 3y agoI don't think that comparison is valid, and in fact, actually comparing them shows how reasonable it is: the HHGtG example is egregious because it is imposed silently, long after the fact, made deliberately invisible and hard to access, and discoverable only after the fact. All of those are false for robots.txt and Common Crawl. These are well-known, easy, old protocols which long predate most of the websites in question, which is completely disanalogous to the HHGtG example. Specifically: robots.txt precedes pretty much every website in existence. It's not some last-minute addition tacked on. Further, it is straightforward: you can deny scraping to everyone with a simple 'User-agent: * / Disallow: /' or nofollow headers (also 1 line in a web server like Apache or nginx) - hardly burdensome, and it rules out all projects, not just Common Crawl. Common Crawl is itself, incidentally, 15 years old, and long predates many of the websites it crawls, its crawler operates in the open with a clear user-agent and no shenanigans, and you can further look up what's in it because it's public. (This is how I know Twitter isn't in it: when people claimed GPT-3 was stealing answers from Twitter, I could just go check.) It is also well known, even many non-webmaster web users know about it because it governs what you'll see in search engines, what will be downloaded by some agents like wget by default, is covered early on in website materials, and so on.