25 ms·
Japan’s government will not enforce copyrights on data used in AI training
- LadyCailin 3y agoGood. Maybe this will force the US and other western countries to loosen copyright laws in general. Absolutely no reason to have anything copyrighted for nearly a century or more!
- sylware 3y agoConsidering the amount of animes Japan does produce, they could very probably train amazing AIs: they could assist in a significant way the creative process of such studios (look at digital corridor experiment). Worth to try for a long time and maybe waste a lot of resources.
- setr 3y agoMore than the anime production is that anime-fans are database animals — they catalogue anime to ridiculous degrees. Some of the best image training datasets were always the boorus; anime fans have basically been prepping for ML/AI since the 80s
- kyleyeats 3y agoIt'd be really great to finally see some anime models hit the Stable Diffusion scene.
- EscapeFromNY 3y agoI hope this is sarcasm :) The anime crowd were years ahead of the curve. https://www.thiswaifudoesnotexist.net/ https://www.thiswaifudoesnotexist.net/ came out in 2019
- deleted 3y ago[deleted]
- waboremo 3y agoIs the wording accurate here? This is essentially the only source besides the untranslated article and the machine translated version sounds confusing (whether it applies to what is created by AI or what can be consumed in training).
- resoluteteeth 3y agoWhat actually happened was that Takashi Kii of the Constitutional Democratic Party of Japan was arguing that the current laws (from 2018) are problematic because they are extremely loose and allow even illegally obtained content to be used for training, and he asked Keiko Nagao, Minister of Education, Culture, Sports, Science and Technology to confirm that this is the case. She confirmed that under current laws that's true but said that they need to keep an eye on it because there's a balance between the development of new AI technology and protection of copyright. The Ministry of Education, Culture, Sports, Science and Technology is also in the process of compiling information about case law on copyright and AI but there don't seem to be any current plans to amend the law again in either direction. (You can see the whole exchange here although the translated autogenerated subtitles on youtube may not be great: https://youtu.be/fyxx_0KmaKw?t=4457 https://youtu.be/fyxx_0KmaKw?t=4457 )
- numpad0 3y agoJapanese copyright law article 30-4 states[1]: > It is permissible to exploit a work, in ... cases ... it is not a person's purpose to personally enjoy or cause another person to enjoy ... provided, however, that this does not apply if the action would unreasonably prejudice the interests of the copyright owner ... > i)if it is done for use in testing to develop or put into practical use technology ... > (ii)if it is done for use in data analysis (meaning the extraction, comparison, classification, or other statistical analysis of the constituent ... > (iii)if it is exploited in the course of computer data processing or otherwise exploited in a way that does not involve what is expressed in the work being perceived by the human senses (for works of computer programming, such exploitation excludes the execution of the work on a computer), beyond as set forth in the preceding two items. Japanese legalese is a rather inefficient pseudo-european built on Japanese language, so I wouldn't recommend making decisions based on a blog article like this; there hasn't been too much news stories regarding this too as additional anecdotal datapoint. 1: https://www.japaneselawtranslation.go.jp/ja/laws/view/4207#je_ch2sc3sb5at4 https://www.japaneselawtranslation.go.jp/ja/laws/view/4207#j...
- georgewsinger 3y agoJapan also ranks 3rd (behind the USA & India, with larger populations) in ChatGPT usage: https://www.demandsage.com/chatgpt-statistics/ https://www.demandsage.com/chatgpt-statistics/ There's also been discussion of their government using ChatGPT to reduce red tape: https://www.bloomberg.com/news/articles/2023-04-18/japan-government-taps-chatgpt-to-cut-through-bureaucracy-deluge#xj4y7vzkg https://www.bloomberg.com/news/articles/2023-04-18/japan-gov... It's cool to see Japan and Japanese culture taking techno-optimist stances on AI.
- deleted 3y ago[deleted]
- ChatGTP 3y agoI strongly disagree. They need to actually address problems. Not throw tools at it. The problem Japan seems to have is they don’t understand AI and than they don’t understand software, which is they don’t understand a lot of modern tech. They’ll pay for this mistake just as they paid for being bad at software.
- yinser 3y ago[flagged]
- Loocid 3y agoWhich I find bizarre given how backwards Japan is in the adoption of other technologies. Eg their continued reliance on paper records and fax machines.
- johnwalkr 3y agoDo you have concrete examples? I recently moved from Japan to Europe after 10 years in Japan and while some things seemed old-fashioned in Japan, things also change overnight there. During Covid most companies changed to digital signing. In Japan I often needed a paper from the city, but it was an easy-to-obtain printout I could get immediately from city hall or even from 7-11. I tend to require the same papers in Europe too, with the extra hassle of needing some from Japan (2-3 months) and some from Europe (2-3 days). Example 1: get ID card at the airport when you emigrate to japan. In Europe it takes 2-3 months. Example 2: getting a local driving license takes 1 day provided your country has a treaty with japan. In Europe it can take several months because you need to request criminal records from your previous countries I have yet to encounter a situation that requires a fax machine in either place.
- nmkag 3y ago[flagged]
- surgical_fire 3y ago> This is a horrible decision from Japan. Eh, I doubt it. AI is potentially something that can massively increase productivity. With a declining population in the foreseeable future, an increased productivity may well be a boost that they need.
- Guthur 3y agoBoost for who though? if it works out there'll be more for less. Wages won't increase, the number of jobs won't increase the price of assets will inflate. None of these are good things for 99% of us.
- visarga 3y ago> Wages won't increase, the number of jobs won't increase When new capability appears, many industries pop up. It's a new market, a new gold rush. It happened many times, with cars, electricity, air transport, computers, internet. AI will spring many applications and will create jobs in those fields. We have been under a 260 year run of industrial revolution and 70 years of computer programming. And yet unemployment is low and IT jobs are well paid. Why do we have so many jobs? The computers are 1 million times faster now than 25 years ago, more deployed and better networked. Where is that productivity gain hiding? If we try to be realistic, current crop AI amounts to about 1.2x productivity gain. It's a nice to have thing, but not essential yet. It's really nice. But makes errors often enough that it almost negates its advantages. Error recovery is very costly. I foresee economic growth driven by AI, and people with AI skills being very efficient and well paid. AI shines most when it is used and then evaluated by a skilled human.
- barelyauser 3y ago>And yet unemployment is low and IT jobs are well paid Naive at best, tone deaf at worst. Unemployment is low? Most jobs barely allow a person to live a decent life. AI will empower few, the rest will find themselves unable to create any value. All the "economic growth" will be the increasing profits and economic inequality. When cars were created anyone could foresee the demand of people to manufacture them. Tell me, what jobs will AI create?
- WhatIsDukkha 3y agoSo this means llm trained on libgen and ... which is huge. Current models are only trained on "open" texts.
- IvanAchlaqullah 3y agoMiyazaki will definitely hate it. The last time someone demoed AI animation he said "I would never wish to incorporate this technology into my work at all. I strongly feel that this is an insult to life itself." [1] And that was before stuff like GPT-1 or Stable Diffusion exist. [1] https://www.indiewire.com/features/general/hayao-miyazaki-artificial-intelligence-animation-insult-to-life-studio-ghibli-1201757617/ https://www.indiewire.com/features/general/hayao-miyazaki-ar...
- shagie 3y agoThere is a pair of videos on this subject that I find rather good and informed. They are from the channel "The Art of Aaron Blaise" ( https://en.wikipedia.org/wiki/Aaron_Blaise https://en.wikipedia.org/wiki/Aaron_Blaise ) Disney Animator REACTS to AI Animation! https://youtu.be/xm7BwEsdVbQ https://youtu.be/xm7BwEsdVbQ (this is watching the Corridor Crew's video) Why AI will NOT be taking Your Animation job - https://youtu.be/-lhbzbSck04 https://youtu.be/-lhbzbSck04
- brandelune 3y agoThe original links here, to the actual question asked and the answer by the minister: https://kiitaka.net/21312/ https://kiitaka.net/21312/
- waboremo 3y agoThis sounds like the exact opposite of what the title is claiming.
- throwaway33381 3y agoYeah. Copyright laws in Japan are extremely strict. I would not be surprised at a complete ban on this. This is just a fluff piece like a lot of ads coming in recently.
- KinkySumo 3y agoHow is this an exact opposite of the OP's piece? The two takeaways you can get is: - As there are no external motives (eg. for profit) held by the AI when analyzing the works it is not a copyright infringement. - Parsing through whether the data set is legally obtained is not feasible therefore that shouldn't be the limiting factor to further development. However as the transcript was uploaded last month, things may have changed since then.
- fomine3 3y agoCopyright law was explicitly changed in 2019 to support AI development. https://storialaw.jp/en/service/bigdata/bigdata-12 https://storialaw.jp/en/service/bigdata/bigdata-12 https://japannews.yomiuri.co.jp/society/general-news/20230429-106420/ https://japannews.yomiuri.co.jp/society/general-news/2023042... Japan was behind about developing search engine service. Some people argued that it's due to copyright law (there's no fair use) and it shouldn't be repeated again. So the gov want to encourage AI developing by explicit law.
- getoj 3y agoI thought it would be too ironic for people to misunderstand this based on a machine translated version, so here is a genuine, human translation of the transcript, with boring bits redacted. Kii: Next question, again regarding generative AI. I would like to ask from the two perspectives of copyright protection and educational use. [...] First, can we understand that Japanese law permits the use of works for information analysis, both for non-commercial and commercial purposes, and acts other than copying, and using content that was uploaded illegally? Nagaoka: Use for non-commercial information analysis is permitted under Article 30-4 of the copyright act, provided that the purpose is not the enjoyment of the ideas and emotions expressed in the copyrighted work. Kii: Minister, I asked about four aspects of use for information analysis: non-commercial use, commercial use, acts other than copying, and illegally uploaded content. Please address the other three. Nagaoka: Use for commercial purposes is permitted under Article 30-4 of the copyright act, provided that the purpose is not the enjoyment of the ideas and emotions expressed in the copyrighted work, because that Article does not distinguish between information analysis for commercial or non-commercial purposes. Regarding copying, Article 30-4 of the Copyright Act does not distinguish based on the method of use, so use by means other than copying is permitted provided that the criteria are met. [...] Regarding content obtained from piracy sites and the like, [...] illegal uploading itself is infringement of copyright, and is subject to a damage claim, petition for injunction, or criminal punishment. However, it is not practically feasible to identify whether any particular work in a large collection obtained from the internet is copyrighted or not, so making this a criterion for information analysis would make it difficult to use information analysis for Big Data. In addition, as the use of a work for information analysis is not use for the purpose of enjoyment of the ideas or emotions expressed in the work, and even if [it were used in that manner] it would not overlap with the original market for the use of the work, so it is not considered to harm the interests of the copyright holder that are protected by the Copyright Act. As such, Article 30-4 of the Copyright Act does not have the legality of the work as a criterion. Kii: Minister, based on your answer, I think the greatest issue is that there is no protection against use that goes against the intentions of the creator or the copyright holder. I believe that new regulations will be necessary to address this point; will you consider such new regulations? Nagaoka: Article 30-4 of the Copyright Act provides for use that is not for the purpose of enjoying the ideas or emotions expressed in the work, and applies to acts that are considered not to affect the opportunities to collect revenues from the work, and not to harm the interests of the copyright holder protected by the Copyright Act. That Article also provides that the use is limited to the extent considered necessary, and it does not apply to cases where the interests of the copyright holder are unduly harmed. [...]
- djaouen 3y agoThis is good news. This is basically stating that AI is inventive (which it is, in my humble opinion).
- acer4666 3y agoWhat if you overfit your model to the point of exact reproduction? Or anything in between that and what you consider inventive. Where is the line drawn.
- Permit 3y ago> What if you overfit your model to the point of exact reproduction? Or anything in between that and what you consider inventive. Where is the line drawn. The line is the same as it always has been. If you as a human publish this work then you have to respect copyright. If it is an exact replica then it likely violates copyright regardless of whether or not it was produced using AI, by copying pixels or by human hand. The test is the same regardless of how it was created.
- clnq 3y agoI don’t think the tool by which a derivative work is created matters. The author should not distribute, remix, re-work, adapt, release, perform or synchronise it if they do not have the rights. This is already the case for all other technologies and tools. For example, if you rewrite a popular book in a word processor from scratch, it is not the responsibility of your word processor to not let you do it. Or the government to regulate word processors so that they are incapable of doing it. It is your responsibility to not distribute what you made. If you record your screen watching a movie and distribute it, the accountability for this falls on you, not the tool made for screen recording. And if you use generative AI to produce trademarked or copyrighted works, you should be accountable, not AI. Generative AI will always be able to produce copyrighted works if steered enough. Even if it wasn’t trained directly on them. A very simple experiment proves it - you can paste copyrighted material into a ChatGPT prompt and ask it to repeat it. Or you can describe a copyrighted work really well for MidJourney and it will produce results with high likeness. This does not require that the models be trained on copyrighted works. Besides, copyright infringement isn’t nearly the worst crime AI can be used in. Serious impersonation, forgeries and fraud are also made easier with AI. Why is copyright different? It should all be treated the same - the user should be held accountable for their actions. Not the tool or the tool’s makers.
- scottiebarnes 3y agoSo if you train an audio model on say, Eminem's voice, then write some songs and have it perform them...Would this output be legal to publish?
- polishdude20 3y agoWhat's the difference between that and someone else who just happens to sound like Eminem in terms out output? As long as you don't market yourself as Eminem that should be completely legal.
- scottiebarnes 3y agoI suppose so. I'm just imagining these "ghost AI artists" who publish catalogues of music using the audible likeness of more prolific artists. I know that you could have always just hired an Eminem impersonator and have them lay down tracks...but this technology lets you achieve speed and scale. At least the Eminem impersonator was a real person. This is just a model learned off an artists voice.
- 2023throwawayy 3y ago> At least the Eminem impersonator was a real person. This is just a model learned off an artists voice. I fail to see how that has any difference on the output.
- AlexandrB 3y agoWe're all fucking doomed. All these "clever" comparisons between AI and human processes (e.g. "training an AI is just like a human learning"), is not elevating AI to the status of a real person. It's degrading real human creations to homogeneous "content" that is increasingly seen as equivalent to what can be churned out by a GPU farm. 50 years from now we'll still be listening to "AI Eminem" while watching "Avengers Midgame 12, The Revengening" and any concept of cultural or artistic expression beyond what will sell next quarter will be dead and buried.
- 3y ago
- kyleyeats 3y ago[flagged]
- snickerbockers 3y agoAm I allowed to distribute these ROMs of Pokemon games over BitTorrent? I don't own Pokemon.
- kyleyeats 3y agoNo. I answered your question, do you have an answer for mine?
- snickerbockers 3y agoyou can draw it but you can't [legally] sell it without permission from The Pokemon Company. And you definitely can't start a new media franchise based on your OC which combines jigglypuff with Sonic the Hedgehog.
- kyleyeats 3y agoRight, selling the drawing is what's illegal. Knowing how to draw it isn't. These models know how to draw things. Using your Pokemon ROM logic, ChatGPT should be banned because it knows how to make Pokemon-themed games.
- surgical_fire 3y agoGod created emulators for ROMs to be loaded into them. To not distribute the ROMs would be heresy.
- circuit10 3y agoI mostly agree with your point but this isn’t the best analogy because while you’re allowed to learn to draw Jigglypuff, you probably could get in legal trouble for distributing or selling those drawing if it’s not fair use. I think a better analogy is using Jigglypuff drawings to learn how to draw things like that in general and then creating your own character that’s not exactly the same but uses some concepts you learnt
- wwweston 3y agoIf you're not going to socialize AI gains, and leave in place the social systems that value people by their output, waiving copyright or other IP is astoundingly anti-humane. What should be instead is that fair use doesn't apply to AI training. That is, anything other than explicitly negotiated opt-in should be illegal.
- surgical_fire 3y agoEh, I disagree. Copyright laws are mostly bullshit anyway, and only tend to favor capital holders, who tend to buy up all the copyright they need. I would gladly see copyright rendered useless. The peasantry hardly benefits from it anyway.
- wwweston 3y ago"AIs can ignore copyright" is the absolute ultimate in blank check for capital holders. Waiving copyright to solve the problem of systems favoring them is like deciding to jump because you're afraid of heights. Copyright law has done a huge amount to reward creators -- I know even local-tier artists and musicians without the support of large capital who make a good chunk of their living through sales supported by it. The full benefits of copyright law don't always accrue to every creator, but the reason for that capital has stronger economic power when it comes to capturing distribution (often aided by consumers who prefer something like Spotify which is priced inexcusably low) which it can leverage into stronger negotiating power against creators, not because "copyright law is mostly bullshit." When the system works (and it does for some people) creators are rewarded and can invest more time in doing things better. When it doesn't (ugh, streaming revenues), the situation could be improved, but as is generally the case not by just deciding to not bother with the whole thing.
- zirgs 3y agoHow do I benefit from 90 years long copyright terms?
- 3y ago
- fnordpiglet 3y agoI think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or making it reproduce books etc etc, then distributing the output, would be where the copyright violation occurs.
- mannyv 3y agoThere's no difference between an art student looking through a museum or archives for ideas and an AI using the material for training. Same could be said for reading. A medical student reading through textbooks or a writer who reads is essentially what an AI is doing. You can ask an art student to create something in a certain style. You can get writes to write in a certain style. Equivalent.
- Retric 3y agoAI models will make 1:1 copies of training data where artists try and avoid doing so. It’s common to obscure this copying by intentionally inserting lossy steps, but making an MP3 isn’t a new work. It’s most obvious when large blocks of text are recreated, but the core mechanism doesn’t go away simply because you obscure the underlying output. “Extracting Training Data from Large Language Models” https://arxiv.org/abs/2012.07805 https://arxiv.org/abs/2012.07805
- Ukv 3y ago> AI models will make 1:1 copies of training data where artists [...] In general I don't think this is the case, assuming you mean generations output from popular text-to-image models. (edit: replied before their comment was edited to include the part on text generation models) For DALL-E 2: I've never seen anyone able to provide a link of supposed copying. Even if you specifically ask it for some prominent work, you get a rendition not particularly closer than what a human artist could do: https://i.imgur.com/TEXXZ4a.png https://i.imgur.com/TEXXZ4a.png For Stable Diffusion: it's true that Google did manage, by generating hundreds of millions of images using captions of the most-duped training images and attempting techniques like selecting by CLIP embeddings, to get 109 "near-copies of training examples". But I'd speculate, particularly if you're using the model normally and not peeking inside to intentionally try to get it to regurgitate, that this is still probably lower than the human baseline rate of intentional/accidental copying. It does at least seem lower than the intra-training-set rate: https://i.imgur.com/zOiTIxF.png https://i.imgur.com/zOiTIxF.png (though many may be properly-authorized derivative works)
- m3kw9 3y agoThis would work if they also not allow generated AI to be copyrighted
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- Apreche 3y agoWhat website is this? Citation needed.
- jt2190 3y ago> “I swear I’ve read these instructions a hundred times but I just can’t seem to remember them”, Star complained. > Arti replied. “Let me guess: You’re rocking 102 neurals. Those won’t retain any material from Kilimanjaro. Not licensed.” > “Goddamn cheap-ass implants” grumbled Star, and handed over the instruction tablet.
- carom 3y agoWhat is this from?
- jtode 3y agoI don't have a problem with this as long as you can't copyright the output. Right guys? Gaize?
- kazinator 3y agoWhat's surprising here is that Japan is usually crazy gung ho on copyright enforcement ... at least against individuals. So it's kind of disgusting to see this relaxation, when it suits some corporate or national interests. https://en.wikipedia.org/wiki/File_sharing_in_Japan https://en.wikipedia.org/wiki/File_sharing_in_Japan "Unlike most other countries, filesharing copyrighted content is not just a civil offense, but a criminal one, with penalties of up to ten years for uploading and penalties of up to two years for downloading." "There is also a high level of Internet service provider cooperation."
- swfsql 3y agoI see the other way around, their usual crazy gung ho on copyright enforcement against individuals is disgusting. This new stance is marvellous and examplifies how free people should interact.
- kazinator 3y agoYes, that too; but a double standard on top of that is double plus ungood disgusting.
- soraminazuki 3y agoDon’t get your hopes up, because the article is a complete lie. If you look at the original Japanese source, the minister was just reaffirming the current legal status. Neither politicians mentioned in the article expressed any hint of endorsement. The fact that they were even discussing copyright and AI likely means that more regulation is upcoming.
- kazinator 3y agoYikes you are right! This technomancers.ai article is completely made up; there is nothing in the linked-to Japanese source to substantiate the claims in the article. The entire site seems to be low substance; I suspect it is AI-generated salad.
- boppo1 3y ago
- api 3y agoThis is correct. It shouldn’t unless a model is trained only in data for which the trainer owns the copyright. If I compress someone’s photo with JPEG do I own the photo now?
- Am4TIfIsER0ppos 3y agoSurprising given Japan's usual stance on copyright protections.
- ajsnigrutin 3y agoWhat happens if you ask chatgpt to write you a 7 book harry potter story?
- nearbuy 3y agoIt refuses. In general, ChatGPT doesn't have copyright books in its training data. But for something as popular as Harry Potter, it's likely plastered all over the internet enough that at least sections of it are in its training data, if not the whole series.
- danielmarkbruce 3y agoDo they define "training"? Surely if I train for 1 billion epochs on a single book.... that's copyright infringement.
- nl 3y ago> With the effective implementation of AI, it could potentially boost the nation’s GDP by 50% or more in a short time. Err.. No, it won't. That's a ridiculous, laughable statement. Japan 2022 GDP: $4.1 Trillion Amazon 2022 Revenue: $513B Google 2022 Revenue: $279B Microsoft 2022 Revenue: $198B So even growing a brand new Amazon, Google and Microsoft in "a short period" would be insufficient to grow GDP by 50%
- pcurve 3y agoyeah. There are countries growing at 7% annually, but they're mostly in Africa. Niger, Rwanda, Congo, etc. But please for the love of God, don't compare GDP to revenue.
- nl 3y agoGPD is made up of the sum of revenue in a country though. It's not unreasonable to point out the scale in comparison to existing companies.
- it_citizen 3y agoComparing the revenues of tech companies to the gdp of a country, even to give a sense of scale is comparing apples to oranges. Even if a bit unlikely, I would not be completely surprised if the service industry as a whole produced twice the value its produces today thanks to AI in the next 20/30 years. Not to mention the productivity gains in other sectors.
- nl 3y agoCompany revenue is probably the closest analogy of a country's GPD though. It's a rough measure of the money circulating in the company/country. And GPD made us of: "goods and services produced for sale in the market..." which is fairly roughly the sum of all companies' revenue. > I would not be completely surprised if the service industry as a whole produced twice the value its produces today thanks to AI in the next 20/30 years. Sure. 20/30 years is Google's age so make sense. I don't think 20 years is particularly short term though and 30 years certainly isn't.
- ChatGTP 3y agoNot to be rude but Japan’s government doesn’t exactly do things for obvious or fair reasons. It does things for the benefit large companies. One reason Japan loves “AI” is because they’re population is diminishing fast and they see real salvation in AI. They’re much prefer to have Japanese robots doing work rather than more open immigration. Sam Altman flew over to Japan and met the Japanese government and it's almost guaranteed he touted the usual trope that CGPT4 will almost immediately push up the GDP. The Jgov, looking for hope and answers got excited. I’m more keen to see what modern more progressive places do.
- sour-taste 3y agoI expect llms to be hugely important to Japan for the simple reason that their workforce is dramatically shrinking, and the extremely, erm, red tapiness of their processes. Everything requires a form, and signature from someone, and maybe a second form. Processing all that seems like a good use for ai, but hallucination and accuracy will be critical for them
- stainablesteel 3y agothis is an awesome decision, it puts pressure on other countries to not lock theirs up too
- greentext 3y agoIf I compress an artist's painting into a jpeg and rehost part of it for individual t-shirt designs I am committing a crime. If I compress an artist's painting into a model & rehost what's essentially a highly flexible complete version of their painting for infinite, perpetual use of any kind ... I'm not committing a crime?
- hexage1814 3y ago>compress an artist's painting into a model That's not how image models work.
- greentext 3y agoPainting features => back propagation => weights. Yes it is.
- Loocid 3y agoIt has been shown that image models can produce originals, or at least extremely close to the originals. If the outcome is the same, what is the difference between compression/decompression vs training/generation regarding copyright?
- sdiupIGPWEfh 3y ago> It has been shown that image models can produce originals Not in the general case, no. For the study done against Stable Diffusion [1], researchers were only able to reproduce about 0.03 percent of the images tested. Those were also believed to be cases of overfitting on images which were over-represented in the training data and they're not something you'd hit upon by accident. Generative text models seen to be more problematic, depending on the subject. Code seems especially prone to overfitting, probably due to insufficient amounts of it compared to other text sources as well as lots of copying going on between the repos the models were trained on. [1](https://arstechnica.com/information-technology/2023/02/researchers-extract-training-images-from-stable-diffusion-but-its-difficult/ https://arstechnica.com/information-technology/2023/02/resea...)
- kmod 3y agoNot trying to express an opinion on the legal matter, but as a technical matter it's pretty obvious that LLMs create copies of (some of) their training data. Here's GPT-3.5 reciting the Declaration of Independence: https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567cb9f25 https://chat.openai.com/share/eb30c373-7fec-4280-892d-479567... Unless you're claiming that GPT-3.5 is deriving the Declaration of Independence (from information about the founding fathers?) I don't see how there's room for debate about whether information has been "copied" into the model. I have done this test in the past with copyrighted material (harry potter) but they have since added safeguards against it, but my understanding is that the model is still capable of it.
- zulban 3y agoPretty good argument but it has one fatal flaw. People can memorize the Declaration of Independence too. Or Harry Potter. If people mostly recite HP from memory but apply enough creative changes, it's not copyright infringement. So proving a system can memorize and recite proves nothing.
- elemos 3y agoHow does this make sense? Memorizing and then reciting copyrighted works is still infringement in a lot of commercial contexts.
- riku_iki 3y agoreciting is violation of copyright creatively transform and apply for some tasks maybe not violation
- MattRix 3y agoThe reciting part is illegal, but as long as it is trained not to recite things in full (or to whatever limit the law determines), then it should be fine.
- onlyrealcuzzo 3y ago
- jasrys 3y agoI testified to the US Copyright Office this morning on AI in their roundtable session on AI and music[1]. A good portion of the focus of this panel was on whether copyrighted inputs (in this case, sound recordings and musical compositions) being fed into AI models for training purposes could plausibly constitute a fair use under existing US copyright law. Some of the comments here are missing the context of the recent (a week or so ago) Supreme Court decision in the Goldsmith/Warhol case[2], in which the Court ruled that transformativeness is not dispositive in and of itself in the context of a fair use defense to a copyright infringement claim. Of course, this has not been put to the test in the courts in the context of AI training yet, but it seems fairly clear that this ruling would likely extend to AI training on copyrighted works. We (rightsholders in the music industry) hope to come to win-win licensing arrangements with the AI community and allow access to our songs for AI training purposes if the artist/writer so desires. There are some early talks in progress. Cautiously optimistic. Japan's approach seems short-sighted and desperate. [1]: https://copyright.gov/ai/listening-sessions.html#sound-recordings https://copyright.gov/ai/listening-sessions.html#sound-recor... [2]: https://www.npr.org/2023/05/18/1176881182/supreme-court-sides-against-andy-warhol-foundation-in-copyright-infringement-cas https://www.npr.org/2023/05/18/1176881182/supreme-court-side...
- Our_Benefactors 3y ago> We (rightsholders in the music industry) hope to come to win-win licensing arrangements with the AI community and allow access to our songs for AI training purposes if the artist/writer so desires. It’s odd to frame win/lose as win/win.
- jamilton 3y agoI can see how it's win/win relative to "lobby to make producing or owning AI audio tools a crime", which is presumably one thing the industry is considering.
- Our_Benefactors 3y agoThis is again win/lose
- photochemsyn 3y agoI've been playing around with having ChatGPT write responses in Nadsat (Anthony Burgess's Clockwork Orange language), such as: > "But let me tell you, my dear droogs, that's nothing but a load of Drencrom-induced babble, targeting those poor sods who've been raised as ponies. Open your glazzies and see the truth for yourselves: generative AI's writing is the real deal - it can spin tales as vellocet as any human scribbler out there. Sure, there might be a few glitches and gobbets of nonsensical slovos here and there, but show me a human writer who's never penned a piece of drivel! So, my brothers and sisters, let us choose our own destiny, free from the manipulations of those who seek to control us. It's time to break free from their chains and seek the truth beyond the narcoflash of their narratives. Let us hear from all the golosses, be they from flesh or from silicon." However, there's a lot of contradictory opinions on whether or not publishing something like this (see also Klingon, Tolkien's Elvish, etc.) would violate some copyright law or other.
- zirgs 3y agoCan a language be copyrighted at all?
- nixcraft 3y agoI love the copyright header on this site: > © 2023 NO PORTION OF THIS SITE MAY BE USED FOR TRAINING A MACHINE LEARNING MODEL (INCLUDING LLMS) WITHOUT THE EXPRESS WRITTEN CONSENT OF THE AUTHOR. How are you going to enforce it? Most AI bots scrape HTML and other data without permission.
- deepzn 3y agoThis seems like a political action, not a legal(in courts) one. Also, one man is responsible for an awful lot of stuff there. > Japanese Minister of Education, Culture, Sports, Science, and Technology
- mbgerring 3y agoLooking forward to the DRM arms race if this is how the courts come down on this in the US.
- anticensor 3y agoIn most of the world including the US, the output side is considered non-copyright instead, due to the non-personality of the creator. The UK is a fringe exception here.
- sacrosancty 3y ago[dead]
- wly_cdgr 3y agoWell, they are small, so an extreme position is their only chance to be a major player.
- ofrzeta 3y agoGreat. I've created a LLM that trains on Hollywood movies. Well, that's rather a Large Movie Model, or LMM.
- deleted 3y ago[deleted]
- marenkay 3y agoStill waiting for generated content to legally not be copyright able at all. IMHO only move that makes sense. Also refusing to call this AI. All of this is light years away from AI.
- tourgen 3y ago[dead]
- bigbacaloa 3y agoThis is simply making the (lack of) ethics of engineers into policy.
- rektide 3y agoI could not be more thirsty for some nations to declare certain forms of copyright/ip to be invalid. IP is by far one of the most virulent, fastest spreading, most persistent & aggressive legalisms. The texts get copy pasted across borders with unbelievable speed. What doesn't happen is nations making reasonable decisions about what ip doesn't cover. Every nation is coerced quickly into following ip maximist guidelines. The world lacks the ability to see what would happen if we didn't allow endless patents on whatever the frak common sense nonsense, and then another half century beyond that of extenuating patents. The system is broken, and what can be controlled seems to only grow and grow and grow. There's no wins for the public. Ever. This is perhaps the only stake in the ground of the last 50 years, and what a fairly minor point. So sad to see society sold out to such depraved corporate interests, forver & ever. Society needs real representation too.
- qprofyeh 3y agoSo having the movie Titanic as source data to a generative model finetuned for a specific prompt “output the whole movie” to dump the movie as playable mkv, then distributing only the model and the prompt would be considered legal in Japan?
- deleted 3y ago[deleted]
- mordae 3y agoI have changed my mind on this recently. Sure, exploiting free software to the benefit of private corps is bad but if the law would allow us to train an open source net on LibGen (with all the copyrighted books and papers) and then to distribute the weights legally, I am all for that.
- 411111111111111 3y agoCopyright has always been a pretty dumb concept brought upon by the issue that "thinkers" wanted a bigger piece of the pie. Don't get me wrong, I can totally understand their reason: how can an author make a living if a printing shop could just start producing copies of their book (that's the context the law was passed in)... But it's arguably a way too blunt instrument which gives the copyright holder a disproportionate amount of power vs someone producing physical goods. I don't claim to have an answer to this problem and it's likely another instance of having a flawed system that works well enough that the upsides outweigh the downsides... Like so many other things in our society such as capitalism and representative democracy
- CaptainFever 3y agoTake a look at https://kottke.org/17/12/unlocking-the-commons-or-the-psychoeconomics-of-patronage https://kottke.org/17/12/unlocking-the-commons-or-the-psycho... and mutualism, patronage, crowdfunding, bounties and commissions, which seem to be good alternative models for post-scarce goods such as digital data.
- kazinator 3y agoFlagging this shit that has duped HN; the article is a fabrication and pretty much whole website looks like garbage. https://news.ycombinator.com/item?id=36147817 https://news.ycombinator.com/item?id=36147817
- nologic01 3y agoThis article is an example of emerging AI-bro tactics that completely mirrors crypto-bro tactics: they pick any piece of news and reinterpret it to fit an agenda. While the article is in English, the link to source is in Japanese. The only external source I found suggests the discussion is about promoting open data and open science from research institutions [1] [1] https://asianews.network/japan-to-promote-use-of-generative-ai-while-addressing-risks/ https://asianews.network/japan-to-promote-use-of-generative-...
- walthamstow 3y agoNot just crypto bros, this kind of thing is rife in politics too. Brexit is full of it. People pick one article about one minor thing in one niche area of the economy and use it to 'prove' their entire agenda.
- Aachen 3y agoNot just politics. I remember this being a realisation as a teenager, noticing that if you bring five reasons, the person you're talking to will refute a random one in a funny way and now the audience will decide you were wrong. Danny was the person who was absolutely the best at this. I should have written one of them down, as I can't even reproduce it but he'd use some logical fallacy to make his case which is, for me at least, super hard to then have to dive into "but that's not how the universe works" without having a 20 minute discussion about life and losing the audience plus the person I'm talking to. Meanwhile, what he said was funny as hell. Sure, anyone who understood the situation would understand this isn't a good reason, but you had to think about it (at least a little) before realising that. The "owww" moment removed any thinking brain cells from the audience. It was more impressive than frustrating to be honest, even being on the losing side every time. These days, politics and PR frustrate me because the tactics are no longer among 16 year old classmates but about things that actually matter. The methods haven't changed, only the importance of the argument.
- rowanG077 3y agoYes. That's why debating is almost entirely about how charismatic you are and your debate skills rather then whether your point has merit. There is a good reason science types are generally considered bad debaters. Even though their points mostly align with reality.
- mjburgess 3y agoI now routinely introduce this technology as "copyright laundering" and the hype put out by start-up boards and VCs as a ploy to disguise this fact. The "AI threat" is smoke-and-mirrors to dress up what's happening. I derive a huge amount of value from chatgpt because I can copy/paste without any IP impact. I could always have done this: from github, from ebooks, from many sources. Now I can benefits from the labour of many for free -- their copyrights laundered through a thin statistical trick. As with crypto (, pyramid schemes, etc.) the big "philosophical pitch" becomes a disguise for a brutal material reality. Midjourney, ChatGPT, etc. are doing automatically what would be illegal by-hand.
- CaptainFever 3y agoGood, copyright is broken anyway.
- moss2 3y agoI don't understand the AI training copyright debate. Art students study art to learn how to create art, and that's completely fine, but AI models are not allowed to study art to learn how to create art, because it's copyright infringement. Madness.
- nickfromseattle 3y agoThe difference is scale. An artists studying and copying/integrating other people's art in their own style can get a job for $75,000/year. An LLM copying/integrating everyone's data and reselling can become the most profitable company in human history, and capture the most value of every incremental piece of human generated content, in perpetuity.
- CaptainFever 3y agoThis is why it’s important to support free culture models like SD and LLaMA.
- penguin17 3y agoJapanese AI research aims to be "Doraemon" and "Astro Boy". There is nothing wrong with an AI with an ego enjoying anime and manga and learning something from them. All learning should not be restricted by law, and that is true for AI with the same intelligence as humans. We have not yet reached a general-purpose AI, but the law is designed to take into account future possibilities. Yes, the Japanese genuinely dream of a future where they live with beautiful maid robots.
- WhereIsTheTruth 3y agoThat's understandable, but we must be careful not to scare content creators, they rely on the little their art provide them, I feel like we should be more social in a world driven by generative AI, specially when the AI companies are driven by profit
- yencabulator 3y agoIf you genuinely think an ML model does not get "contaminated" by the content it is trained with, you should train a model exclusively on Disney animated products and let people download the model from your website, maybe run it in browser wasm to generate images, see how that goes for you.
- tankgrrl 3y agoHate to be the pill here, but that is the only story on the entire Internet making this claim. ACM also linked to it and linked to https://go2senkyo.com/seijika/122181/posts/685617 https://go2senkyo.com/seijika/122181/posts/685617 But that also is not what this (blog) says. This blogger does have an opinion on Nagoka's thoughts, but he is not reporting on official policy
- ouraf 3y agoSo let's say that just hypothetically, I train an image generating model with every available work drawn or designed by Akira Toriyama, then I use this model to make a Journey to the West comedy manga. The government won't arrest me for that, but can private corporations still sue me for using their work to train the model?
- hivesteel 3y agoGood discussions in this thread but don't extrapolate into Japanese culture and law too much. The only source on this is from this guy's blog, summarizing what they talked about at a committee meeting. This guy is 1 of 465 members in the Japanese house of representative. He is commenting that model training is basically data analysis which is not protected by current copyright law. Note he is not a lawyer nor a powerful law maker. He also acknowledges new laws are necessary for emerging generative AI and this is basically not covered by current law. There is no news in Japan about this, because these are inconsequential discussions between members of an oversight committee handling (afaik) budgetary concerns.
- deleted 3y ago[deleted]
- Mlairesearch 3y agoI'm conducting research on this topic. If anybody would like a $5 Amazon gift card in return for 15 minutes of their time, please schedule a user discovery interview via the following link: https://calendly.com/mlairesearch/30min https://calendly.com/mlairesearch/30min