25 ms·
AI weights are not open “source”
- meindnoch 3y agoAccording to whom? Weights are a type of program, which are interpreted by the neural network runtime. Same as Java bytecode interpreted by the JVM runtime.
- adamnemecek 3y agoJava bytecode is not "open source". At least for Java bytecode there are decompilers.
- eigenket 3y agox86 machine code is a type of program, which is interpreted by the processor, but distributing the binary of my program doesn't make it open source.
- kfarr 3y agoBingo, did a ctrl+f to find binary as that seems like the closest analogy here.
- slowmovintarget 3y agoWeights are data, not a type of program. A computer program is a set of instructions that may be executed. Weights are values that may be loaded by a program, but are not a program in and of themselves.
- jstanley 3y agoIt's a very difficult distinction to make. Would you consider a Python program to be data rather than program just because it is text input to the python interpreter instead of machine code for the CPU?
- slowmovintarget 3y agoIt is not at all a difficult distinction. Weights are literally numbers computed as output. They are not instructions. The semantics of those numbers even when emplaced (loaded) in an artificial neural net is such that they do not execute. They are not instructions. LLM engines and diffusers perform searches where the weights are used to calculate additional output. Is source code, like Python text, data? Yes. All code is data. But not all data are source code. If I gave you a web request log, you would not assert it is a program. If I gave you a CSV file with time-series values from a sensor, you would not assert it is a program. If I hand you a database of contact information, you would not assert it is a program. Weight files are the equivalent of CSV files. They are are a dump of parameter values computed from training. They are not a program. The definition of computer program is well worn. So is the definition of source code, and the definition of parameters. Weights are parameters.
- jstanley 3y agoThe difference between code and data only exists in our minds. There is no distinction. Both code and data make the computer do things (and, yes, both code and data only make the computer do things if other conditions are permitting, for example if executed with the right interpreter, or loaded with the right type of viewer). Anything that can be expressed as code can be expressed as data, and vice versa.
- xigoi 3y agoIf a program has to consist of instructions, then source code written in a declarative language is not a program.
- pravus 3y ago> Weights are literally numbers computed as output. They are not instructions. They are instructions if you consider the LLM system itself to be a kind of weird, indirect virtual machine. Each number can be mapped to a set of instructions that are executed. Even your CPU uses numbers (machine codes) to execute. Join me in saying: ...code is data is code is data is code is data...
- graypegg 3y agoThey’re not data though, they’re coefficients. They are the only thing that significantly differentiates one model from another. If I told you the economy can be accurately modelled by GDP(x) = Ax + B But I don’t define A And B for you because it’s proprietary, you haven’t learned anything other than what you can glean from the structure of the model itself (it’s linear, there’s only a single input etc) If most of these models are similarly structured, I’d say the weights are the program.
- slowmovintarget 3y agoThe nature of the data as proprietary or not, important or not, is not relevant. Parameters, or actual arguments, are values; data. Not instructions. Valuable data is still data. It's significance doesn't magically turn it into source code.
- golemotron 3y agoNo, declarative programs exist. They are not instructions. There is no real line between code and data. This is an observation that runs all the way from Turing Machines in computability theory to the Von Neumann architecture and homoiconicity in Lisp. What we call 'data' is just code that needs a cleverer interpreter.
- Izkata 3y agoLess into theory and more into "wait wtf": Some of the older projects I've worked on were written by people who loved database-driven stuff, to the point they did things like put perl code into one table column (with sentinel values you had to find/replace before `eval`ing the code) and sql into another table that retrieved values for those find/replaces, both retrieved and executed by some really generic code. Code or data: Well... both.
- killjoywashere 3y agoBut not "raw" data. They are derived from other data and a program. If this was a collaboration where one collaborator did the processing and one sourced the data, they would likely both claim some amount of ownership of the trained weights. At a minimum, it would be an active area of negotiation that the attorneys would take notice of. Source: have negotiated these agreements.
- slowmovintarget 3y agoA curated data set is still a data set. I imagine it is not settled law, but there's a clear argument to be made that regardless of the difficulty in curating the data set, it's still a data set. Can it be licensed and sold. Yes, surely. Is it proper to pretend an open source license is sufficient protection, probably not.
- programmarchy 3y agoThis is a distinction without a difference. Code is data and data is code.
- mrguyorama 3y agoMaybe they've only worked with machines using a Harvard Architecture
- slowmovintarget 3y agoAll source code is data. Not all data is source code. Data may be encoded, but that doesn't make it source code either.
- earleybird 3y agoWeights are data in the same way that instruction codes in memory is data.
- slowmovintarget 3y agoValues for the variables do not the function make.
- daniel-cussen 3y agoThat's the least of it. In Lisp the distinction between code n data is blurred all the time. In F18 assembly i frequently have "double entendres" which are used as code or as literals depending on the entry point. I think at least once there was code and data in the same entry point. Assembly n Lisp are both homoiconic, after all. N verb at the end of the sentence, are you transliterating German, or a two-foot green Jedi master full of wisdom?
- rockinghigh 3y agoWhen people talk about weights, they talk about a network of weights that takes an input and computes an output. There is really not much difference between a saved model and a program.
- tensor 3y agoI think the point here is that by being explicit you avoid the need to have this argument.
- cdelsolar 3y agowho wrote that program?
- thepangolino 3y agoI've always seen weights as akin to configuration files.
- Zetobal 3y agoIf my own data is in the dataset even when I didn't give consent is it a collaborator dataset?
- tensor 3y agoIf you posted your data into a service where the TOS allows this use then yes.
- iLoveOncall 3y ago[flagged]
- bee_rider 3y ago> The ethical license category applies to licenses that allow commercial use of the component but includes field of endeavor and/or behavioral use restrictions set by the licensor. I don’t love the name, “ethical license” sounds like a description of the license: this license is ethical. Really this sort of license imposes a particular ethical framework on the user. Not to throw shade, though. It is actually hard to come up neutral sounding name for this sort of license I think. I keep thinking of things like “morality encumbered license,” but that sounds ridiculously euphemistic in a weird way.
- iandanforth 3y ago"Opinionated" is how I think about it.
- bee_rider 3y agoThat might be a good pick, IMO the word has negative connotations elsewhere, but in tech circles seems basically neutral.
- version_five 3y agoYes I was going to say the same thing. It's a branding that has been applied by the license's proponents, and I personally reject a lot of what they call "ethics" as well as the idea of whatever monitoring and enforcement the restrictions entail - maybe calling it a religious license would be better.
- 93po 3y agoI'd argue any licensing of IP is unethical. I'd use the word "conditional"
- mellosouls 3y agoHmm. Makes a few unsubstantiated claims, with hand-wavy appeals to risks that our private corp overlords are presumably protecting us humble users from, now that they've built their product on open source and data by closing it down and changing terminology to suit. There's an intelligent discussion to be had, and I think this otherwise-reasonable article could be part of it if it toned down the presumption and condescension a little.
- ianbutler 3y agoI'm not sure OCV gets to decide any of this. Just like I don't think OSI trying to be the sole dictator of the term "Open Source" works out long term. My opinion is always received controversially about things like this, but terms evolve to meet the common usage by the people. If people are calling this "Open Source", and there are more people who want to call this "Open Source", than people who don't; unless you intend to legally bar them from using the term, with actual action, like a lawsuit or something then eventually this will will also be encompassed by the term "Open Source" as people know it like it or not. Yes I know this term is currently defined explicitly by OSI, no I don't think language prescriptivism wins out regardless how hard they try with it, and since I haven't seen any of the hundreds of quasi Open Source, but not really, companies get dragged to court over usage of the term, this is all toothless complaining in my view. As to their actual point, I might actually agree with them if it were only the weights being shared. In most cases the configuration is also shared which allows popular frameworks to instantiate the model and then execute it for either inference or further training making the release fully suitable for modification and rerelease. I don't need the exact implementation of FlashAttention they used if I can load the model into Huggingface and use theirs, or mine or whatever. Edit: This obviously doesn't apply to the models who have restrictions placed on usage just in case people think I mean every instance of sharing a model. Those are obviously restricted use and I agree it muddies the term.
- barbariangrunge 3y agocompletely off topic, but funny: I misread "opencoreventures" as "opencorevultures"
- Makhini 3y agoFunny
- ronsor 3y agoModel weights are not source code, but data. Arguably because of how they are generated, they are not even copyrightable at all.
- adamsmith143 3y agoCorporate data is of course protect-able. Otherwise why don't you just open up all your databases so anyone can access them?
- DannyBee 3y agoThey are only protectable by copyright you the degree they are creative works of authorship. Copyright is not usually how these are protected.
- WrongAssumption 3y agoCopyright protection is what gives protection when you put something out into the public. The desire to not publish something is evidence against having these protections, because people know they are not copyright able, so for that reason and others they keep it private. You just presented evidence against your position.
- ftxbro 3y ago> "Model weights are not source code, but data." OK but I mean it's functionally a kind of machine code for a strange machine with a neural transformer architecture, like a 'binary blob'. It's outside of the paradigm where machine code is created only by compilation of copyrightable source code written by humans following their creative "aha moment".
- Makhini 3y agoWhat if you change the weights slightly? Kaboom, not breaking the copyright anymore.
- bskap 3y agoThen it's a derivative work and copyright law covers that too.
- jkeisling 3y agoThe article makes a good point: we should prevent “open-washing” and draw a distinction between well-intentioned restrictive licenses like “Open”RAIL and true open source. However, I worry the name “ethical source” is itself a bit question-begging. While outfits like Bloom may believe in good-faith ethical principles, their definition of ethics isn’t necessarily everyone’s. If restricted models are “ethical”, is releasing open weights “unethical”? Conversely, is releasing a model with PII or artist styles in it “ethical” if a few known use cases are forbidden? There’s no one right answer. Labeling any one set of restrictions as “ethical” off the bat makes discussion harder and puts open source on the back foot to justify “not being ethical”. Better to just call them “restricted models” or “guarded models”, and leave it to individuals to decide if these restrictions are beneficial or not.
- A4ET8a8uTh0 3y agoI think the more interesting aspect of all this is that the confusion created by this new business model ( not sure to classify it so business model had to do ) appears to be largely intentional. The subject matter is complicated to begin with experts being niche of a niche of a niche and the assumption that the general public can even understand it ( and whether it can even dumbed down to digestible sound bites ) is, in my mind, very optimistic. Now, courts are not typically stacked with dummies, but again how many are well versed in issues of technology? All in all, I don't disagree with the point you raised, but I worry that all this will only further muddy the water for the general population.
- pmoriarty 3y ago"Now, courts are not typically stacked with dummies, but again how many are well versed in issues of technology?" Even if they are well versed in issues of technology that does not mean they'll make what any given one of would consider a good decision, as plenty of people well versed in issues of technology disagree with each other on these issues. Nothing guarantees that on, on any issue, really, as you can always find people who disagree.. and if they happen to be judges, they get to decide unless another higher judge overrule them.. and that judge has the same problem as the first.
- tiffanyg 3y agoAI licensing is extremely complex. Unlike software licensing, AI isn’t as simple as applying current proprietary/open source software licenses. AI has multiple components—the source code, weights, data, etc.—that are licensed differently. Are you joking? This isn't wrong, per se, but it's worded as though written by someone with only the most casual / cursory interaction and knowledge of this area of law / commerce (e.g., including licensing, copyright, trademark / service mark, patent, etc.) ... until perhaps quite recently. Yes, the AREA IS complicated. No, so-called "AI" is not introducing all sorts of novel issues, structures, etc. "AI" has some nuances distinct from much of what has come before (happens basically every time more significant tech comes along) and some possibly more unique questions related to economics, ethics, philosophy, and the like, but the relevant areas of law and practice have often been complicated and sort of "bleeding edge", even going back before the industrial revolution. Big money, powerful tech, large-scale economic forces, etc. = lots of maneuvering, legislation, litigation, etc. = complicated "rules of the game". Drawing the distinction vs. software in general is reasonable - but, the rather click-baity headline and "I just learned about 'IP' law and bah gawd y'all are doin' it wrong" tone to the start of this article suggest, to me, that this isn't likely to be the best article to use as a reference to learn about these issues.
- larodi 3y agoI was like going to write ‘are u joking’, but you make the same point so well. This article is at best oversimplifying and misleading. Besides I doubt this ‘my weights your weights’ thing is a thing at all.
- cpcallen 3y agoI'm disappointed that the article is only making the (somewhat pedantic) distinction between source code and weights. From the quotation marks in the headline I hoped that it would instead be making the distinction between human-readable source code and machine-readable compiled form. For example, IMHO (IANAL) an AI code-completion tool that had been trained on GPL software is (or should be) only be legal to distribute if it is accompanied by the training code _and all the code ingested during training_ (or an offer to provide such code upon request).
- version_five 3y agoThis is an interesting point. If you read the OSI open source definition, specifically on source code (quoted below) I'm inclined to treat the training data as part of the source code for the purpose of determining whether to consider any model open source. 2. Source Code The program must include source code, and must allow distribution in source code as well as compiled form. Where some form of a product is not distributed with source code, there must be a well-publicized means of obtaining the source code for no more than a reasonable reproduction cost, preferably downloading via the Internet without charge. The source code must be the preferred form in which a programmer would modify the program. Deliberately obfuscated source code is not allowed. Intermediate forms such as the output of a preprocessor or translator are not allowed. https://opensource.org/osd/ https://opensource.org/osd/
- habitue 3y agoOne thing I don't see discussed enough is that, ok let's say the weights are unencumbered, and the source is under an OSI license: the point of open source licenses and free software was to expose the *human understandable* meaning of the final program. That's why distributing binaries isn't allowed even though technically all of the functionality is present in the machine code. AI weights are basically binary blobs. We don't know what they mean, there is really no source code for them. The best we can do is various black box manipulations on them like LoRA, etc, similar to what we can do to a binary blob.
- deleted 3y ago[deleted]
- phkahler 3y ago>> AI weights are basically binary blobs. We don't know what they mean, there is really no source code for them. No. You can do further training on them. If they are something less than code I don't think it's going to warrant all this talk about licensing. GPL, MIT, or some proprietary should cover it.
- habitue 3y agoYou can do further training on them, just like you can patch a binary blob. There are some surgeries you can do to the weights, and there are analyses you can do to poke at them and try to understand them, but ultimately they weren't created from a human understandable spec, and without a ton of reverse engineering work the weights by themselves aren't human understandable: hence the "source" component is missing. The source code that generated the weights is one step removed from the kind of source code we'd need to interpret a bunch of AI weights. It's really meta-source code
- Traubenfuchs 3y ago[dead]
- ndriscoll 3y agoThe complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like a rather important point to bring up when discussing the licensing landscape for weights...
- WanderPanda 3y agoOf course weights are copyrightable. Otherwise nothing is copyrightable
- enlightens 3y agoRecipes, for example, are not copyrightable in the US. Neither are some of the concepts behind creating a fillable form. It's not an all-or-nothing system. https://www.copyright.gov/circs/circ33.pdf https://www.copyright.gov/circs/circ33.pdf
- throwaway98721 3y agoWhy is it unestablished? Is a document not copyrightable based on its contents? Weights are just a different kind of a document.
- Conscat 3y agoNo, a document's contents aren't inherently copyrightable. They have to be a creative work or a method of production, and part of that basically means it has to be human generated content (as opposed to computer or animal generated). AI weights might be considered a method of production, but that isn't clear yet.
- throwaway98721 3y ago[flagged]
- morpheuskafka 3y ago> AI also poses socio-ethical consequences that don’t exist on the same scale as computer software, necessitating more restrictions like behavioral use restrictions There's plenty of software that has, or could have, similar restrictions. Consider software that allows you to plan vantage points for a shooting or estimate the impact of using explosives at various locations. And the government regulates all sorts of software for export/download because it has military use--everything from development tools to high performance chips that could be used to crunch numbers for a nuclear program, CAD software that can help you build (or destroy) a bridge, etc. The CPUs and GPUs themselves are regulated at certain performance levels, I think. None of this is really new to AI.
- FrustratedMonky 3y agoAre the weights in our brain copyrightable? Might want to get ahead of the curve on this one. How would this work? Would I get a tattoo with a license spelling out covering the contents of my body?
- DannyBee 3y agoHas to be fixated (unchanging) and in a tangible medium.
- RobotToaster 3y agoSo I just need to cryogenically freeze my brain in order to copyright it?
- TheRealPomax 3y agoSo, Open Data. Got it. This is the same category as config files that are kept up to date by a program as it runs. - Is it "a program"? Very clearly not. - Is it source code? You can argue either way. The program won't work without it, but "this specific one" is not required for the program to do something, and that ambiguity means you probably don't want to call it "source code" because it's too vague. - Is it data used by a program in order to perform its task? Absolutely. It even uniquely defines the program behaviour, and so is a thing onto itself within the context of the program it's used by.
- worksonmine 3y ago> Unlike software licensing, AI isn’t as simple as applying current proprietary/open source software licenses. AI has multiple components—the source code, weights, data, etc.—that are licensed differently. Software also has multiple components, often the same as the ones listed by the author. But what do I know, to me AI is just another example of software.
- mensetmanusman 3y agoWeights are an information asset that require millions in capital and burned-out GPUs to mine and refine.
- daniel-cussen 3y agoYeah that's my business, http://fgemm.com http://fgemm.com , coming soon. Paper is coming out v soon however.
- TrackerFF 3y agoWeights are just matrices with values between a certain range. So are digital images - just matrices with values. Images are covered by copyright laws, so why shouldn't weights also be?
- zarzavat 3y agoWeights might be copyrightable but in no universe are they copyrightable by OpenAI, Google, etc just because they did the training and spent money on GPUs. The only people who can possibly own the copyright, if any such copyright exists, are the authors of the training data. I find this whole discussion about copyright of weights almost absurd, the incredible amount of deference given to our corporate lords is such that we are “hallucinating” new forms of IP protection for NN weights that have never existed in any kind of statue or case law and cut completely against the grain of all the law that currently exists.
- mikewarot 3y ago>Weights might be copyrightable but in no universe are they copyrightable by OpenAI, Google, etc just because they did the training and spent money on GPUs. I don't see why not. If you took all the same training data, you would not get the same weights. Especially if RLHF was used to tune those weights. The weights are not a set of facts, they are the result of work, sometimes millions of dollars of work. Surely they deserve copyright protection if they are ever "published". If not, then they are a trade secret, and other rules apply.
- adamsmith143 3y agoThe question shouldn't be whether the weights are copyrightable but whether they are protected under other electronic communication/data privacy laws.
- seydor 3y agoIf it is extremely complex, then it can only be modeled by an AI
- Topfi 3y agoThis post did cover many of the same ideas I have been ruminating on concerning model weights and the nomenclature of current efforts. That's also why I generally tend to stick with calling these[0] "local/self hosted models" for the time being. A major reason for my reluctance is that I see weights far closer to binary than code, making a distinction important and current FOSS concepts not really applicable. Of course, this all hinges on the idea that weights by themselves are inherently protected by current copyright, which still seems to be an unsettled topic, hotly debated by both laypeople and legal professionals. Authors generally are afforded copyright on their work by default, and weights raises so many questions concerning authorship that have never been considered. This being such a contested issue, which will require new laws and/or precedent (depending on the legal system), is very problematic. Regardless of where you live, generally courts and government entities are not famous for their speedy reaction to new things, so clarity may take a while, at which point the industry might have already settled on some agreement that then may be adopted as a basis for actual legislation, which would likely favor financially well baked entities already actively lobbying for their interests, such as OpenAI. Some have also pointed out that this is arguing semantics, and I am tempted to agree in principle, but also want to emphasize that I feel this is a situation where that can be valuable. Should weights in some way be afforded copyright protection, clear nomenclature will be needed. Putting some thought into this now is definitely not the worst idea. I very strongly feel that the specific word "ethical" as part of defining licenses is not the best idea, though. "Ethical" can carry vastly different connotations, depending on a myriad of factors, many of which would go beyond the use-focused definition laid out in the post. Due to this, I'd argue for "behavioral" or "restricted use" over "ethical", as both more clearly state what the intended effect is in cases such as Open RAIL-M[1]. Part of my strong feelings on the use of the word "ethical" come from the fact that with weights and training data, there has been a lot of discussion concerning both rights of and considerations for creators whose published works have been used to create those weights. Due to this, the use of "ethical" referring to a group of licenses could give some the impression that this may indicate that the training data used was "ethically sourced", i.e. in agreement with the original creator. This is something that in my eyes should also have clear labeling, though with weights being very hard to reliably trace back to source data, it currently seems impossible to verify, making this essentially just a good faith effort. [0] https://huggingface.co/tiiuae/falcon-40b-instruct https://huggingface.co/tiiuae/falcon-40b-instruct [1] https://drive.google.com/file/d/16NqKiAkzyZ55NClubCIFup8pT2jnyVIo/view https://drive.google.com/file/d/16NqKiAkzyZ55NClubCIFup8pT2j...
- robomartin 3y agoCan someone give me a legal answer to this? People, from early school, all the way up to university, use copyrighted materials to learn various topics and obtain degrees. This trains our brains using the work of others. The same is true as we navigate life. We learn various skills and subjects consuming the work of others. And, yes, in the case of most people, we use that training to pursue various careers, obtain work and get paid for it. How can there be a claim of infringement on the part of LLM's and not on every person who has ever used a book, website, article, video or publication to learn something?
- famouswaffles 3y agoThis is an argument yes. A model could certainly be considered transformative enough to be fair use.
- blharr 3y agoI am not a lawyer. But isn't this quite simple? Copyrighted materials are either licensed specifically for a human or it's implied that a human will use them to learn. Naturally, human memory is going to distort and change that information over time. But as soon as you use it in an AI, which has superhuman capabilities of memory, that would go out the window.
- robomartin 3y ago> Copyrighted materials are either licensed specifically for a human or it's implied that a human will use them to learn. I don't think that's a part of copyright law at all. Maybe in the future, not today. Which makes sense, since these laws precede AI by a long time.
- palata 3y agoBecause humans are not machines. Otherwise we would send machines to prison when they kill someone, right? I see it like this: Say you write a book. I assume you would find it obvious that I am not allowed to copy your book, replace your name by mine, and sell it, right? That's the point of copyright. Now say I don't just copy-paste your book, but I run it through a software that replaces some words with synonyms (without losing quality or meaning), and I sell it all the same. Are you fine with that? I would tend to say that I am still abusing your copyright on your book. Generative AIs can do exactly that, and the people using generative AIs don't have a simple way to check if the output they got is a slightly-modified copy of copyrighted material or not. All we know is that the AI is a machine taking the words in your book, processing them automatically, and generating a new text. Those are not humans who learned about the world and write down their thoughts, but machines that copy-pasted-and-modified words.
- cf141q5325 3y agoA focus on licensing ignores that there are security incentive to not run just any weights you find floating around the net. Getting exploited through miss-aligned networks is a very real threat and really hard to combat.
- amelius 3y agoJust like you can't de-compile a binary without loss of information, "source" means that you can reconstruct it, so the training data should be available as well as the code that was used to train it, and the build script that invoked it.
- horsawlarway 3y agoIf anything - this entire conversation just highlights (Over and Over and Over and Over again) how absolutely bonkers abusive our current copyright laws are. The vast majority of small individuals are compelled by contract to surrender their rights to large corporations. Those large corporations then abuse the ever loving fuck out of those rights. The express intent of copyright is now a sad joke. Personally - I'm pretty over the entire show. This system is generating an incredible amount of inequality. New and novel content is absolutely NOT getting made, and these laws are creating vicious infights that drain resources from well intentioned companies & individuals and pass them along to complete scam corporations. We are told stories as children that we cannot retell in our own voices decades later to our own children. I am firmly ready to burn this copyright system to the fucking ground. It's been 300 years since the Statute of Anne - I'm ready for a different game.
- PartiallyTyped 3y agoThere is also the whole patent / copyright trolling issue too. The fact that $BIG_CORP can hire armies of lawyers to freeze competitors and beat them to market by filing frivolous lawsuits is yet another example of insanity in the whole system.
- chongli 3y agoI recently watched the documentary Fire in the Blood (2013) [1] about the use, by big pharma, of patents and WIPO to obstruct access to affordable antiretrovirals (ARVs) in Africa during the worst years of the AIDS epidemic, leading to over ten million deaths. All of this when the African market for these medications represented less than 1% of the total market, in dollars. It’s absolutely infuriating! [1] https://en.wikipedia.org/wiki/Fire_in_the_Blood_(2013_film) https://en.wikipedia.org/wiki/Fire_in_the_Blood_(2013_film)
- drdaeman 3y agoIt's a problem with legal system (not unique to any specific country, mind you, the problem is global), not patent or copyright system specifically. It grew incredible amounts of complexity so pro se became a sad joke in all but simplest cases, and there's no incentive to fix it - quite the opposite, everyone in the system is all for keeping the status quo, because it generates money.
- dahart 3y ago> Some people have the perspective that if a license isn’t open source, it’s proprietary. I think it’s more nuanced than that and believe there are three more license types worth naming: non-commercial NDA, non-commercial public, and ethical. It’s very useful to remember the U.S. government definition of commercial software: it is software that “Has been sold, leased, or licensed to the general public” [1] This means that a “non-commercial license” is a bit of an oxymoron to a lot of people. Their definition of commercial includes all software with a license, and does not depend on whether the software costs money. (Perhaps not entirely unlike how FSF does not define “free software” based on whether it costs money.) [1] https://www.acquisition.gov/far/2.101 https://www.acquisition.gov/far/2.101
- TZubiri 3y agoAgreed, output weighs are target code, and no one would argue the contrary. Companies pretending to publish source code is nothing new. Stallman defines source code as "the preferred way in which developers modify the program" I wrote for wikipedia once that "Stallman's definition thus contemplates JavaScript and HTML's source-target ambivalence, as well as contemplating possible future forms of software production, like visual programming languages, or datasets in Machine Learning." So the datasets could be a form or source code, but the most appropriate source code would be the code that crawls or downloads the dataset and modifies it. Clear as water
- low_tech_punk 3y agoThe lack of freedom to modification makes it not "open" either. Comparing to traditional software, weights are actually worse than binary. You can't "decompile" the weights into the training source code so there is no way for the community to make useful changes to them.
- light_hue_1 3y agoI think this is very shortsighted. Weights are a program. CUDA is an interpreter for that program. One day we will be able to decompile these programs into something more human understandable.
- kmeisthax 3y ago>While the RAIL organization suggests adding the word “Open” to RAIL licenses that include similar open-access and free-use as open source (i.e. OpenRAIL-M), this is confusing since the license is not open source so long as it includes usage restrictions. A better name would be EthicalRAIL-M. Using the term “ethical” to describe this category license clearly indicates its functional difference from open source licenses. I don't even think we should be using the word "ethical" because it implies that anything more permissive is unethical. We should call these morality clause licenses. The question of whether or not we should have morality clauses involved is complicated. Most bad actors do not give a shit about the licensing status of the code they are using. And these licenses also cause headaches for people who want to follow the rules[0] and avoid copyleft trolling[1]. On the other hand, the morality clauses in OpenRAIL-M are relatively straightforward and non-obnoxious. [0] This also applies to "non-commercial" licensing, since that is a concept entirely foreign to copyright law. As far as I'm concerned the 'NC' clause in Creative Commons just means 'OK to torrent'. [1] A practice in which people abuse copyleft licenses to try and extract licensing agreements for minor license violations. The forgiveness periods added to GPLv3 and later versions of Creative Commons are specifically to prevent this behavior.
- c7b 3y agoImho the weights are the real meat for most typical models, you can run with them and continue training them with your own code. It's not even guaranteed that the original code would be very useful for that. But if you are going to make that distinction, for which you can make a case I think, shouldn't you include a third dimension, 'data'? The code alone is hardly useful if you want to rebuild the weights, but all it tells you is that they're loading their proprietary data and then using PyTorch to set up and train the model. You can't reproduce anything using just that. So the real equivalent of open source would be imho either open weights, or open data plus code plus weights (the latter are arguably redundant, but still practical to include). Given that the size of that repo will typically be gigantic, I think open weights is the case we should really be focusing on. I'd rather have a paper explaining the model together with the weights, rather than code that I can't run anyway, if I'm designing an algorithm to continue training the model.
- soultrees 3y agoJust a thought, what would happen if all copyrights were abolished or the generative AI revolution we’ve been seeing will continue to the point where almost everything machine generated to a point, and therefore, open season if derivative copyright isn’t protected, then what would actually happen to the US economy? I don’t buy the argument that people just won’t innovate anymore as there won’t be an incentive anymore just doesn’t cut it. There are multiple motivations that exist simultaneously, for example - governments have a motivation to technologically advanced compared to peer nations, humans have an inherent desire to create, power, notoriety, etc etc. So in that case, let’s just abolish the first layer of incentive that actually uncovers more greed than anything by absurd copyright laws and just open the flood gates and get rid of all copyright. We need some real innovation and all this babble about who owns an ‘idea’ is way too restricting.