20 ms·
"Eventually though, open source Linux gained popularity – initially because it allowed developers to modify its code however they wanted ..." I find the langua
by the8thbit 2y ago
"Eventually though, open source Linux gained popularity – initially because it allowed developers to modify its code however they wanted ..."
I find the language around "open source AI" to be confusing. With "open source" there's usually "source" to open, right? As in, there is human legible code that can be read and modified by the user? If so, then how can current ML models be open source? They're very large matrices that are, for the most part, inscrutable to the user. They seem akin to binaries, which, yes, can be modified by the user, but are extremely obscured to the user, and require enormous effort to understand and effectively modify.
"Open source" code is not just code that isn't executed remotely over an API, and it seems like maybe its being conflated with that here?
- bilsbie 2y agoCan’t you do fine tuning on those binaries? That’s a modification.
- the8thbit 2y agoYou can fine tune the models, and you can modify binaries. However, there is no human readable "source" to open in either case. The act of "fine tuning" is essentially brute forcing the system to gradually alter the weights such that loss is reduced against a new training set. This limits what you can actually do with the model vs an actual open source system where you can understand how the system is working and modify specific functionality. Additionally, models can be (and are) fine tuned via APIs, so if that is the threshold required for a system to be "open source", then that would also make the GPT4 family and other such API only models which allow finetuning open source.
- whimsicalism 2y agoI don't find this argument super convincing. There's a pretty clear difference between the 'finetuning' offered via API by GPT4 and the ability to do whatever sort of finetuning you want and get the weights at the end that you can do with open weights models. "Brute forcing" is not the correct language to use for describing fine-tuning. It is not as if you are trying weights randomly and seeing which ones work on your dataset - you are following a gradient.
- the8thbit 2y ago"There's a pretty clear difference between the 'finetuning' offered via API by GPT4 and the ability to do whatever sort of finetuning you want and get the weights at the end that you can do with open weights models." Yes, the difference is that one is provided over a remote API, and the provider of the API can restrict how you interact with it, while the other is performed directly by the user. One is a SaaS solution, the other is a compiled solution, and neither are open source. ""Brute forcing" is not the correct language to use for describing fine-tuning. It is not as if you are trying weights randomly and seeing which ones work on your dataset - you are following a gradient." Whatever you want to call it, this doesn't sound like modifying functionality in source code. When I modify source code, I might make a change, check what that does, change the same functionality again, check the new change, etc... up to maybe a couple dozen times. What I don't do is have a very simple routine make very small modifications to all of the system's functionality, then check the result of that small change across the broad spectrum of functionality, and repeat millions of times.
- Kubuxu 2y agoThe gap between fine-tuning API and weights-available is much more significant than you give it credit for. You can take the weights and train LoRAs (which is close to fine-tuning), but you can also build custom adapters on top (classification heads). You can mix models from different fine-tunes or perform model surgery (adding additional layers, attention heads, MoE). You can perform model decomposition and amplify some of its characteristics. You can also train multi-modal adapters for the model. Prompt tuning requires weights as well. I would even say that having the model is more potent in the hands of individual users than having the dataset.
- thayne 2y agoThat still doesn't make it open source. There is a massive difference between a compiled binary that you are allowed to do anything you want with, including modifying it, building something else on top or even pulling parts of it out and using in something else, and a SaaS offering where you can't modify the software at all. But that doesn't make the compiled binary open source.
- bilsbie 2y agoYou make a good point but those are also just limitations of the technology (or at least our current understanding of it) Maybe an analogy would help. A family spent generations breeding the perfect apple tree and they decided to “open source” it. What would open sourcing look like?
- the8thbit 2y ago"You make a good point but those are also just limitations of the technology (or at least our current understanding of it)" Yeah, that is my point. Things that don't have source code can't be open source. "Maybe an analogy would help. A family spent generations breeding the perfect apple tree and they decided to “open source” it. What would open sourcing look like?" I think we need to be weary of dilemmas without solutions here. For example, let's think about another analogy: I was in a car accident last week. How can I open source my car accident? I don't think all, or even most things, are actually "open sourcable". ML models could be open sourced, but it would require a lot of work to interpret the models and generate the source code from them.
- gowld 2y agoBe charitable and intellectually curious. What would "open" look like? GNU says "The GNU GPL can be used for general data which is not software, as long as one can determine what the definition of “source code” refers to in the particular case. As it turns out, the DSL (see below) also requires that you determine what the “source code” is, using approximately the same definition that the GPL uses." and offers these categories, for example: https://www.gnu.org/licenses/license-list.en.html#NonFreeSoftwareLicenses https://www.gnu.org/licenses/license-list.en.html#NonFreeSof... * Software Licenses * * GPL-Compatible Free Software Licenses \ * * GPL-Incompatible Free Software Licenses \ * Licenses For Documentation * * Free Documentation Licenses \ * Licenses for Other Works * * Licenses for Works of Practical Use besides Software and Documentation * * Licenses for Fonts * * Licenses for Works stating a Viewpoint (e.g., Opinion or Testimony) * * Licenses for Designs for Physical Objects
- the8thbit 2y ago
- jpadkins 2y ago> vs an actual open source system where you can understand how the system is working and modify specific functionality. No one on the planet understands how the model weights work exactly, nor can they modify them specifically (i.e. hand modifying the weights to get the result they want). This is an impossible standard. The source code is open (sorta, it does have some restrictions). The weights are open. The training data is closed.
- the8thbit 2y ago> No one on the planet understands how the model weights work exactly Which is my point. These models aren't open source because there is no source code to open. Maybe one day we will have strong enough interpretability to generate source from these models, and then we could have open source models. But today its not possible, and changing the meaning of open source such that it is possible probably isn't a great idea.
- orthoxerox 2y agoOpen training dataset + open steps sufficient to train exactly the same model.
- the8thbit 2y agoThis isn't what Meta releases with their models, though I would like to see more public training data. However, I still don't think that would qualify as "open source". Something isn't open source just because its reproducible out of composable parts. If one, very critical and system defining part is a binary (or similar) without publicly available source code, then I don't think it can be said to be "open source". That would be like saying that Windows 11 is open source because Windows Calculator is open source, and its a component of Windows.
- orthoxerox 2y agoThat's what I meant by "open steps", I guess I wasn't clear enough.
- the8thbit 2y agoIs that what you meant? I don't think releasing the sequence of steps required to produce the model satisfies "open source", which is how I interpreted you, because there is still no source code for the model.
- blackeyeblitzar 2y agoHere’s one list of what is needed to be actually open source: https://blog.allenai.org/hello-olmo-a-truly-open-llm-43f7e7359222 https://blog.allenai.org/hello-olmo-a-truly-open-llm-43f7e73...
- Yizahi 2y agoThey can't release training dataset if it was illegally scrapped all over the web without permission :) (taps head)
- jsheard 2y agoI also think that something like Chromium is a better analogy for corporate open source models than a grassroots project like Linux is. Chromium is technically open source, but Google has absolute control over the direction of it's development and realistically it's far too complex to maintain a fork without Googles resources, just like Meta has complete control over what goes into their open models, and even if they did release all the training data and code (which they don't) us mere plebs could never afford to train a fork from scratch anyway.
- skybrian 2y agoI think you’re right from the perspective of an individual developer. You and I are not about to fork Chromium any time soon. If you presume that forking is impractical then sure, the right to fork isn’t worth much. But just because a single developer couldn’t do it doesn’t mean it couldn’t be done. It means nobody has organized a large enough effort yet. For something like a browser, which is critical for security, you need both the organization and the trust. Despite frequent criticism, Mozilla (for example) is still considered pretty trustworthy in a way that an unknown developer can’t be.
- Yizahi 2y agoIf Microsoft can't do it, then we can reasonably conclude that it can't be done for any practical purpose. Discussing infinitesimal possibilities is better left to philosophers.
- candiddevmike 2y agoNone of Meta's models are "open source" in the FOSS sense, even the latest Llama 3.1. The license is restrictive. And no one has bothered to release their training data either. This post is an ad and trying to paint these things as something they aren't.
- JumpCrisscross 2y ago> no one has bothered to release their training data If the FOSS community sets this as the benchmark for open source in respect of AI, they're going to lose control of the term. In most jurisdictions it would be illegal for the likes of Meta to release training data.
- exe34 2y agothe training data is the source.
- JumpCrisscross 2y ago> the training data is the source Sure. But that's not going to be released. The term open source AI cannot be expected to cover it because it's not practical.
- diggan 2y agoSo because it's really hard to do proper Open Source with these LLMs, means we need to change the meaning of Open Source so it fits with these PR releases?
- JumpCrisscross 2y ago> because it's really hard to do proper Open Source with these LLMs, means we need to change the meaning of Open Source so it fits with these PR releases? Open training data is hard to the point of impracticality. It requires excluding private and proprietary data. Meanwhile, the term "open source" is massively popular. So it will get used. The question is how. Meta et al would love for the choice to be between, on one hand, open weights only, and, on the other hand, open training data, because the latter is impractical. That dichotomy guarantees that when someone says open source AI they'll mean open weights. (The way open source software, today, generally means source available, not FOSS.)
- causal 2y ago"Open weights" is a more appropriate term but I'll point out that these weights are also largely inscrutable to the people with the code that trained it. And for licensing reasons, the datasets may not be possible to share. There is still a lot of modifying you can do with a set of weights, and they make great foundations for new stuff, but yeah we may never see a competitive model that's 100% buildable at home. Edit: mkolodny points out that the model code is shared (under llama license at least), which is really all you need to run training https://github.com/meta-llama/llama3/blob/main/llama/model.py https://github.com/meta-llama/llama3/blob/main/llama/model.p...
- ajxlasA 2y agoReally? I have to check out the training code again. Last time I looked the training and inference code were just example toys that were barely usable. Has that changed?
- deleted 2y ago[deleted]
- aerzen 2y agoLLAMA is an open-weights model. I like this term, let's use that instead of open source.
- gowld 2y agoCan a human programmer edit the weights according to some semantics?
- sebastiennight 2y agoIt is possible to merge two fine-tunes of models from the same family by... wait for it... averaging or combining their weights[0]. I am still amazed that we can do that. [0]: https://arxiv.org/abs/2212.09849 https://arxiv.org/abs/2212.09849
- 2y ago
- mkolodny 2y agoLlama’s code is open source: https://github.com/meta-llama/llama3/blob/main/llama/model.py https://github.com/meta-llama/llama3/blob/main/llama/model.p...
- apsec112 2y agoThat's not the training code, just the inference code. The training code, running on thousands of high-end H100 servers, is surely much more complex. They also don't open-source the dataset, or the code they used for data scraping/filtering/etc.
- the8thbit 2y ago"just the inference code" It's not the "inference code", its the code that specifies the architecture of the model and loads the model. The "inference code" is mostly the model, and the model is not legible to a human reader. Maybe someday open source models will be possible, but we will need much better interpretability tools so we can generate the source code from the model. In most software projects you write the source as a specification that is then used by the computer to implement the software, but in this case the process is reversed.
- blackeyeblitzar 2y agoThat is just the inference code. Not training code or evaluation code or whatever pre/post processing they do.
- patrickaljord 2y agoIs there an LLM with actual open source training code and dataset? Besides BLOOM https://huggingface.co/bigscience/bloom https://huggingface.co/bigscience/bloom
- osanseviero 2y agoYes, there are a few dozen full open source models (license, code, data, models)
- stale2002 2y agoOk call it Open Weights then if the dictionary definitions matter so much to you. The actual point that matters is that these models are available for most people to use for a lot of stuff, and this is way way better than what competitors like OpenAI offer.
- deleted 2y ago[deleted]
- the8thbit 2y agoThey don't "[allow] developers to modify its code however they want", which is a critical component of "open source", and one that Meta is clearly trying to leverage in branding around its products. I would like them to start calling these "public weight models", because what they're doing now is muddying the waters so much that "open source" now just means providing an enormous binary and an open source harness to run it in, rather than serving access to the same binary via an API.
- sailingparrot 2y agoFeels a bit like you are splitting hair for the pleasure of semantic arguments to be honest. Yes there are no source in ML, so if we want to be pedantic it shouldn't be called open source. But what really matters in the open source movement is that we are able to take a program built by someone and modify it to do whatever we want with it, without having to ask someone for permission or get scrutinized or have to pay someone. The same applies here, you can take those models and modify them to do whatever you want (provided you know how to train ML models), without having to ask for permission, get scrutinized or pay someone. I personally think using the term open source is fine, as it conveys the intent correctly, even if, yes, weights are not sources you can read with your eyes.
- wrs 2y agoCalling that “open source” renders the word “source” meaningless. By your definition, I can release a binary executable freely and call it “open source” because you can modify it to do whatever you want. Model weights are like a binary that nobody has the source for. We need another term.
- input_sh 2y agoOpen Source Initiative (kind of a de-facto authority on what's open source and what not) is spending a whole lot of time figuring out what it means for an AI system to be open source. In other words, they're basically trying to come up with a new license because the existing ones can't easily apply. I believe this is the current draft: https://opensource.org/deepdive/drafts/the-open-source-ai-definition-draft-v-0-0-8 https://opensource.org/deepdive/drafts/the-open-source-ai-de...
- downWidOutaFite 2y agoOSI made themselves the authority because they hated Richard Stallman and his Free Software movement. It's just marketing.
- gowld 2y agoRMS has no interest in governing Open Source, so your comment bears no particular relevance. RMS is an advocate for Free Software. Free Software generally implies Open Source, but not the converse. RMS considers openness of source to be a separate category from the freeness of software. "Free software is a political movement; open source is a development model." https://www.gnu.org/licenses/license-list.en.html https://www.gnu.org/licenses/license-list.en.html
- ab5tract 2y agoAre you really pretending that OSI and the open source label itself wasn’t a reactionary movement that vilified free software principles in hopes of gaining corporate traction? Most of us who were there remember it differently. True open source advocates will find little to refute in what I’ve said.
- cheema33 2y ago> True open source advocates will find little to refute in what I’ve said. No true Scotsman https://en.wikipedia.org/wiki/No_true_Scotsman https://en.wikipedia.org/wiki/No_true_Scotsman OSI helped popularize the open source movement. They not only make it palatable to businesses, but got them excited about it. I think that FSF/Stallman alone would not have been very successful on this front with GPL/AGPL.
- Zambyte 2y ago> If so, then how can current ML models be open source? The source of a language model is the text it was trained on. Llama models are not open source (contrary to their claims), they are open weight.
- thayne 2y agoI think it would also include the code used to train it
- pphysch 2y agoThat would be more analogous to the build toolchain than the source code, but yes
- tshaddox 2y agoSurely traditional “open source” also needs some notion of a reproducible build toolchain, otherwise the source code itself is approximately useless. Imagine if the source code was in a programming language of which the basic syntax and semantics were known to no one but the original developers. Or more realistically, I think it’s a major problem if an open source project can only be built by an esoteric process that only the original developers have access to.
- pphysch 2y agoSource code in a vacuum is still valuable as a way to deal with missing/inaccurate documentation and diagnose faults and their causes. Raw training datasets similarly has some value as you can analyze it for different characteristics to understand why the trained model is under/over-representing different concepts. But yes real FOSS should be "open-build" and allow anyone to build a test-passing artifact from raw source material.
- moffkalast 2y agoYou can find the entire Llama 3.0 pretraining set here: https://huggingface.co/datasets/HuggingFaceFW/fineweb https://huggingface.co/datasets/HuggingFaceFW/fineweb 15T tokens, 45 terrabytes. Seems fairly open source to me.
- gorgoiler 2y agoOne counterpoint is that major publications (eg New York Times) would have you believe that AI is a mildly lossy compression algorithm capable of reconstructing the original source material.
- actinium226 2y agoIt's not?
- _flux 2y agoI believe it is able to reconstruct parts of the original source material—if the interrogator already knows the original source material to prompt the model appropriately.
- halflings 2y agoTraining code is only useful to people in academia, and the closest thing to "code you can modify" are open weights. People are framing this as if it was an open-source hierarchy, with "actual" open-source requiring all training code to be shared. This is not obvious to me, as I'm not asking people that share open-source libraries to also share the tools they used to develop them. I'm also not asking them to share all the design documents/architecture discussion behind this software. It's sufficient that I can take the end result and reshape it in any way I desire. This is coming from an LLM practitioner that finetunes models for a living; and this constant debate about open-source vs open-weights seems like a huge distraction vs the impact open-sourcing something like Llama has... this is truly a Linux-like moment. (at a much smaller scale of course, for now at least)
- kemiller 2y agoI dunno — if an open source project required, say, a proprietary compiler, that would diminish its open source-ness. But I agree it's not totally comparable, since the weights are not particularly analogous to machine code. We probably need a new term. Open Weights.
- 0-_-0 2y agoThere are many "compilers", you can download The Pile yourself.
- nothrowaways 2y agoWeight is the new code.
- nomel 2y agoI think saying it's the new binary is closer to the truth. You can't reproduce it, but you can use it. In this new version, you can even nudge it a bit to do something a little different. New stuff, so probably not good to force old words, with known meanings, onto new stuff.
- GreenWatermelon 2y agoThe model is more akin to a python script than a compiled C binary. This is how I see it: Training Code and dataset are analogous to the developer who wrote the script Model and weights are end product that is then released Inference Code is the runtime that could execute the code. That would be e.g. PyTorch, which can import the weights and run inference.
- nomel 2y ago> The model is more akin to a python script than a compiled C binary. No, I completely disagree. Python is near pseudo-text source. Source exists for the specific purpose of being easily and completely understood, by humans, because it's for and from humans. You can turn a python calculator into a web server, because it can be split and separated at any point, because it can be completely understood at any point, and it's deterministic at every point. A model cannot be understood by a human. It isn't meant to be. It's meant to be used, very close to as is. You can't fundamentally change the model, or dissect it, you can only nudge it in a direction, with the force of that nudge being proportional to the money you can burn, along with hope that it turns out how you want. That's why I say it's closer to a binary: more of a black box you can use. You can't easily make a binary do something fundamentally different without changing the source. You can't easily see into that black box, or even know what it will do without trying. You can only nudge it to act a little differently, or use it as part of a workflow. (decompilation tools aside ;))
- GuB-42 2y agoI like the term "open weights". Open source would be the dataset and code that generates these weights. There is still a lot you can do with weights, like fine tuning, and it is arguably more useful as retraining the entire model would cost millions in compute.
- szundi 2y agoOpen source = reproducible binaries (weights) by you on your computer, IMO. Strategy of FB is that they are good to be a user only and fine ruining competitor’s business with good enough free alternatives while collecting awards as saviors of whatever.
- j_maffe 2y agoNot sure what you mean by "they are good to be a user only." Whatever their strategy is, this is great for the community.
- ric2b 2y agoIf that were the definition then any software you can install on your computer would be open source. It makes open source lose nearly all meaning. Just say "open weights", not "open source".
- rmbyrro 2y agoIf you think about LLMs as a new kind of programming runtime, the matrices are the source.
- beloch 2y agoIt's no secret that implementing AI usually involves far more investment into training and teaching than actual code. You can know how a neural net or other ML model works. You can have all the code before you. It's still a huge job (and investment) to do anything practical with that. If Meta shares the code their AI runs on with you, you're not going to be able to do much with it unless you make the same investment in gathering data and teaching to train that AI. That would probably require data Meta won't share. You'd effectively need your own Facebook. If everyone open sources their AI code, Meta can snatch the bits that help them without much fear of helping their direct competitors.
- the8thbit 2y agoI think you're misunderstanding what I'm saying. I don't think its technically feasible for current models to be open source, because there is no source code to open. Yes, there is a harness that runs the model, but the vast, vast amount of instructions are contained in the model weights, which are akin to a compiled binary. If we make large strides in interpretability we may have something resembling source code, but we're certainly not there yet. I don't think the solution to that problem should be to change the definition of open source and pretend the problem has been solved.
- seoulmetro 2y agoUnfortunately open source really just means an open API these days. The API is heavily intertwined with closed source.
- langcss 2y agoComing up with the words and concepts to describe the models is a challenge. Does the training data require permission from the copyright holder to use? Are the weights really open source or more like compiled assembly?
- shdjkKA 2y agoOf course you are right, I'd put it less carefully: The quoted Linux line is deceptive marketing. - If we start with the closed training set, that is closed and stolen, so call it Stolen Source. - What is distributed is a bunch of float arrays. The Llama architecture is published, but not the training or inference code. Without code there is no open source. You can as well call a compiler book open source, because it tells you how to build a compiler. Pure marketing, but predictably many people follow their corporate overlords and eagerly adopt the co-opted terms. Reminder again that FB is not releasing this out of altruism, but because they have an existing profitable business model that does not depend on generated chats. They probably do use it internally for tracking and building profiles, but that is the same as using Linux internally, so they release the weights to destroy the competition. Isn't price dumping an anti trust issue?
- bjornsing 2y agoThe term “source code” can mean many things. In a legal context it’s often just defined as the preferred format for modification. It can be argued that for artificial neural networks that’s the weights (along with code and preferably training data).
- kashyapc 2y agoI agree; there's a lot of muddiness in the term "open source AI". Earlier this year there was a talk[1] at FOSDEM, titled "Moving a step closer to defining Open Source AI". It is from someone at the Open Source Initiative. The video and slides are available in the link below[1]. From the abstract: "Finding an agreement on what constitutes Open Source AI is the most important challenge facing the free software (also known as open source) movement. European regulation already started referring to "free and open source AI", large economic actors like Meta are calling their systems "open source" despite the fact that their license contain restrictions on fields-of-use (among other things) and the landscape is evolving so quickly that if we don't keep up, we'll be irrelevant." [1] https://fosdem.org/2024/schedule/event/fosdem-2024-2805-moving-a-step-closer-to- https://fosdem.org/2024/schedule/event/fosdem-2024-2805-movi... defining-open-source-ai/
- rbits 2y agoYou release all the technology and the training data. Everything that was used to create the model, including instructions. I'm not sure if facebook has done that
- roguas 2y agoNo, open source means that sources are open, typically for inspection, modification etc. Also here it can be considered the case. Likely in order to claim "true open source", they would have to share dataset? But even this might not be enough for truely open source model? This dataset is nothing but another artifact. So how did they arrive at this dataset, now they have to share pipelines and infra... .. the thing is, we have not dealt with llm much, it's hard to say what can be considered open source llm just yet, so we use that as metaphore for now