7 ms·
It's amazing to me that "open source" has been so diluted that it is now used to mean "we will give you an opaque binary and permission to run it on your own co
by helpfulclippy 2y ago
It's amazing to me that "open source" has been so diluted that it is now used to mean "we will give you an opaque binary and permission to run it on your own computer."
- paxys 2y agoHaha yup. Going by the current definition of "open source" in AI 100% of software created before the cloud era would have been considered open source.
- mFixman 2y agoI can't believe Microsoft finally made Windows open source.
- zelphirkalt 2y agoYay, let's "fine-tune" and share the result with everyone!
- desdenova 2y agoEvery binary is open source if you can read assembly.
- dartos 2y agoYeah blame the crowds of newbies calling llama open source bc it was free after being leaked.
- hexomancer 2y agoIf I publish some c++ code that has some hard-coded magic values in it, can the code not be considered open source until I also publish how I came up with those magic values?
- z3c0 2y agoI don't know if that compares to an AI model, where the most significant portions are the data preparation and training. The code DeepSeek released only demonstrates how to use the given weights for inferencing with Torch/Triton. I wouldn't consider that an open-source model, just wrapper code for publicly available weights. I think a closer comparison would be Android and GApps, where if you remove the latter, most would deem the phone unusable.
- mohsen1 2y agoif you publish only the binary it's not open source if open the source then it is open source if you write a book/blog about how you came up with the ideas but didn't publish the source it's not open source, even if you publish the blog+binaries
- mistercheph 2y agomodel weights != binaries
- fragmede 2y agowhy not?
- jay_kyburz 2y agoIts like the image you generated in Photoshop released as creative commons, not the Photoshop source code.
- fragmede 2y agothat adds to model weights == binaries tho
- bityard 2y agoIt depends on what those magic numbers are for. If they represent pure data, and it's obvious what the data is (perhaps a bitmap image), then sure, it's open source. If the magic values are some kind of microcode or firmware, or something else that is executed in some way, then no, it is not really open source. Even algorithms can be open source in spirit but closed source in practice. See ECDSA. The NSA has never revealed in any verifiable way how they came up with the specific curves used in the algorithm, so there is room for doubt that they weren't specifically chosen due to some inherent (but hard to find) weakness. I don't know a ton about AI, but I gather there are lots of areas in the process of producing a model where they can claim everything is "open source" as a marketing gimmick but in reality, there is no explanation for how certain results were achieved. (Trade secrets, in other words.)
- Ukv 2y ago> If the magic values are some kind of microcode or firmware, or something else that is executed in some way, then no, it is not really open source. To my understanding, the contents of a .safetensors file is purely numerical weights - used by the model defined in MIT-licensed code[0] and described in a technical report[1]. The weights are arguably only really "executed" to the same extent kernel weights of a gaussian blur filter would be, though there is a large difference in scale and effect. [0]: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/inference/model.py https://github.com/deepseek-ai/DeepSeek-V3/blob/main/inferen... [1]: https://arxiv.org/html/2412.19437v1 https://arxiv.org/html/2412.19437v1
- TeMPOraL 2y agoCode is data is code. Fundamentally, they are the same. We treat the two things as distinct categories only for practical convenience. Most of the time, it's pretty clear which is which, but we all regularly encounter situations in which the distinction gets blurry. For example: - Windows MetaFiles (WMF, EMF, EMF+), still in use (mostly inside MS Office suite) - you'd think they're just another vector image format, i.e. clearly "data", but this one is basically a list of function calls to Windows GDI APIs, i.e. interpreted code. - Any sufficiently complex XML or JSON config file ends up turning into an ad-hoc Lisp language, with ugly syntax and a parser that's a bug-ridden, slow implementation of a Lisp runtime. People don't realize that the moment they add conditionals and ability to include or refer back to other parts of config, they're more than halfway to a Turing-complete language. - From the POV of hardware, all native code is executed "to the same extent kernel weighs of a gaussian blur filter" are. In general, all code is just data for the runtime that executes it. And so on. Point being, what is code and what is data depends on practical reasons you have to make this distinction in the first place. IMHO, for OSS licensing, when considering the reasons those licenses exist, LLM weights are code.
- reedciccio 2y agoThe Open Source Definition is quite clear on its #2 requirement: `The source code must be the preferred form in which a programmer would modify the program. Deliberately obfuscated source code is not allowed.` https://opensource.org/osd https://opensource.org/osd
- ChadNauseam 2y agoArguably this would still apply to deepseek. While they didn’t release a way of recreating the weights, it is perfectly valid and common to modify the neural network using only what was released (when doing fine-tuning or RLHF for example, previous training data is not required). Doing modifications based on the weights certainly seems like the preferred way of modifying the model to me. Another note is that this may be the more ethical option. I’m sure the training data contained lots of copyrighted content, and if my content was in there I would prefer that it was released as opaque weights rather than published in a zip file for anyone to read for free.
- jonex 2y agoIt takes away the ability to know what it does though, which is also often considered an important aspect. By not publishing details on how to train the model, there's no way to know if they have included intentional misbehavior in the training. If they'd provide everything needed to train your own model, you could ensure that it's not by choosing your own data using the same methodology. IMO it should be considered freeware, and only partially open. It's like releasing an open source program with a part of it delivered as a binary.
- DonHopkins 2y agoIt's not that they want to keep the training content secret, it's the fact that they stole the training content, and who they stole it from, that they want to keep secret.
- blackeyeblitzar 2y agoIt’s because prominent people with large followings are confusing the terms on purpose. Yann LeCun of Meta and Clem Delangue of Hugging Face constantly use the wrong terms for models that only release weights, and market them to their huge audiences as “open source”. This is a willful open washing campaign to benefit from the positivity that label generates.
- seberino 2y agoI agree it would be nice to have the training specifics. Nevertheless everything DeepSeek released is under the MIT license right? So you can go set up a cloud LLM, fine tune it, and, do whatever else you wish with it right? That is pretty significant no?
- fragmede 2y agoIt is, but words mean things. If I said I got you a puppy and gave you a million dollars instead, that'd be nice, but what about the puppy?
- ogrisel 2y agoIt's better to be specific: - open-source inference code - open weights (for inference and fine-tuning) - open pretraining recipe (code + data) - open fine-tuning recipe (code + data) Very few entities publish the later two items (https://huggingface.co/blog/smollm https://huggingface.co/blog/smollm and https://allenai.org/olmo https://allenai.org/olmo come to mind). Arguably, publishing curated large scale pretraining data is very costly but publishing code to automatically curate pretraining data from uncurated sources is already very valuable.
- Palmik 2y agoAlso open-weights comes in several flavors -- there is "restricted" open-weights like Mistral's research license that prohibits most use cases (most importantly, commercial applications), then there are licenses like Llama's or DeepSeek's with some limitations, and then there are some Apache 2.0 or MIT licensed model weights.
- seberino 2y agoWait timeout. I thought DeepSeek's stuff was all MIT licensed too no? What limitations are you thinking of that DeepSeek still has?
- Palmik 2y agoI am referring to this one: https://huggingface.co/deepseek-ai/DeepSeek-V3/blob/main/LICENSE-MODEL https://huggingface.co/deepseek-ai/DeepSeek-V3/blob/main/LIC... It is a bit more permissive than Llama's it seems (no MAU threshold it seems).
- seberino 2y agoWow. Your link is frustrating because I thought everything was under the MIT license. Why did people claim it is MIT licensed if they sneaked in this additional license?
- 2y ago
- cedws 2y agoEven with the training material what good is it? The model isn’t reproducible, and even if it were you’re not going to spend the money to verify the output.
- deegles 2y agoI guess something like a kickstarter campaign would be needed to get together the millions of dollars needed per training run
- Gigachad 2y agoWho would fund that? What would be the point?
- mistercheph 2y agoFrontier models will never be reproducible in the freedom-loving countries that enforce intellectual property law, since they all depend on copyrighted content in their training data.
- barnabee 2y ago> The model isn’t reproducible Not necessarily[0], it's a WIP, but: https://github.com/huggingface/open-r1 https://github.com/huggingface/open-r1 [0] Surely they won't end up with the exact same weights, but it should be possible to verify something about the model and approach
- fragmede 2y agowhy not? if we could get a version of ChatGPT that wasn't censored and would tell me how to make meth, or an censored version of deepseek that wanted to talk about tank man, you don't think the Internet would come together and make that happen?
- mcbuilder 2y agoSurely the architecture released as a HF transformers python file counts as "open source". https://huggingface.co/deepseek-ai/DeepSeek-R1/raw/main/modeling_deepseek.py https://huggingface.co/deepseek-ai/DeepSeek-R1/raw/main/mode... Yes training is left as an exercise to the user, but it's outlined in the paper, and a good ML engineer should be able to get started with it, cluster of GPUs not included
- cma 2y agoThere was an article saying they used hand-tuned PMX instead of CUDA so it might be a bit hard to match just from the paper without some good performance experts.
- LiamPowell 2y agoCUDA isn't so bad that hand writing PTX will give you a huge performance improvement, but when you're spending a few million dollars on training it makes sense to chase even a single digit percentage improvement, maybe more in a very hot code-path. Also these articles are based on a single mention of PTX in a paper.
- cma 2y agoThe mention is here: "3.2.2. Efficient Implementation of Cross-Node All-to-All Communication In order to ensure sufficient computational performance for DualPipe, we customize efficient cross-node all-to-all communication kernels (including dispatching and combining) to conserve the number of SMs dedicated to communication. The implementation of the kernels is codesigned with the MoE gating algorithm and the network topology of our cluster. To be specific, in our cluster, cross-node GPUs are fully interconnected with IB, and intra-node communications are handled via NVLink. NVLink offers a bandwidth of 160 GB/s, roughly 3.2 times that of IB (50 GB/s). To effectively leverage the different bandwidths of IB and NVLink, we limit each token to be dispatched to at most 4 nodes, thereby reducing IB traffic. For each token, when its routing decision is made, it will first be transmitted via IB to the GPUs with the same in-node index on its target nodes. Once it reaches the target nodes, we will endeavor to ensure that it is instantaneously forwarded via NVLink to specific GPUs that host their target experts, without being blocked by subsequently arriving tokens. In this way, communications via IB and NVLink are fully overlapped, and each token can efficiently select an average of 3.2 experts per node without incurring additional overhead from NVLink. This implies that, although DeepSeek-V3 13 selects only 8 routed experts in practice, it can scale up this number to a maximum of 13 experts (4 nodes × 3.2 experts/node) while preserving the same communication cost. Overall, under such a communication strategy, only 20 SMs are sufficient to fully utilize the bandwidths of IB and NVLink. In detail, we employ the warp specialization technique (Bauer et al., 2014) and partition 20 SMs into 10 communication channels. During the dispatching process, (1) IB sending, (2) IB-to-NVLink forwarding, and (3) NVLink receiving are handled by respective warps. The number of warps allocated to each communication task is dynamically adjusted according to the actual workload across all SMs. Similarly, during the combining process, (1) NVLink sending, (2) NVLink-to-IB forwarding and accumulation, and (3) IB receiving and accumulation are also handled by dynamically adjusted warps. In addition, both dispatching and combining kernels overlap with the computation stream, so we also consider their impact on other SM computation kernels. Specifically, we employ customized PTX (Parallel Thread Execution) instructions and auto-tune the communication chunk size, which significantly reduces the use of the L2 cache and the interference to other SMs." It's definitely not the full model written in PTX or anything, but still some significant engineering effort to replicate, from people commanding 7-figure salaries in this wave, since the training code isn't open.
- JumpCrisscross 2y ago> amazing to me that "open source" has been so diluted It’s not and I called it [1]. We had three options: (A) Open weights (favoured by Altman et al); (B) Open training data (favoured by some FOSS advocates); and (C) Open weights and model, which doesn’t provide the training data, but would let you derive the weights if you had it. OSI settled on (C) [2], but it did so late. FOSS argued for (B), but it’s impractical. So the world, for a while, had a choice between impractical (B) and the useful-if-flawed (A). The public, predictably, went with the pragmatic. This was Betamax vs VHS, except in natural linguistics. There is still hope for (C). But it relies on (A) being rendered impractical. Unfortunately, the path to that flows through institutionalising OpenAI et al’s TOS-based fair use paradigm. Which means while we may get a definition (not exactly (B), but (A) absent use restrictions) we’ll also get restrictions on even using Chinese AI. [1] https://news.ycombinator.com/item?id=41047269 https://news.ycombinator.com/item?id=41047269 [2] https://opensource.org/ai/open-source-ai-definition https://opensource.org/ai/open-source-ai-definition
- sho_hn 2y agoWe absolutely had a choice (D), in that no one was forced to call it "open source" at all, which was arguably done to unfaithfully communicate benefits that don't exist. This is the part that riles people up, and that furthermore is causing collateral damage outside the AI bubble, and is nothing like Betamax vs. VHS. If you want to prioritize pragmatism, that every discussion of this includes a lengthy "so what open source do you mean, exactly?" subthread proves this was a poor choice. It causes uncertainly that also makes it harder for the folks releasing these models to make their case and be taken seriously for their approach. We should probably call them "free to run", if the "it's cheap" connotation of "freeware" needs to be avoided. Or maybe "open architecture" to appreciate the Python file that utilizes the weights more.
- JumpCrisscross 2y ago> We absolutely had a choice (D), in that no one was forced to call it "open source" at all Technically yes, practically no. You’re describing a prisoner’s dilemma. The term was available, there was (and remains) genuine ambiguity over what it meant in this context, and there are first-mover advantages in branding. (Exhibit A: how we label charges). > causing collateral damage outside the AI bubble, and is nothing like Betamax vs. VHS Standards wars have collateral damage. > We should probably call them "free to run", if the "it's cheap" connotation of "freeware" needs to be avoided. Or maybe "open architecture" Language is parsimonious. A neologism will never win when a semantic shift will do.
- Palmik 2y agoExcept the "binary" is not really opaque, and can be "edited" in exactly the same way it was produced in the first place (continued pre-training / fine-tuning).
- seberino 2y agoI'm not an expert but didn't they release the weights under MIT license? So you can make your own LLM with complete control right? I agree it would nice to know the details of their training, but, simply calling this drop an "opaque binary" is seriously underselling it no?