6 ms·
I don't think people make the distinction like that. The open source vs non open source distinction boils down to, usually, can you use it for commercial use.
by make3 1y ago
I don't think people make the distinction like that. The open source vs non open source distinction boils down to, usually, can you use it for commercial use.
what you're saying is just that it's non reproducible, which is a completely valid but separate issue
- piperswe 1y agoBut where's the source? I just see a binary blob, what makes it open source?
- jacob019 1y agoThe weights are the source. It isn't as though something was compiled into weights. They're trained directly. But I know what you mean, it would be more open to have the training pipeline and souce dataset available.
- timschmidt 1y agoThe weights seem much more like a binary to me, the training pipeline the compiler, and the training dataset the source.
- jumski 1y agoCome here to write this - perfect analogy!
- reedciccio 1y agoIt's very imperfect analogy though these things can't be rebuilt "from scratch" like a program, the training process doesn't seem to be replicable anyway. Nonetheless, full data disclosure is necessary, according to the result of the years-long consultation led by the Open Source Initiative https://opensource.org/ai https://opensource.org/ai
- timschmidt 1y ago> the training process doesn't seem to be replicable anyway The training process is fully deterministic. It's just an algorithm. Feed the same data in and you'll get the same weights out. If you're speaking about the computational cost, it used to be that way for compilers too. Give it 20 years and you'll be able to train one of today's models on your phone.
- reedciccio 1y agoCan you point at the research that says that the training process of a LLM at least the size of OLMo or Pythia is deterministic?
- timschmidt 1y agoCan you point to something that says it's not? The only source of non-determinism I've read of affecting LLM training is floating point error which is well understood and worked around easily enough.
- reedciccio 1y agoSearch more, there is a lot of literature discussing how hard the problem of reproducibility of GenAI/LLMs/Deep Learning is, how far we are from solving it for trivial/small models (let alone for beasts the size of the most powerful ones) and even how pointless the whole exercise is.
- timschmidt 1y agoIf there's a lot, then it should be easy for you to link an example right? One that points toward something other than floating point error. There simply aren't that many sources of non-determinism in a modern computer. Though I'll grant that if you've engineered your codebase for speed and not for determinism, error can creep in via floating point error, sloppy ordering of operations, etc. These are not unavoidable implementation details, however. CAD kernels and other scientific software do it every day. When you boil down what's actually happening during training, it's just a bunch of matrix math. And math is highly repeatable. Size of the matrix has nothing to do with it. I have little doubt that some implementations aren't deterministic, due to software engineering choices as discussed above. But the algorithms absolutely are. Claiming otherwise seems equivalent to claiming that 2 + 2 can sometimes equal 5.
- 1una 1y agoI won't call it "binary blob". Safetensors is just a simple format for storing tensors safely: https://huggingface.co/docs/safetensors/index https://huggingface.co/docs/safetensors/index
- otabdeveloper4 1y agoYou can fine-tune their weights and release your own take. E.g. see all the specialized third-party models out there based on Qwen. "Open-source" is the wrong word here, what they mean is "you can modify and redistribute these weights".
- yetihehe 1y agoYou can also reverse engineer and modify closed source programs (see mods for games). Weights are like compiled version of source data.
- macrolime 1y agoNot legally. That's the difference.
- timschmidt 1y agoSure you can. It's often legally protected activity. You're just limited to distributing your modifications without the original work.
- macrolime 1y agoFor some games maybe, but software often has a clause forbidding reverse engineering
- timschmidt 1y agoChatGPT says that such clauses are typically void in the EU, though they may apply in some cases in the US. Even in the US, the triennial DMCA rule-making has granted broader exemptions for good-faith security research every cycle since 2016. https://chatgpt.com/share/6838c070-705c-8005-9a88-83c9a5550a1e https://chatgpt.com/share/6838c070-705c-8005-9a88-83c9a5550a...
- otabdeveloper4 1y ago
- microtonal 1y agoThere is work to try to reproduce (the original) R1: https://huggingface.co/open-r1 https://huggingface.co/open-r1
- alpaca128 1y agoThere's already established terms and licenses for non-commercial use. Like "open weights". Open source has the word "source" in it for a reason, and those models ain't open source and have nothing to do with it.
- ben_w 1y agoTook me until this thread to remember that in the 90s we had "freeware".