4 ms·
It's entirely reproducible from the available documentation (which is why you see vLLM, SGLang, MLX etc all racing to produce optimized implementations). (As a
by nl 2mo ago
It's entirely reproducible from the available documentation (which is why you see vLLM, SGLang, MLX etc all racing to produce optimized implementations).
(As an aside, this is why the "open weights are not open source" thing is a complete misunderstanding. The weights themselves along with the documentation give you enough to fine tune the LLM. You can't rebuild it from scratch, but you can't do this even with the data anyway (because of randomness!))
- esperent 2mo ago> open weights are not open source" thing is a complete misunderstanding > You can't rebuild it from scratch There's is extremely clear and misunderstanding-free. Open weights is not open source.
- FinchNova12 2mo agoI agree that if you have the weights you can use/train a model with the same architecture, and that you won't get the exact weights on your own due to randomness. But isn't data an extremely important part of your ability to effectively train/finetune? It might be much harder to get close to the level of the open weight model if you don't have the data that made it, which is why I think the open weights vs open source distinction is useful.
- nl 2mo agoYou fine-tune the open-weight model without access to the original weights. If you want to train it from scratch you need data, yes. But presumably if you are doing that there is a reason you want to do it. You lose nothing without access to the original data - you can do every single modification without it. That is unlike open source where you (mostly) need to source code to modify it beyond what the original designed originally thought.
- spider-mario 2mo agoIsn’t this a bit like saying “ ‘open object files are not open source’ is a complete misunderstanding” because “You can’t rebuild the executable from scratch, but you can’t do this even with the source code anyway (because of build nondeterminism / compiler versions / etc.)”?
- desterothx 2mo agonot really, compiling isnt a heuristic problem, it has a lot less randomness involved
- eru 2mo agoDepends on your compiler. You could have a compiler that deliberately uses randomised algorithms. They are often faster and easier to understand and write. Though in practice you can get all the benefits of both determinism and (that kind of) randomisation by using a PRNG and saving the seed you are using. It's an open question roughly on par with P vs NP whether true randomisation is ever necessary, or whether PRNGs are enough. So far we haven't found any problem or algorithm where true RNG is necessary and good PRNG ain't enough.
- pcmasterr 2mo agoAlso, most (optimizing) compilers ”optimize” the code for a fixed amount of time, leading to better optimized binaries on faster computers. That’s why developers should have as fast computers money can buy!
- eru 2mo ago> That’s why developers should have as fast computers money can buy! I don't see the connection? Most local builds are done with debugging on and optimisation turned off anyway. And what you deliver to your customers is usually something you produce on your CI/CD server, not what's on any developer's machine. (And if you want reproducible builds https://en.wikipedia.org/wiki/Reproducible_builds https://en.wikipedia.org/wiki/Reproducible_builds you can't optimise for a specific wall clock time.) > Also, most (optimizing) compilers ”optimize” the code for a fixed amount of time, leading to better optimized binaries on faster computers. That’s why you should give your developers computers that have slow clocks!
- NooneAtAll3 2mo agothere's a reason reproducible builds are a thing and most compilations aren't
- eru 2mo ago> (As an aside, this is why the "open weights are not open source" thing is a complete misunderstanding. The weights themselves along with the documentation give you enough to fine tune the LLM. You can't rebuild it from scratch, but you can't do this even with the data anyway (because of randomness!)) They could give you the random seeds? (Assuming you carefully train in such a way to remove other sources of randomness, like concurrent execution.)
- alightsoul 2mo agoDo AI labs even keep track of seeds? They don't even keep the checkpoints most of the time.
- eru 2mo agoIt's moot in practice, because there's additional sources of randomness from executing in parallel. Though in principle saving random seeds is a lot less hassle than keeping entire checkpoints around: your random seeds would fit on a floppy disk or even a tweet. The checkpoint is basically as big as the model.
- embedding-shape 2mo ago> You can't rebuild it from scratch > It's entirely reproducible from the available documentation You have a very interesting understanding of "reproducibility", I'll give you that :) But even with that, there are plenty of technical details (especially in regards to the training process) missing from the tech report that leads to these weights not being reproducible in any sense of that word.
- nvme0n1p1 2mo ago> You have a very interesting understanding of "reproducibility", I'll give you that :) He was talking about two different things, hence the parentheses. The architecture is reproducible, not the model weights.
- embedding-shape 2mo ago> The architecture is reproducible But what's the point of even saying that? Of course it is, otherwise how is it supposed to run in the runtimes? You cannot release model weights that others can run, without also releasing the model architecture, it's in the code at the very least...
- davidguetta 2mo agothe beginning of the arguments was good but the "because of randomness" is wild
- nl 2mo agoWhy? Have you ever tried reproducing even a small neural network exactly if you train on GPUs on more than one machine? I have and it is pretty close to impossible, and I'd argue actually impossible at scale.
- monocasa 2mo agoThere's more to open source than reproducibility. For instance introspection which is even more important with weights since detection of backdoors in models is NP hard IIRC.