5 ms·
No, it's not open source till someone can actually reproduce it. That's the hardest part. For now it's open weights open dataset. Which is not the same.
by bubaumba 2y ago
No, it's not open source till someone can actually reproduce it. That's the hardest part. For now it's open weights open dataset. Which is not the same.
- diggan 2y agoThat's... Not how open source works? The "binary" (model weights) is open source and the "software" (training scripts + data used for training) is open source, this release is a real open source release. Independent reproduction is not needed to call something open source. Can't believe it's the second time I end up with the very same argument about what open source is today on HN.
- dboreham 2y agoBut wouldn't failure to achieve independent reproduction falsify the open claim? Similar to you publish the source for Oracle (the database), but nobody can build a binary from it because it needs magic compliers or test suites that aren't open source? Heck when the browser was open-sourced, there was an explicit test where the source was given to some dude who didn't work for Netscape to verify that he could actually make a working binary. It's a scene in the movie "Code Rush".
- bubaumba 2y agoYou are missing key points here. "reproduce" means produce the same. Not just train similar model. I can simplify the task, can you convincingly explain how the same model can be produced from this dataset? We can start simple, how you can possibly get the same weights after the first single iteration? I.e. the same as original model got. Pay attention to randomness, data selection, initial model state. Ok, if you can't do that. Can you explain in believable way how to prove that given model was trained on give dataset? I'm not asking you for actually doing all these things, that could be expensive, only to explain how it can be done. Strict 'open source' includes not only open weights, open data. It also includes the word "reproducible". It's not "reproduced", only "reproducible". And even this is not the case here.
- worewood 2y agoHow often do people expect to compile open-source code and get _exactly_ the same binary as the distributed one? I've seen this kind of restriction only on decompilation projects e.g. the SM64 decompilation -- where they deliberately compare the hashes of original vs. compiled binaries, as a way to verify the decompilation is correct. It's an unreasonable request with ordinary code, even more with ML where very few ones have access to the necessary hardware, and where in practice, it is not deterministic.
- e12e 2y agoI expect that if I compile your 3d renderer, and feed it the same scene file you did - I get the same image?
- TylerE 2y agoWhy would you expect that? 3D renderers are not generally deterministic. Many will incorporate, for instance, noise algorithms. They will frequently not produce byte-identical renders on the same hardware using the same binary.
- e12e 2y agoSame recognizable image? Like if you look at the povray benchmark image on Linux and Windows? For an llm that should (ideally) translate to similar answers?
- bubaumba 2y ago> How often do people expect to compile open-source code and get _exactly_ the same binary _Always_, with the right options. And that's the key point. If distributed code is different it means it may be infected or altered in other way. In other words it cannot be trusted. The same with models, if they are not reproducible or verifiable they cannot be trusted. Trust is the main feature of open source. Calling black box with attached data 'open source', even 'the first' is a bit of a stretch. It's not reproducible and not verifiable. And it's definitely not the first model with open data. To be correct you should add 'untrusted' if you want to call this thing 'open source'. Like with Meta's models who knows what it holds. PS: finally I'm negative, fanboys don't like it ;-)
- wrs 2y agoThe interesting part of the product we’re taking about (that is, the equivalent of the executable binary of an ordinary software product) is the weights. The “source” is not sufficient to “recompile” the product (i.e., recreate the weights). Therefore, while the source you got is open, you didn’t get all the source to the thing that was supposedly “open source”. It’s like if I said I open-sourced the Matrix trilogy and only gave you the DVD image and the source to the DVD decoder. (Edit: Sorry, I replied to the wrong comment. I’m talking primarily about the typical sort of release we see, not this one which is a lot closer to actually open.)
- littlestymaar 2y ago> The “source” is not sufficient to “recompile” the product (i.e., recreate the weights). Therefore, while the source you got is open, you didn’t get all the source to the thing that was supposedly “open source”. What's missing?
- wrs 2y agoWell, I’m not experienced in training full-sized LLMs, and it’s conceivable that in this particular case the training process is simple enough that nothing is missing. That would be a rarity, though. But see my edit above — I’m not actually reacting to this release when I say that.
- littlestymaar 2y agoOK, so you just like to be a contrarian…
- Jabrov 2y agoWhat’s the difference?
- avaldez_ 2y agoReproducibility? I mean what's the point of an open technology nobody knows if it works or not.
- frontalier 2y agothe goal posts moved?!