4 ms·
Debian packager of Leela Zero here. Most of the questions raised in the thread were basically irrelevant to Leela Zero, which does everything "correct" from th
by infinity0 8y ago
Debian packager of Leela Zero here.
Most of the questions raised in the thread were basically irrelevant to Leela Zero, which does everything "correct" from the point-of-view of a strict interpretation of free software:
- open freely licensed data set
- open freely licensed training code
The only issue relevant for user software freedom, is that the results of the training process can't be reproduced easily.
I was irritated at the thread because 90% of it was theorising about some alternative situation where the data set was not free, or the training code was not free. That is not the situation we have with Leela Zero, we can skip that discussion and focus on what's actually at issue, i.e. the reproduction / verification of the training process.
- moviuro 8y agoThat's really just the reproductible build issue 2.0 I suppose no one could just lend you the infrastructure needed to rebuild the model (sponsors, or otherwise supporters of Debian?)
- chubot 8y agoWhat are the issues with reproducing the training process? It requires too many computing resources (which might require non-free cloud services), or it requires non-free software like GPU drivers, or both?
- infinity0 8y agoFor Leela Zero it's basically just the computing resources, the software is done using an abstract but openly-specified GPU execution system OpenCL, similar (as far as I understand) to an instruction set architecture (ISA). Whether the abstract execution system is implemented in a proprietary or free way is a separate issue. For example these days it's hard to execute x86 instructions without running some sort of proprietary microcode and/or management engine. For comparison, OpenCL has many competing implementations, including the free open source MESA graphics drivers.
- deleted 8y ago[deleted]
- NhanH 8y agoDoes "training process" concept in this discussion includes the whole process of starting from scratch and reproduce leela zero from tabula rasa, or is it just the training of the latest network from available training data?
- GistNoesis 8y agoHow do you deal with evolving software versions (and dataset versions, and build environment). Do you have to retrain the whole weights to make sure the training would results in exactly the same weights. Maybe some unit cases/non-regression tests between versions to check that there are no breaking changes and that results are fully deterministic?
- infinity0 8y agoThat's a very interesting question, especially since the training data is also (partly) generated by the software in a self-feedback loop. For the case of Leela Zero, most of it has been generated by Leela Zero itself based on its own algorithm, but some of the data has been supplied from the data set generated by the Facebook ELF Go engine. (They used it as a shortcut to catch up to the level of ELF, and then surpass it.) In order to reproduce the current best-weights then, one would have to record exactly what versions of Leela Zero and ELF was used to generate which subsets of data, and which subsets were used to create further subsets of the data. I don't think anyone has kept that information around, so I'd guess the current best-weights will not actually be reproducible, ever. In future, one can imagine other software that could keep track of this information, and then be actually able to reproduce a particular resulting set of weights. However let's step back a bit. On a high-level, nobody actually cares that the results are not fully deterministic, they only care that it is a faithful representation of what the source code does. This is true both for software determinism and for weights-model determinism. Being deterministic is a (relatively) easy property, which when we achieve it, allows us to verify that the results (binary software, trained weights) don't contain backdoors or other unpleasantness that's not visible in the source code. But the latter is what we "actually" care about. If we can achieve the latter property without achieving determinism, then we are also mostly satisfied. That would involve being able to examine the model directly and see what it does, and see that it doesn't contain backdoors or other things. I can't even begin to imagine how to achieve this, it is a hard problem and the solution to this, would also solve the criticism of these AI weights/models being opaque and not really contributing much to human knowledge. Determinism is still useful for other purposes though. If you know exactly how something was produced, you have much greater control and understanding of how to tweak it, which might actually help us with the aforementioned goal of deeply-understanding these weights from a human perspective.