3 ms·
The term "absorbed" was not for the people in the field, but for people who don't know what folding means. IMHO it's a better metaphor then "learning", because
by Shamar 5y ago
The term "absorbed" was not for the people in the field, but for people who don't know what folding means.
IMHO it's a better metaphor then "learning", because learning is a _subjective_ experience that everyone does and using that term lead inevitably to anthropomorphisation.
"Absorb" match the insight of filters and pipelines, that can be easily understood from any CS student, any "ML expert", any lawyer and any other citizen.
____
As for the network, my argument is simple: if I get back the source dataset from the executable, I think we can agree that such dataset is projected on the numerical matrices that such executable record.
Now where is the dataset?
You might argue that it is recorded _only_ into the gradients logged there (the gradients applied to one single "neuron" for each "layer"), but if so you could reconstruct the source dataset from the logs alone, and in fact, you cannot. You need both the "model" and those gradients in the correct order (and the encodings of inputs and outputs, obviously).
You might ask: "fine, but how much of the source dataset is projected into the gradients and how much is projected into the model?"
To answer, we need to consider that
- the vector space that constitutes the executable is non-linear (the "model" part) and hierarchical (the vectors of the gradients are not independent neither between layers nor between samples)
- (initialization apart) all the values (and the operative value) that the "model" contains comes from the source dataset
Thus I argue that a substantial portion of the source dataset is contained in the "model".
This does not exclude that another substantial portion of the source dataset is also contained into the few logged gradients!
And in fact I've never stated that the "model" contained the whole source dataset.
But if the portion contained into the "model" was negligible, you would be able to get back the sources from those logged gradients alone with negligible errors.
AFAIK, it is not possible, but if you can, please teach me how!
I'm always more than happy to be proven wrong if I can learn how to do something that I previously thought impossible!
- r-zip 5y ago> And in fact I've never stated that the "model" contained the whole source dataset. Apologies for seeming rude, but I feel that the abstract is disingenuous (I'm assuming you're the author of the article). The abstract states: > we provide a ... decompiler that reconstruct the source dataset from the cryptic matrices that constitute the software executed by them. But that's not what's happening here. Instead, what's happening is (correct me if I'm wrong) the decompiler uses the gradient information along with the network itself (which is very close to the penultimate network) to reconstruct the input. If we consider for instance, MSE loss, that reconstruction appears trivial given all the information available. As I said before, this reconstruction (while interesting) does not show copyright violation, because the training process information is not available once the network is deployed. Obviously the model contains information about its training set. If it didn't, it would be useless. I'm not saying there aren't obvious copyright issues, and I'm also not saying that the approach to recover the training set is not interesting. I'm just saying that the overall copyright argument has a major gap (and there are more direct alternative arguments).