2 ms·
[AI/ML researcher/engineer here] At best you can partly do this, and even that is debatable. Here's a simple example: say Google builds a linear regression f
by _dps 10y ago
[AI/ML researcher/engineer here]
At best you can partly do this, and even that is debatable.
Here's a simple example: say Google builds a linear regression for the incremental likelihood to click a social post based on the distance of poster from you in a social network. By merely revealing the (single) coefficient in this model, they've revealed something about the underlying data set that they may not have intended.
It gets exponentially (I use this term deliberately) more difficult with multivariate nonlinear models. I have no doubt that substantial information about a data set could be reverse engineered from complex ML models, even if we don't yet know how. But even in the case of basic multivariate regression, the model directly encodes both the mean values and the covariances of all the variables. So that's a fair amount of disclosure already.
In some sense all models are (usually lossy) data compression, and it's just a matter of understanding the compression logic to go backwards from facts about the model to facts about the source data.