3 ms·
My area of expertise isn't AI, so correct me if this is wrong, but isn't the important part the model? Would it be possible to open-source the model without rev
by bjacobel 10y ago
My area of expertise isn't AI, so correct me if this is wrong, but isn't the important part the model? Would it be possible to open-source the model without revealing the data which it was trained on?
(For applications where the problem you're trying to solve is relatively similar to the ones Apple/Google/Amazon are trying to solve - conversational AI)
- _dps 10y ago[AI/ML researcher/engineer here] At best you can partly do this, and even that is debatable. Here's a simple example: say Google builds a linear regression for the incremental likelihood to click a social post based on the distance of poster from you in a social network. By merely revealing the (single) coefficient in this model, they've revealed something about the underlying data set that they may not have intended. It gets exponentially (I use this term deliberately) more difficult with multivariate nonlinear models. I have no doubt that substantial information about a data set could be reverse engineered from complex ML models, even if we don't yet know how. But even in the case of basic multivariate regression, the model directly encodes both the mean values and the covariances of all the variables. So that's a fair amount of disclosure already. In some sense all models are (usually lossy) data compression, and it's just a matter of understanding the compression logic to go backwards from facts about the model to facts about the source data.