3 ms·
So, I am specifically talking about the scenario where all your labelling functions are highly correlated and there is little or no ground truth data to come up
by tadkar 6y ago
So, I am specifically talking about the scenario where all your labelling functions are highly correlated and there is little or no ground truth data to come up with empirical weights for each of the labelling functions. An example is the scenario where you have the label functions: x>5, x>4.99, x>5.01 for some feature x. I am really struggling to see how Snorkel can correct for the correlation, especially given the relatively simple generative model in section 2.2 of the paper. https://arxiv.org/pdf/1711.10160.pdf https://arxiv.org/pdf/1711.10160.pdf
- polm23 6y agoThe Snorkel paper doesn't cover this in depth, the math is all in this paper: https://arxiv.org/abs/1703.00854 https://arxiv.org/abs/1703.00854 I can't say I followed all the proofs, but it seems that under certain limited assumptions about labelling functions they prove their generative function can do well. Reading Snorkel it initially sounded like magic in the bad way, but this does make it clear that if your labelling functions are garbage or have certain kinds of problems there's nothing they can do about it. Even leaving aside the generative model I think the focus on function-based data bootstrapping is great, which is why I've been following Snorkel's projects for a while.