9 ms·
One of the things I’m really curious about is how Snorkel deals with poor labelling functions. More generally, labelling functions are another data source for t
by tadkar 6y ago
One of the things I’m really curious about is how Snorkel deals with poor labelling functions. More generally, labelling functions are another data source for the model and are just as susceptible to corruption and other real world issues like completeness, bias, counter factual issues and repetition. Perhaps even more so because these are manually constructed. For example, you can imagine that the person writing labelling functions writes effectively the same rule many times over. My understanding of the paper is that Snorkel would then weight this repeated labelling very heavily.
I think weak supervision techniques (at least the ones under the Hazy Research umbrella) require a degree of skill in machine learning that is easy to underestimate if you just think about the problem as an issue of domain understanding (or writing labelling functions in Snorkel terms)
- sanxiyn 6y agoMy understanding is different. I think you are talking about correlated data source. Snorkel's surprising innovation is that it does NOT overweigh correlated data source.
- tadkar 6y agoSo, I am specifically talking about the scenario where all your labelling functions are highly correlated and there is little or no ground truth data to come up with empirical weights for each of the labelling functions. An example is the scenario where you have the label functions: x>5, x>4.99, x>5.01 for some feature x. I am really struggling to see how Snorkel can correct for the correlation, especially given the relatively simple generative model in section 2.2 of the paper. https://arxiv.org/pdf/1711.10160.pdf https://arxiv.org/pdf/1711.10160.pdf
- polm23 6y agoThe Snorkel paper doesn't cover this in depth, the math is all in this paper: https://arxiv.org/abs/1703.00854 https://arxiv.org/abs/1703.00854 I can't say I followed all the proofs, but it seems that under certain limited assumptions about labelling functions they prove their generative function can do well. Reading Snorkel it initially sounded like magic in the bad way, but this does make it clear that if your labelling functions are garbage or have certain kinds of problems there's nothing they can do about it. Even leaving aside the generative model I think the focus on function-based data bootstrapping is great, which is why I've been following Snorkel's projects for a while.