4 ms·
I would do something like this: 1. Get a few thousands book titles that are "long enough" so that the probability they appear in a sentence without meaning the
by ar7hur 9y ago
I would do something like this:
1. Get a few thousands book titles that are "long enough" so that the probability they appear in a sentence without meaning the book is low
2. Search these spans in HN comments and use this corpus to train Stanford's CoreNLP NER
3. Run the NER on all comments
4. Check on Openlibrary or another book DB that the extracted spans are real books titles
- mfalcon 9y agoThis is known as bootstrapping or semi supervised learning, just in case you want to look for some theory behind it.
- achompas 9y agoHow so? I don't see reference to sampling with replacement or a suggestion to re-use the unsupervised results to further improve the model. Seems like I'm missing something...
- mfalcon 9y agoYou're creating a small classified dataset without manually labeling them, in order to train a supervised learning model. I'm not an expert, but I think that a boostraping technique doesn't imply continually improvement of the model.
- achompas 9y agoAhh, I see. I was confused about whether semi-supervised approaches rely on using predictions on the unlabeled data to improve model performance. Wikipedia seems to suggest this is a key component which isn't mentioned in OP: > Semi-supervised learning may refer to either transductive learning or inductive learning. The goal of transductive learning is to infer the correct labels for the given unlabeled data. Agreed on bootstrap, but in the proposed approach you're not artificially expanding your sample size by sampling with replacement.