3 ms·
One problem that you will run into is that X (the phrase in the comment) might be a permutation of a book title. Or perhaps even a title of a book being mention
by demonshalo 9y ago
One problem that you will run into is that X (the phrase in the comment) might be a permutation of a book title. Or perhaps even a title of a book being mentioned in a comment where the author did not intend on it being interpreted as such. This is mostly due to ambiguity & information density in language as well as how we assign names/titles to things.
Ex. If I write a comment saying "The number of our clients went from zero to one instantly", in this sentence, "Zero to One" will match Peter Theil's book.
If you are willing to put up with such "noise" in the output, then you don't need to train a thing or even use ML. Just chunk the given piece of text and look up all the tokens permutations of given length (from 1 up to N) in your SQL Entities Database.
You can generally find this kind of Entities Databases by aggregating a lot of datasets from all over the web or even use something like google books datasets:
https://books.google.co.uk/ https://books.google.co.uk/
https://storage.googleapis.com/books/ngrams/books/datasetsv2.html https://storage.googleapis.com/books/ngrams/books/datasetsv2... - This is an ngram set but there should be one with only book titles somewhere. Can't find it atm but you can search for it yourself!
You could even (if you are brave enough) try to use wikipedia's dumps to mine book titles from articles