13 ms·
Using this methodology more generally, it would be interesting to use NLP to identify the part of speech that is redacted to narrow the word search space. In t
by mlady 9y ago
Using this methodology more generally, it would be interesting to use NLP to identify the part of speech that is redacted to narrow the word search space.
In this example, the POS would be an adjective, and since the subject noun is plural, it would be more likely the adjective is a number
- EGreg 9y agoWhat?
- probably_wrong 9y agoNLP: Natural language processing. Part of speech (POS): the elements that, when combined, form a sentence (adjectives, verb, substantive, etc). What the OP suggests is to run existing tools to identify what POS are missing (as in "in this blank space you could only have a substantive and two to three adjectives"), and use that to reduce the search space.
- abpavel 9y agoOr better yet, rank the resulting sentences for each of the possible fits using existing speech API, setting a cutoff to filter out nonsensical results. This might even yield surprises.
- FLUX-YOU 9y agoYou can't really predict if there's going to be an aside in the redacted sentence though.
- FLUX-YOU 9y agoYou might get some out of predicting the first and last words and word type of a redaction (based on the words next to the redaction), but it's only cutting down on brute force space. That makes short redactions more dangerous for declassification than entire paragraphs as a general rule because you have no context to start from, but that's probably common sense to people doing redactions.