4 ms·
I _love_ to see projects like this. I was working on a similar problem this weekend, but with whole words instead of abbreviations I had made in a dictionary,
by zerojames 3y ago
I _love_ to see projects like this.
I was working on a similar problem this weekend, but with whole words instead of abbreviations I had made in a dictionary, and in the general case, fine-tuned on any given corpus of text.
I wanted to know: could I write an autocorrect that is "fine-tuned" on a given corpus of text? Use case: I write a lot of docs with long phrases (i.e. "data augmentation"). Could I automate them?
I arrived at:
1. Calculate "surprisal" of unigrams and bigrams (entropy) from a general dataset (an NYT corpus), give a boost to words in the "fine-tuned" index;
2. Create a trie data structure that is weighed by surprisals. The more surprising a word, the more weight it gets.
3. Use that as advanced autocomplete.
I got a working solution here: https://github.com/capjamesg/autowrite/blob/main/autocomplete.py https://github.com/capjamesg/autowrite/blob/main/autocomplet...
(No docs yet -- coming in the next few days. Leave a GitHub Issue if you want to chat about it!)
- itake 3y agoThis sounds very similar to: https://github.com/wolfgarbe/SymSpell https://github.com/wolfgarbe/SymSpell Or maybe I am confused what you're saying. Are you trying to find more creative ways of saying the same phrase?
- zerojames 3y agoI definitely need to write more docs on this project! I will share them when I'm done. It looks like SymSpell is doing spelling correction, which is part of what my code does. The "surprisal" (information theory: Shannon information) component is doing statistical next word prediction. For example, given a corpus of data in which "data augmentation" is a common phrase, "da" could complete to "data augmentation". "data augmentation" starts with "da", and would be determined as a likely candidate for the next word because it was common in the dataset on which the statistical model was "fine-tuned". This Wikipedia page covers the concept in more depth: https://en.wikipedia.org/wiki/Information_content https://en.wikipedia.org/wiki/Information_content
- zerojames 3y agoBy the way, this project is _awesome_! My programming language [1] does OCR correction via heuristics (possible because there is a restricted grammar) but this project looks like it could help _a lot_ with improving the accuracy of the OCR output. Thank you! [1] https://visionscript.dev/paper/ https://visionscript.dev/paper/