4 ms·
From the FAQ: "The master list includes 60,469 words and 304,275 nonwords. From these you get a random sample of 70 words and 30 nonwords. The word list should
by larkeith 7y ago
From the FAQ: "The master list includes 60,469 words and 304,275 nonwords. From these you get a random sample of 70 words and 30 nonwords. The word list should contain almost all English words (non-inflected forms)."
So just multiply by 60k for how many you know. You've also got the info you need for stuff like stddev.
- ebg13 7y agoLol. "should contain almost"
- imtringued 7y agoThat's not how that works. To obtain an absolute number the words have to be ranked by frequency and then the 100 words have to be picked so that frequent words appear first and less frequent words appear at the end. The estimator then knows whether you know a 90%, 99%, 99.9% or 99.9% percentile words and uses that to estimate that you know enough words to cover x% of all words that will appear in a text. Then you use that percentile on your word ranking to count how many words are below that percentile. The posted test site does nothing of the sort. Imagine the test only asked you a single 90% percentile word. vocabulary.ugent.be would tell you: You know 100% of the words. The process I described would tell you: You know 90% of the words. Obviously both tests become more accurate as the number of words increases but the vocabulary.ugent.be test requires you to test all 60k words to get an accurate result. The second test will probably need less than 1000 words to be dead on. http://testyourvocab.com/ http://testyourvocab.com/ seems to be a pretty good implementation of the approach I described.
- larkeith 7y agoThe test you described is useful for determining how many words in a text you will know - i.e. what percentage in a given set of English. That's probably a more useful number, but it's not what parent asked: "I was hoping for [...] how many English worsd [sic] I know." > Obviously both tests become more accurate as the number of words increases but the vocabulary.ugent.be test requires you to test all 60k words to get an accurate result. They could never get an accurate result for words known in a body wih this test - it needs a separate dataset of word frequency. That's not what they're testing.