4 ms·
Hi, author here. First, it's she, not he. Second, because I hadn't heard about that obvious test at all back then. I never pretended to do a superserious scie
by sole 14y ago
Hi, author here.
First, it's she, not he.
Second, because I hadn't heard about that obvious test at all back then. I never pretended to do a superserious scientific analysis but rather answer to the questions that came to my mind, by using a computer to validate hypothesis.
But many thanks for the pointers & suggestions, though! I was thinking about rebuilding this to make it realtime+interactive so that more than one text could be analysed/visualised, so I'll make sure to introduce the new tests.
- jmilloy 14y agoHave fun with it! This seems like a great example to encourage further analysis because you still don't have a testable quantitative hypothesis for why it was harder to read. Makes me want to start coming up with metrics, too! Of course there's a big list of readability metrics, but it's way more fun to discover them on (y)our own. That said, I too was a little surprised when you had a heading titled "The classic graphs" without actually having perspective in the field for what would be classic.
- scrumper 14y agoHi author, I really liked the visualisations, especially the black hole thing. I think I may have misunderstood the purpose of your article: I first read it as a serious attempt to use textual analysis to do a comparison of the comprehensibility of Tolkein's best-known works, in which case you fell way short of the mark by going no further than word counting. I think it started off like that, but on re-reading it I see that you say you hit a brick wall (paraphrasing) and decided to have some fun visualising your results so far; something you did very nicely. So perhaps I got the wrong end of the stick. I thought about what you've taken on here. Doing things at the word level is pretty easy. Taking into account grammatical structure to get things like sentence lengths and clauses per sentence, or breaking words up into syllables (needed for F-K, for example) is considerably harder. I'm interested to hear what you come up with. Flesch & Flesch-Kincaid are US inventions so perhaps not obvious to everyone. On he/she: I agonised over that for a good ten minutes over my breakfast. I read around your blog and your twitter page this morning for clues as to the appropriate pronoun. In the end I couldn't tell, so I went with 'he' because I stereotyped you. I nearly changed my wording to use constructions like "the author," "they," and various other mealy-mouthed alternatives but they were too ugly. So I did try, but I got it wrong. I apologise.
- Evbn 14y agoIsn't FK determined almost completely by sentence length? That is what I recall from messing with MS Word docs in high school.
- scrumper 14y agoFK combines average sentence length (total words/total sentences), average syllables per word and some fixed coefficients to come up with an equivalent school grade. Flesch Reading Ease score, which is what I actually meant, does the same but with different coefficients to come up with a more granular difficulty score, usually in the range of 30-100. They're both pretty arbitrary. The more I read up on this subject the more respect I have for the author's own attempts at an originality score. It's all subjective ultimately.
- saraid216 14y ago> On he/she While I consider defaulting to 'he' to be legitimate and acceptable, I actually prefer the zie/zir gender-neutral pronouns when I think about it. http://santiago.mapache.org/nonfiction/essays/zie.html http://santiago.mapache.org/nonfiction/essays/zie.html If I ever begin to agonize about the gender of the person I'm talking about, that's enough to kick me over into using GNPs.
- shardling 14y agoThe "originality" index bothered me, because as the work of a length grows, you'd generically expect less words to be introduced -- exactly what the results show. The idea makes sense, but I'm unclear on how to actually measure it in a way that's normalized by page count.
- Evbn 14y agoBinomial distribution mumble mumble logarithm of page count mumble.
- andreasvc 14y agoThere's the concept of a vocabulary growth curve. It shows the number of words occurring once as a function of the amount of text, e.g., how much new words in the first 1000 words? how much in first 2000 words? etc.
- shardling 14y agoBut is there a good way to boil that down to a single metric? (i.e., can we parametrize these curves with a single value?)
- andreasvc 14y agoI wouldn't know off the top of my head, but the concept was introduced by Harold Baayen. There are free PDFs online of his with lots of interesting quantitative techniques.
- hsmyers 14y agoNice work! So my question is where did you get the text to analyze?
- sole 14y agoYou can find the text for pretty much any popular work if you look in the proper places :-)
- ratzinho87 14y agoWould it be possible for you to generate a HD version of the "Graphical representation of words frequency" image? It would look great as a poster.