3 ms·
The purpose of this metric is to give authors a practical tool they can use to measure the degree of direct sensory language they use in their writing, and comp
by benjismith 8y ago
The purpose of this metric is to give authors a practical tool they can use to measure the degree of direct sensory language they use in their writing, and compare it with authors they admire. It's not a value judgement.
There are other admirable qualities of writing (emotional, conceptual, etc) that are orthogonal to "vividness", and in the long-run, I plan on developing metrics for those qualities as well.
For example, take a look at the works of Jane Austen. Here's the analysis of "Pride and Prejudice":
http://prosecraft.io/library/jane-austen/pride-and-prejudice/ http://prosecraft.io/library/jane-austen/pride-and-prejudice...
She's a brilliant writer, and her prose is highly emotional, but it isn't especially vivid, according to my definition of "vividness", which is: prose that evokes a sensory experience (with colors, textures, flavors, aromas, sounds, and bodily sensations).
There are plenty of ways to write a brilliant novel, and some of them involve vivid sensory writing. But Jane Austen's brilliance comes from her handling of emotional relationships.
For a modern example of the same phenomenon, take a look at one of my favorite authors, Nick Hornby:
http://prosecraft.io/library/nick-hornby/about-a-boy/ http://prosecraft.io/library/nick-hornby/about-a-boy/
http://prosecraft.io/library/nick-hornby/juliet-naked/ http://prosecraft.io/library/nick-hornby/juliet-naked/
I've read every one of his novels, and they're all about human relationships, but the prose itself isn't very vivid. Nothing wrong with that, though. It's just a measurement.
As an author, it's helpful to be mindful of these kinds of measurements. The same thing is true of "passive voice". Using a lot of passive voice is still a legitimate way of writing, but it's helpful for an author to be aware of the literary voice they're crafting:
https://blog.shaxpir.com/thoughts-on-passive-voice-705fa4dbd291 https://blog.shaxpir.com/thoughts-on-passive-voice-705fa4dbd...
- ajmarcic 8y agoMy impression is that your "vividness" metric is closed source. [0] Your metric is wholly subjective without a derivation and formula. We have no clue what's being measured. Your results are susceptible to "Yeah, well, that's just like your opinion man". [0] https://blog.shaxpir.com/writing-vivid-prose-33283e861358 https://blog.shaxpir.com/writing-vivid-prose-33283e861358
- benjismith 8y agoThe formula is easy, as explained in the article: 1) For any word in the 10,000-word vividness dictionary, add its score to the sum. 2) Divide by the total word count. The complete word list in the vividness dictionary, and the scores of each word, are in constant flux, based on the results of a massive machine-learning algorithm, driven by the set of novels in the corpus. Every time we add new novels, the results change slightly, but as the corpus, but the word-list and scores will eventually converge, and perhaps then we'll publish the dataset :)
- ajmarcic 8y agoThe formula involves: 1) A mystery heuristic to arrive at your 10,000 'vivid' words 2) A mystery 'vivid' word numerical rating "...depending on the intensity of the sensory experience it invokes" 3) A mystery "linguistic algorithm"/"massive machine-learning algorithm" Imagine clicking on an HN post titled "I made a raytracer in Python". Imagine the post contains pretty examples of the renders. Unfortunately it describes the raytracer as "cutting edge" but has no code for repeatability. Worse yet, the post doesn't recount any critical reasoning employed in the process of building a raytracer.
- benjismith 8y agoThe algorithm for building the lexicon and scoring the vividness of each word is still evolving pretty rapidly, so I'm not quite ready to publish the exact details yet, but I'll eventually write about it in-depth... Until then, here's a high-level overview: 1) Start with a human-curated list of several hundred vivid words. Be sure to include words that invoke all the senses: sight, sound, touch, smell, taste, and bodily sensation. These are the "seed words". 2) Scan through the entire corpus and find all instances of those seed words. 3) Vivid words tend to occur in clusters, so find all words that tend to occur in close proximity to the original seed words. 4) Create a list of new candidate words, and suggest them to a human reviewer. 5) The human review accepts or rejects each of the candidate words. 6) The accepted candidate words become new seed words for the next iteration. Eventually, you'll have somewhere in the neighborhood of 10,000 words :) 7) The score of each word is based on a modified TF/IDF metric, with a few extra finishing touches (like incorporating sentiment scores from the "Hedonometer" project at the University of Vermont Computational Story Lab). That's where the algorithm is at right now. I've been iterating on this basic premise for over a year, and I'm pretty happy with the current results. It's not perfect, but it's useful. As far as I'm concerned, it's a pretty mature "version one" of the vividness metric. But it has several weaknesses I want to address in a "version two" sometime soon. The biggest weakness is that TF/IDF is only a shallow proxy for intensity of vividness. For example, looking through the 500 million words in the prosecraft corpus, the word "blood-red" occurs 972 times, yielding a vividness score of 6.1. But the word "magenta" occurs only 515 times, yielding a vividness score of 8.3. It's more rare, so the model thinks it's more vivid. To me, that seems wrong. Because the word "blood" has special bodily meaning, beyond just a prefix for a color-word, my intuition says the word "blood-red" should be scored as significantly more vivid than "magenta". I have some ideas for addressing that deficiency by using the output of "version one", alongside a new "sensory hierarchy" model to train a new "version two" classifier. But that new model is still in its very early conceptual stage, and I'm not ready to write about it yet :) In the meantime, I consider "version one" a legitimately useful tool for working authors. Glad you asked!
- theluketaylor 8y agoIt's interesting you use Jane Austen as an example, as I have found her books to be filled with more than enough detail to picture the scenes quite distinctly. I find it particularly interesting that your algorithm has identified the argument between Elizabeth and Lady Catherine as the most passive page, when in fact it's one of the most tense exchanges between any two characters and is dripping with sarcasm on both sides. They use the passive voice to insulate the argument and add a veneer of social acceptability. Have you done any analysis to see how your tools handle irony? The description makes it seem words are taken at face value, but really good prose usually operates on several levels.