7 ms·
Big warning: their summary is "Among high-legibility fonts, a study found 35% difference in reading speeds between the best and the worst." This is completely
by ad8e 4y ago
Big warning: their summary is "Among high-legibility fonts, a study found 35% difference in reading speeds between the best and the worst."
This is completely wrong and comes from an abuse of statistics. See the original research at https://dl.acm.org/doi/10.1145/3502222#d1e6428 https://dl.acm.org/doi/10.1145/3502222#d1e6428
An understandable explanation: imagine having 5 dice. You roll each die 4 times, then compare the highest sum to the lowest sum. Then you report that the highest die rolls 35% higher than the lowest. This is what the authors did, with each die being a font. But the experiment does not actually show evidence of any difference. If you rolled 500 dice, this method could claim that the highest dice are 200% higher than the lowest, even though all dice are still equal.
The original authors seem aware of this shortcoming, but did it anyway: "we are somewhat stretching the applicability of a Cohen’s d analysis for this data". This is likely because they did not know of a better method. But it is wrong to be pushing this analysis. The main author is from industry, so perhaps he was unaware that this effect can be corrected for, or that this type of misleading claim is malpractice. But someone in the chain of publishing - the journal editors, the reviewers, the large author list, or Jakob Nielsen who is promoting this - should have caught this. It is their main result!
In the absence of legitimate statistics, the article's circumstances point to a failure to detect measurable differences between fonts. There are two ways fonts may be better than each other:
1. across the population, so that one font is better for everyone
2. personalized, so that different fonts are better for different people
The first case should be easily detectable, and the second case should see some correlation between preference and speed because of a familiarity bias (https://dl.acm.org/doi/10.1145/3502222#d1e7351 https://dl.acm.org/doi/10.1145/3502222#d1e7351). These are the Bayesian expectations I walked in with. Neither of these appear supported by the article, although I have only skimmed it.
To be clear, the experiment does not give evidence that fonts perform equally either. It looks more likely that the experiment design failed.
- wolverine876 4y agoReading the article, what valuable things can we learn? Imperfection exists in everything in the universe - even the ideas in our minds we sometimes imagine to be ideal and perfect, but they turn imperfect as soon as we write them down (this cursed keyboard!). Imperfection does not destroy all value, or we live caves and aren't communicating using an imperfect alphabet and language, over imperfect signals, using imperfect power, etc etc. (in fact, we would be dead in the caves). Reading and learning from imperfection is the defintion of 'reading and learning', because there's nothing else to read. If I listened to HN comments, very little research would have value, very little information would be worth reading. The top comment is almost always of this nature; it's depressing to me that it still happens; we all should be very familiar by now with social media and these patterns, yet we keep following them. > I have only skimmed it [the OP]. Maybe that should be at the top of the comment. Imagine an OP which presented a detailed analysis and then, at the bottom, said 'I only skimmed the thing I analyzed' - imagine what the top comment would say.
- ad8e 4y agoI object strongly to your comment. I find every part of it to be absurd. > Reading and learning from imperfection is the defintion of 'reading and learning', because there's nothing else to read. > If I listened to HN comments, very little research would have value, very little information would be worth reading. The top comment is almost always of this nature; it's depressing to me that it still happens; we all should be very familiar by now with social media and these patterns, yet we keep following them. Just as "laymen blanket dismissing productive scientists" is a tired and sad trend, so too is this - because now there is a delicious and ironic reversal. I am the productive scientist here, whose day job involves extracting the useful content from scientific articles by analyzing their methods and results. When I read, it is my pride to be able to determine when the imperfections of an article destroy its content, or if these imperfections can be worked around and to what degree. Like other academics who can do the same, I believe this is a skill that requires training and intelligence, and its outcome is to make subjects understandable that would otherwise be a maze of bad results. People like me are confident at playing the Many Labs replication guessing game. I read useful articles every day, and the original article is not one of those; it falls squarely in the "awful" bucket. I am not the uninformed layman making pithy snarks which have no specific relevance to the situation at hand - criticisms with general applicability, copied from elsewhere. You are. You are the layman complaining about what the actual scientists do and how they advance the field. Even your comment complaining about this trend is reflected adroitly upon yourself, in a very ironic manner. Notice how my original comment uses deep field-specific knowledge of fonts, science, and statistics to make my point, while you only use broad strokes about social media trends, as you lack such knowledge. In fact, this remarkable and joyous irony about a person complaining about uninformed criticism, while making exactly that uninformed criticism himself, was what led me to make this comment, as otherwise I would not have bothered to respond. > > I have only skimmed it [the OP]. > Maybe that should be at the top of the comment. Imagine an OP which presented a detailed analysis and then, at the bottom, said 'I only skimmed the thing I analyzed' - imagine what the top comment would say. My comment was based on the parts I read, and it still stands correctly - unless you have something meaningful to say? For example, if you believe yourself to also possess the skill to analyze articles and extract their value, it's open to you to do a complementary analysis. So far, you have not even addressed any of the points, only pushed some dismissals which are themselves easily dismissed: a programmer can contribute to the Linux kernel without having read every line. In fact, I looked up the original article's author, and his affiliation is not only with Adobe. He is also a PhD candidate at Brown University, under advisor Jeff Huang. This makes my opinion far more negative, but not of the PhD candidate, who only has my sympathy. Rather, his advisor bears the blame here, as the advisor has the responsibility to control ethics and give guidance on the unfamiliar. PhD students depend on their advisors to know the way forward, especially on statistics and scientific culture, which they cannot navigate the conventions of on their own. Jeff's behavior here evokes complex emotions in me, related to the ethics of science and responsibility, leading to this conclusion: Brown University should no longer allow Jeff Huang to advise students, and Jeff Huang's articles should be either ignored, or looked over by data thugs if their results need to be relied on. Or, if Jeff's incursion into experimental science is him trying something unfamiliar, then perhaps he too is a victim of his own ignorance and bears no ethical blame or responsibility, but he should stay away until he learns better.
- SamBam 4y agoPersonally, I think that you could deduce that the conclusions were likely invalid by simply looking at the ordering of the "best fonts" list. They're is simply no metric by which any reasonable person could arrange that set of fonts in that order. Serif vs sans? Random. Condensed vs spaced out? Random. Heavy vs light? Random. Obviously there can be subtle interactions that we haven't understood yet, but there is simply no hypothesis presented as to what features could be important. A study with "statistical significant" (perhaps) results and no hypothesis ought to at least be replicated before we even discuss it.
- smegsicle 4y agowell it doesn't seem completely off the wall- to my eye the top seven look pleasant to read, with the possible exception of oswald, where the bottom three- avenir next, avant garde, and open sans- look a tad obnoxious curious because open sans looks so normal, but yet so subtly bad ..
- SamBam 4y agoThe fact that Helvetica and Arial are far apart, yet completely indistinguishable by the average non-font-nerd is another indication to me that this ordering is fairly random.