6 ms·
really surprised this isn't getting more action on HN. the implications for CS are immense.
by oneJob 11y ago
really surprised this isn't getting more action on HN. the implications for CS are immense.
- andolanra 11y agoNo, the implications are actually quite small. This article is a very bad pop-science summary of relatively banal linguistic research. Linguists have studied linguistic universals for a long time, which are properties that all human languages have. For example, one could try to imagine (in the style of Borges) a language which had no nouns, and in which all sentences are formed of relationships between verbs—but no natural language has this feature: all natural languages have nouns and verbs. There are also implicational universals, which are of the form, if [some language] has property X, then it will also have property Y, and tendencies, which are broad driving trends that might have individual exceptions. An example of the latter is that languages that place the verb at the end of the sentence usually have postpositions rather than prepositions, but this has exceptions (e.g., Latin.) What's being studied here is a tendency in sentence structure: languages usually structure their syntax such that they can minimize the dependency length, or the distance between syntactically releated words in a sentence. This has long been hypothesized, but this paper gives evidence for it in the form of a large cross-language survey. Which is cool! But by no means does it have major implications for CS in any way. (At least, no more than any of the copious previous research on linguistic universals.) EDIT: I should also add that this area of research is not new. In fact, linguist Joseph Greenberg published an article called 'Some universals of grammar with particular reference to the order of meaningful elements' in 1963. This is continuing research and, while good research, not particularly groundbreaking or pioneering.
- AnimalMuppet 11y agoI think it has implications for language design. Some syntaxes are going to be better than others, based on this. It also means that, for a function that takes several parameters, some parameter orders are better than others.
- canjobear 11y agoYes, specifically you should order arguments so the order will be on average from short to long. I try to write my code this way. map, filter, and reduce are terrible from this perspective. Unless you have do blocks like in Ruby or Julia! Also, dplyr's %>% pipe operator is a great way to reduce dependency length in R code.
- jonahx 11y agoThanks for all your answers here! Could you elaborate on the above? Why does it imply short to long orderings are better? Also, I don't follow your comment about map/filter/reduce being bad except unless you have ruby-esque do blocks. Are you referring to something like map(<big function>, array)?
- canjobear 11y agoYeah, map(<big function>, array) creates a dependency that exists from when you read "map" to when you read the name of the array, potentially spanning a very long function. But if you have map(array) do <big function>, then you only have a dependency from "map" to the beginning of the function. In general, if you have a function call f(a, ..., y, z), when you parse that (mentally, or in a shift-reduce parser) you have to keep the function name f in memory all the way to z. So you want to make a, ..., y as short as possible. Similarly, dependency length minimization predicts that in English people will want to order expressions from short to long after a verb or preposition. There is a lot of evidence for this preference; it's been documented since the 1930s. If there were a programming language where the function name came after the arguments, like (a, b)f, then the best order would be long-to-short. Similarly, the DLM prediction for verb-final languages like Japanese is that people will prefer long-to-short orders. It appears that this preference does exist, but it is much weaker than the short-to-long preference among speakers of English-like languages.
- andolanra 11y agoThere's a hidden premise in your argument that I would argue is false, and that's that programming languages should attempt to emulate the same principles as natural languages. Programming languages and natural languages do not at all fill the same niche, and there are many cases in which natural languages optimize for things that would be bad in programming languages. For example, natural languages are infamously redundant—for example, gender agreement between nouns and adjectives and even (in some languages) verbs—but that's because they developed so that they could be understood even if you were shouting over the wind or otherwise didn't hear part of the sentence. Programming languages have no such restrictions, and as such, optimizing a programming language for the same kind of redundancy as a natural language would lead to needless tedium like int x = int_addition(int 2, int 3); but in the context of a programming language, this kind of redundancy ends up being needless bookkeeping without presenting any of the same advantages of redundancy in natural language. That doesn't mean that your conclusions are wrong—I think some parameter orderings are better than others! But I think that's true for reasons orthogonal to the findings in this paper.
- canjobear 11y agoThe cool thing about dependency length minimization is that you can use it as a principle to derive many of Greenberg's word order universals, in addition to sentence-by-sentence preferences. You can also use it to derive the fact that natural language expressions are usually well-nested (though I'm somewhat dubious: it seems like there are other good possible explanations).
- shakethemonkey 11y agoIt's much older than 1963. Erasmus was researching this in the 15th/16th century, but John Wilkins' 1668 book is probably the most well known of the early works[1]. [1] https://books.google.com/books?id=BCCtZjBtiEYC https://books.google.com/books?id=BCCtZjBtiEYC
- andolanra 11y agoWilkins' book describes a universal language, which is not at all the same thing as linguistic universals. Wilkins was trying to derive a language in which each word functioned as an index into a universal ontology of concepts, so that the concept represented by a word could be deduced by breaking apart the structure of the word itself. This is an interesting (if quixotic) experiment, but it's really concerned with building an a priori language. The study of linguistic universals is the study of properties of natural languages: for example, all languages have pronouns is a linguistic universal, because it is a property that is true of all natural human languages. This is clearly not something that Wilkins was working towards: he was building a new language for the purpose of perfecting and clarifying communication. His Real Character had little—if anything—to do with analysis of the properties of natural language, and therefore also has little to do with the study of linguistic universals.
- ars 11y agoNo, not really. They did not look at enough languages for this to be anything except a starting point for additional research. As it is you can not draw any conclusions from this.
- oh_sigh 11y agoWhy don't you tell us what the implications for CS are?
- crimsonalucard 11y agonah, the implications are mostly biological. The discovery itself only implies that the language module in our brains are bounded by identical rules across all people. Therefore the discovery leads only to greater insight toward an existing biological system rather then a fundamental theoretical property of language. As a result, the implications aren't as closely intertwined with CS.
- coldtea 11y agoWhat implications for CS? Even the implications for the study of natural languages are not that great (and oversold by the article).