3 ms·
You could measure the efficiency of a language by computing how many bits are required to store the semantics of a word/phrase/sentence losslessly. Assuming ide
by chewxy 3y ago
You could measure the efficiency of a language by computing how many bits are required to store the semantics of a word/phrase/sentence losslessly. Assuming ideal encoding to bits of course. Think of things like perplexity and the like.
However, traditional way of computing efficiency of compression would not be useful for a meaningful analysis of the efficiency of a language. Barring issues like having an ideal encoding to bits, or even having the concept of "efficiency" being rigorously defined, there are problems just from the outset.
Take context for example.
All useful compression methods have some sort of decompression key involved. This could be the dictionary, or the bitmap or the know-how (for cases like RLE). In natural langauges, the compression/decompression key is stored in a distributed fashion across the minds of a society.
"Darmok and Jalad at Tanagra" is a VERY efficient compression for what is presumably a very long story about two hunters who met at an island and fought a beast together, but it is only efficient to the people who speak that language. The "local" efficiency (to the population who speak the language) is very high, but the "global" efficiency isn't.
So we must account for efficiency in terms of the size of the compressed concept as well as the compression key. And from my experience, it's a sorta lumpy kinda world out there.
- jdmichal 3y ago> You could measure the efficiency of a language by computing how many bits are required to store the semantics of a word/phrase/sentence losslessly. Assuming ideal encoding to bits of course. Think of things like perplexity and the like. This has been done! The answer is about 39 bits a second. https://www.science.org/content/article/human-speech-may-have-universal-transmission-rate-39-bits-second https://www.science.org/content/article/human-speech-may-hav...
- tsukikage 3y agoIt is customary, when comparing performance of compression algorithms, to include the size of the tool needed for decompression in the compression benchmarks, since otherwise one can simply smuggle the uncompressed data in the decompression tool. ISTM a similar principle would need to apply here: learning the "Darmok and Jalad at Tanagra" language would involve absorbing many volumes of history and mythology where for the usual sort of language a dictionary, grammar reference and maybe a book of common idioms would suffice. Whatever metric is used to compare languages for efficiency should reflect this.
- shanusmagnus 3y agoYou can't store "the semantics" losslessly because you can't definitively say what the semantics of an utterance even are, unless you're using some reduced definition of the term, or a pre-selected frame, or a computer language.
- bentcorner 3y ago> "Darmok and Jalad at Tanagra" is a VERY efficient compression for what is presumably a very long story about two hunters who met at an island and fought a beast together, but it is only efficient to the people who speak that language. The "local" efficiency (to the population who speak the language) is very high, but the "global" efficiency isn't. I suppose image macros/memes are the modern equivalent. Social context enables readers to "decompress" the meme. [Drake top]: "Two hunters who met at an island and fought a beast together" [Drake bottom]: "Darmok and Jalad at Tanagra"