4 ms·
IMO, uncompressed bytes is a better representation, because it can be used to compare relative expressive power for the particular problem. I'd bet Python clean
by morepedantic 3y ago
IMO, uncompressed bytes is a better representation, because it can be used to compare relative expressive power for the particular problem. I'd bet Python cleans house here, but the write-only languages are a wild card.
- dawnofdusk 3y agoWhy would uncompressed bytes be better? Using a good compression algorithm better approximates the statistical entropy of the code which is at least correlated with e.g., Kolmogorov complexity.
- iopq 3y agobecause it often sits uncompressed on the drive
- deleted 3y ago[deleted]
- morepedantic 3y agoBecause humans read and write the uncompressed code. Gzip will hide problems like copy paste.
- igouy 3y agoAnd hide differences due to label-length personal-preferences.
- lifthrasiir 3y ago"Expressive power" is a very subjective term, and uncompressed size is a bad proxy as it includes too many variables specific to coding conventions. Compressed size with a stupid enough algorithm (here gzip) is meant to reduce these variables. The true Kolmogorov complexity in comparison can't be computed, and too smart algorithms can start to infer enough about the language itself.
- morepedantic 3y agoKolmogorov complexity is absolutely the wrong metric. It doesn't account for big-O, timing, and many other production requirements. You'll never have a perfect metric here, but human readable size of code base is well justified. Do you write minified javascript?
- shpx 3y agoWhat I'm asking about is how much energy and time (computation) by a human brain it takes to emit or ingest each program because a human working 40 hour weeks from 18-65 will have 100,000 hours of working time, which at a typing speed of 250 characters per minute and a reading rate of 1500 characters per minute is a total career budget of like a billion characters emitted and 10 billion characters ingested. For emitting programs, if we assume the program is already fully formed in my brain and I'm just transcribing it, then we would like an accurate physical model of my hands moving over a QWERTY keyboard that can tell us how many joules and milliseconds I will use for each keystroke to type the given sequence of symbols. Ideally we would have a complete model or simulation of my brain so we can measure how many neurons need to fire for each keystroke as well. We don't have that, so we could measure my average milliseconds per character and multiply it by the total count of characters, but a language that is just made up of one symbol function names is probably harder to type than one made out of English words and there's also all the typing mistakes I will make. This is what the simple compression algorithm (gzip) is an attempt to normalize for, that a more verbose but more predictable language is as fast to write and read than an overly terse one. gziping is an imitation of a complex model.
- igouy 3y agohttps://benchmarksgame-team.pages.debian.net/benchmarksgame/how-programs-are-measured.html#source-code https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
- lifthrasiir 3y ago> Do you write minified javascript? Yes [1] [2] [3]? [1] https://js1024.fun/demos/2020/46/readme https://js1024.fun/demos/2020/46/readme [2] https://js1024.fun/demos/2022/18/readme https://js1024.fun/demos/2022/18/readme [3] https://github.com/lifthrasiir/roadroller/blob/442caa4/index.mjs#L1106-L1225 https://github.com/lifthrasiir/roadroller/blob/442caa4/index... Jokes aside, human doesn't read each (uncompressed) byte anyway. The number of tokens would have been much better than the number of bytes, but even this is unclear because a single token can have multiple perceived words (e.g. someLongEnoughIdentifier) and multiple tokens can even be perceived as a single word for some cases (e.g. C/C++ `#define` is technically two tokens long, but no human would perceive it as such). I would welcome a more realistic estimate than the gzipped size, but I'm confident that it won't be the number of uncompressed bytes.