3 ms·
This actually sparked an idea for me. Could code complexity be measured as cumulative entropy as measured by running LLM token predictions on a codebase? Notabl
by Scene_Cast2 2y ago
This actually sparked an idea for me. Could code complexity be measured as cumulative entropy as measured by running LLM token predictions on a codebase? Notably, verbose boilerplate would be pretty low entropy, and straightforward code should be decently low as well.
- jeffparsons 2y agoNot quite, I think. Some kinds of redundancy are good, and some are bad. Good redundancy tends to reduce mistakes rather than introduce them. E.g. there's lots of redundancy in natural languages, and it helps resolve ambiguity and fill in blanks or corruption if you didn't hear something properly. Similarly, a lot of "entropy" in code could be reduced by shortening names, deleting types, etc., but all those things were helping to clarify intent to other humans, thereby reducing mistakes. But some is copy+paste of rules that should be enforce in one place. Teaching a computer to understand the difference is... hard. Although, if we were to ignore all this for a second, you could also make similar estimates with, e.g., gzip: the higher the compression ratio attained, the more "verbose"/"fluffy" the code is. Fun tangent: there are a lot of researchers who believe that compression and intelligence are equivalent or at least very tightly linked.
- 8note 2y agoInterpreting this comment, it would predict low complexity for code copied unnecessarily. I'm not sure though. If it's copied a bunch of times, and it actually doesn't matter because each usecase of the copying is linearly independent, does it matter that it was copied? Over time, you'd still see copies being changed by themselves show up as increased entropy
- malfist 2y agoCode complexity can already be measured deterministically with cyclomatic complexity. No need to use an AI fuzzy logic at this. Especially when they're bad at math.
- contravariant 2y agoThere's nothing fuzzy about letting an LLM determine the probability of a particular piece of text. In fact it's the one thing they are explicitly designed to do, the rest is more or less a side-effect.
- david-gpu 2y ago> Could code complexity be measured as cumulative entropy as measured by running LLM token predictions on a codebase? Notably, verbose boilerplate would be pretty low entropy, and straightforward code should be decently low as well. WinRAR can do that for you quite effectively.