3 ms·
But you start with a constant cost of 8 bits. For 255 characters you lose 9.4 bits which is bigger than the cost of an 8 bit size field. For 65535 characters
by std_throwaway 11y ago
But you start with a constant cost of 8 bits.
For 255 characters you lose 9.4 bits which is bigger than the cost of an 8 bit size field.
For 65535 characters you lose 378 bits which is much bigger than the cost of a 16 bit size field.
On the other hand you have only one string type for arbitrarily sized strings and most strings are short.
What would really be needed to be more space efficient for short and long strings is a size field with a variable length encoding. E.g. 1 byte for string lengths up to 64, 2 bytes for string lengths up to 8k, etc.
- vardump 11y ago> What would really be needed to be more space efficient for short and long strings is a size field with a variable length encoding. E.g. 1 byte for string lengths up to 64, 2 bytes for string lengths up to 8k, etc. It's just decoding that requires pipeline busting branches. Well, I guess it's fine if you take branch mispredict for the large string case. Maybe 1 byte with highest bit zero, 0-127, 2 bytes for 128-32767, etc? The irony is of course now we have variable length encoding for the string size... I think just using size_t instead makes more sense than saving a byte or two per string. Memory is cheap, but CPUs aren't getting much faster. When size is an issue and nothing else prevents it either, just LZ4 (or similar memcpy order of magnitude speed compression) the string.
- std_throwaway 11y agoI think it depends on the use case. In embedded systems bits consume energy and therefore increase cost. The battery of a satellite is much more expensive than a Chinese coal power plant. Yes, "size_t" should be just about right for everybody except the guys who _need_ to program in C/C++ and Assembler due to the hard requirements of their systems. Not very many people but very important systems.
- Dylan16807 11y agoIf you want to go variable-width, probably best to keep it as simple as possible. 0-254 signal a short string, 255 signals that a size_t (or similar) follows. That way the overhead is never high, and you can handle the two formats with a very small number of instructions.
- vardump 11y agoYou'll have these duplicated compiler-inlined checks all over your binaries now, so your memory requirements might actually be higher. Start of your string data is now (almost) always unaligned, some penalty might apply. You'll potentially get pathological mispredicted branches when the string is being accessed.