5 ms·
In the article he says it is just one byte extra, but that's clearly not the case, since we probably want strings longer than 255 characters. To be practical, e
by supersillyus 15y ago
In the article he says it is just one byte extra, but that's clearly not the case, since we probably want strings longer than 255 characters. To be practical, even in the sort term, you'd probably need at least two bytes, and that's still a bit limiting. NUL-terminated strings keep you from having to worry about the size of the length specifier and what byte ordering when storing it.
He notes that other languages of the day didn't go with NUL-terminated strings. It could also be interpreted, considering that C is more widely used still than other languages from the time, that by doing something different, C made the right choice.
- spullara 15y agoWhat you would likely do is use an extension bit so there would be no fixed maximum length for the strings. This would of course add a bunch of overhead that you may or may not make up with the other advantages of knowing the length of strings.
- tptacek 15y agoThis is what DER does to encode ASN.1, and it is basically a nightmare of epic proportions. The 16-bit tag at the beginning of strings is a reasonable suggestion. 65k charstars are slow and unwieldy anyways.
- coldnose 15y ago* If you're appending data to a string, and the variable-length increases by a byte, will you have to memmove() the entire existing string down 1 byte? * Is every programmer responsible for detecting this condition? * (This will make manual string manipulation very complicated and dangerous.) * Suppose you're concatenating two strings, such that the sum of the lengths requires an additional byte. This could cause a buffer overflow. How would a strcat() function avoid causing a buffer overflow here? * Does every string need a maximum-length counter too? * Can you access a random element in the string without having to dereference and decode the length? On the other hand, if you use a constant-sized length, * What happens when you overflow the maximum length? * Can you erroneously create a shorter string by appending text? * How should string libraries handle this condition? By abort()ing? By returning a special error code? Does every string manipulation need to be wrapped in an if() to detect the special error? How should the programmer handle this condition? In either case, * Can you tokenize a string in-place? * Can an attacker read a program's entire address space by finding an address that begins 0xffffffff and treating it as a string?
- spullara 15y agoI think most of these are answered with the suggestion that no one would have made strings this way without making a matching library that handled the concerns you have here. Honestly char[] strings are much less useful in a UTF-8 world anyway.
- shaggyfrog 15y agoI read it as "one byte more than just using one byte for NUL". In other words, two bytes.
- burgerbrain 15y agoEven so, 2^16 chars is still absurdly limiting. Furthermore, address/length pairs complicate otherwise very simple programming tasks such as creating string tokenizers (with NULL-terminated strings, you just need to drop in a \0 where the deliminator was).
- astral303 15y agoI disagree that 2^16 chars is absurdly limiting. I remember coding in some language that had a 255-char String limit (maybe that was some kind of Pascal) and while that was somewhat limiting, it was not an issue in some 90% of the strings. 2^16 pretty much takes care of 99% of string usage, especially in the earlier days of programming languages. Anything over 2^16 and using a more specialized data structure for a buffer would've probably been more than acceptable.
- burgerbrain 15y ago2^16 bytes is a mere 64KB. Sure you can get away with small strings if you have to, but in a world where that isn't something that you have to put up with it would be quite frustrating. For example, say you need to preform some sort of text editing style task, and insert a few chars into the middle of a file. If one of the internal representations of the file happens to be one contiguous char* , then all you have to do is one quick memmove to make some room. With a length-prefixed representation the best case scenario is you do the memmove as before, then also update the length (no biggy really, since you probably keep that around somewhere anyway). However, if you have a 2^16 restriction and have a file larger than that you're suddenly can't use a contiguous piece of memory. This would complicate numerous things including searching, splitting, and (potentially) insertion. Not having a contiguous piece of memory also complicates the process of laying any number of data structures on top of the file data. Even further, it causes issues when you want to just memmap in a file, unless you want all your files to be perpended with the number of chars in them, which causes even more issues...
- nyellin 15y agoFrom the article's comments: "Rather than ptr+len, I would use a ptr+ptr format, that automatically ports to any size word/memory/address-space without any need for adjustment. -- Poul-Henning Kamp" "Encode the size of the length value into itself. One way to do this is to set aside the high bit, giving you seven bits of length value storage in each byte. The high bit is set for all bytes of the length value, except the last one. If there is only one byte of length value, its high bit is not set. So for example, strings of length 127 or less would have one byte of length value. Strings of length 128 to 16383 would require two bytes of length value, etc. This way the length value can have arbitrary values, yet still consuming no more space than necessary. Loading and storing the length values could be efficient with hardware support. I would suggest storing the length little-endian. -- F" "Length-prefixed strings do have many advantages, but I hate to think of the interoperability issues. These days you're probably safe with a 32-bit length, but there'd still be systems around using 16 bits, and you'd try to talk to them on a network and things would be all crazy. Not to mention endinanness issues. As well, null-terminated strings have some optimizations you can do by pointing into part of the string. e.g. you can have the strings "foo bar" and "bar" occupying the same memory space by pointing the latter at the middle of the former. It's common to do this when parsing a string, incrementing a pointer to the part you're interested in. An alternative might be to have length + pointer, pointing to a string somewhere else (which can still be null-terminated for compatibility), instead of length as a prefix. It's worth noting that the Lua language does something like this. You might also be interested to know that a buffer overflow exploit was indeed one of the earliest tricks hackers used to get into the PS3 system. A device known as PSJailbreak exploited such a vulnerability in the kernel's USB device handling to take over the system. (And getting into the console was sort of the first step of the PSN issues, since 1. Sony reacted very poorly and ticked off a lot of hackers, and 2. PSN was designed with the assumption that anything coming from a PS3 could be trusted.)" And more.