10 ms·
Are the any good resources that explain the concept of strings in C, particularly why they’re considered to be so difficult to manage? I’m interested in the lan
by the-printer 4y ago
Are the any good resources that explain the concept of strings in C, particularly why they’re considered to be so difficult to manage? I’m interested in the language, and that along with its safety concerns seem to be the two most frequent complaints against it that I read about online.
- e-dant 4y agoC strings are pointers to memory. There are semantics and assumptions encouraging null-character delimited strings, but not every API follows those rules (just got done working with a Windows API that doesn’t). Often, you have to both null-delimit your string and store its length somewhere. That’s the dangerous part. Messing either of those up, or passing your string to an API that messes that up, is not safe. C strings are pointers to memory, either the stack or the heap, and follow exactly the same rules as everything else in that chaotic space: Not many.
- the-printer 4y agoThank you for this. C programming sounds almost like some sort of combat sport. Riveting.
- lelanthran 4y ago> Thank you for this. C programming sounds almost like some sort of combat sport. Riveting. I've done it for decades; it isn't really as bad as hype-attracting headlines would have you believe. Munitions control, aircraft management systems, industrial automation systems, and many more life-critical systems were programmed in C for decades with comparatively little danger from the language intrinsics leading to death. It's easy to look at the stats and say "there's a few dozen CVEs annually due to C footguns", but that's a few dozen out of hundreds of millions of deployed systems that are written in C. In practice, very few lines of C code bypass the type system, so you get much fewer bugs than an equivalent system in the more usual dynamic programming languages (Python, Javascript, etc).
- thesnide 4y agoWondering if the big influx of C derived CVE are old or new code. If it is new code, I'm also wondering about the brain damage that those safe languages causes. Yes, it is better to have memory safe languages. But it encourages sloppiness as "nothing can happen". Then those folks aren't fit to write anything else. Which closes the feedback loop on inefficient but safe languages. Which becomes the same thing in airplanes. Pilots don't really know how to fly without instruments anymore.
- wadd1e 4y ago>Which becomes the same thing in airplanes. Pilots don't really know how to fly without instruments anymore. Well that's just a blatantly wrong generalisation you made there, curious as to where you got that from. Consider looking up how pilot training is done before making such assumptions. Even though modern airplanes make heavy use of technology, there are emergency scenarios where lots of instruments may not work, and pilots receive more than enough training to fly an airplane in that scenario just to give one example among tons of others. edit: grammar
- thesnide 4y agoIt seems there's a difference between theory and practice. https://www.businessinsider.com/too-many-pilots-cant-handle-an-emergency-2014-12 https://www.businessinsider.com/too-many-pilots-cant-handle-... Now, I'm not an expert into pilots statistics, so my example might be off, but I do see a worrisome pattern in my daily work (software engineering). Blind reliance on those "frameworks". Which isn't bad in itself, but no-one really knows how they work anymore. They just assume. And that leads to lots of cargo cults. Which ranges from inefficient to outright dangerous.
- pjmlp 4y agoIt is more like a combat sport, doing martial art moves, while trying to juggle knives between moves.
- marssaxman 4y agoMore like fire-performance: it looks dangerous, and it does require some finesse, but it's really satisfying when you get in the flow, and burns are both less frequent and less serious than you might imagine as an onlooker.
- jstimpfle 4y ago"Strings" are quite an abstract concept. They are a linear sequence of characters. But there are a number of ways to represent them - the simplest of which is a contiguous memory allocation, but depending on the use case you'd need more complex schemes. There are also different ways to do the necessary memory management (e.g. allocate statically at compile time vs dynamically at run time). One of the most complex representations is probably the string rope datastructure - a balanced tree of string chunks, supporting efficient insertion and removal anywhere in the string. Specific to C, as well as lots of low-level APIs, is only that strings are often expected to be contiguously laid out in memory and terminated with a NUL (0) byte. So you need to make sure that you always terminate with a NUL after writing to string storage. Other than that, strings aren't any harder than other aspects of programming with manually managed memory. Maybe motivated from higher-level dynamic or managed languages, is the popular idea that strings should always be allocated dynamically (like std::string for example), and support operations like string-append with automatic reallocation if the currently allocated memory isn't enough to store the new string. In practice, that's not true at all - unless you are in a domain where lots of small intermediate strings are generated. This is pretty inefficient anyway and there is likely no point to use C in this case. By far most strings in most domains are either completely static (use string literals), or are created once in a sequence of append operations and then never changed again. I get by, doing many different things from GUI apps to networking to parsers and interpreters, without any sophisticated string type. All I do is define some printf-like APIs to do logging, for example. Those typically just use a fixed size buffer for the formatting, and then flush that buffer to e.g. stderr. or flush it to a dynamically allocated memory buffer, but there almost never is a need to reallocate that string later.
- Quentak 4y agoI have written a short article explaining why null terminated strings as they exist in C cannot represent proper ASCII and UTF-8 because of the null terminator. It's not a full explanation of how strings work but it might be helpful for you. https://kttnr.net/blog/null-terminated-strings-are-incorrect/ https://kttnr.net/blog/null-terminated-strings-are-incorrect...
- lifthrasiir 4y agoSQLite does store a null character in strings, it has lots of documented [1] issues in the API level though. [1] https://www.sqlite.org/nulinstr.html https://www.sqlite.org/nulinstr.html
- Quentak 4y agoThanks for this link. How do you get the null byte into the string? Is it through casting blob to string? The way I have encountered this is when using the C API in which string arguments for prepared statements are passed as char pointers. If those contain the null byte then the string is cut off. Allowing null characters and then mishandling them is worse than not allowing them.
- Someone 4y agoThe problem is that C doesn’t have strings; it has functions that treat sequences of non-zero bytes followed by a zero bytes as if they are strings. So, you can’t ask it to create a string that contains the result of appending a string to another one. If you want to append two ‘strings’, you have to create a buffer large enough to hold the result, and then copy in the two sequences of bytes. And even for doing that, the library functions aren’t optimal. The basic “append this string’s data to that string, assuming there’s enough space to do so” function is strcat. It walks the first string to find the zero byte, but to “create a buffer large enough to hold the result” you already must do that. See for example https://stackoverflow.com/questions/21880730/c-what-is-the-best-and-fastest-way-to-concatenate-strings https://stackoverflow.com/questions/21880730/c-what-is-the-b...
- jstimpfle 4y agoYou can use snprintf to easily achieve any concatentation you'd like. len = snprintf(buffer, buffersize, "%s%s%d", string_1, string_2, int_1); if (len + 1 /*NUL*/ > buffersize) { // not enough space } You can also use this to dynamically allocate any formatted string len = snprintf(NULL, 0, ....); buffersize = len + 1; /* NUL */ buffer = allocate(buffersize); snprintf(buffer, buffersize, ... /*same args as before*/);
- tom_ 4y agoI like the printf family too. Any time you're doing a bunch of strcat or whatever it's almost always massively easier to use a format string to get the same result. Very easy to get the desired width/precision/alignment, and if you need numbers, printf has your back. It even does the bounds checking for you! (And how often do you get that in C.) It won't be as fast, but it's almost always not a problem, and the nice thing about C and C++ is that the char-by-char route is still available when it is. I like to use asprintf, when available: https://man7.org/linux/man-pages/man3/asprintf.3.html https://man7.org/linux/man-pages/man3/asprintf.3.html - and when not available, I add it, along the lines of the snippet you present. Here's something I've found a useful upgrade to asprintf, as it frees the passed-in buffer after expanding the format string. You can just pass the same char ** repeatedly and it'll update the char * appropriately each time. int xasprintf(char**p,const char *fmt,..) { int n=0; char *p2=nullptr; if(fmt) { va_list v; va_start(v); n=asprintf(&p2,fmt,v); va_end(v); } if(n>=0) { free(*p); *p=p2; } return n; }
- WalterBright 4y ago> why they’re considered to be so difficult to manage? Back in the 90s, I was very experienced with C strings and managing them. Then I chanced to look at BASIC again, and realized that strings in BASIC were so simple and intuitive. Why couldn't C be like that? When I started on the design of D, I decided that it had to make strings as easy to do as BASIC did. And D does. The trouble with C strings is the 0 termination of them. This means: 1. to get the length of the string, you have to scan it. This is expensive. 2. when manipulating strings, a common error is to get off by one in the storage because of the 0 termination 3. you cannot get a subset of the string without making a copy. Not only is a copy expensive, but then you have to keep track of the memory for it 4. there's no way to check for buffer overflows D's design, which uses a phat pointer (length, ptr) for strings, solves these problems.
- tored 4y agoA year ago I picked up the BASIC dialect PureBasic. Pleasant surprise actually, the syntax of the PureBasic dialect is a bit archaic, but if you accept that it is much easier and faster to get anything done compared to C (and C++). Personally I find low level topics easier to grok in PureBasic than in C even though they mirror the same concepts. PureBasic has Unicode strings built in. It is a bit shame that BASIC has such a bad reputation, there are many BASIC dialects that does the job well still today. https://www.purebasic.com/documentation/reference/ug_string.html https://www.purebasic.com/documentation/reference/ug_string....