25 ms·
C Strings and my slow descent to madness
- coldpie 4y agoIt's unfortunate the author put the arrays-are-pointers thing so early in the doc, as that's a very beginner-to-C mixup and really nothing at all to do with strings. Otherwise, yep. It's pretty bad. C is a great language, but its string handling is definitely garbage. You get used to it pretty quick, and it's not hard to write a handful of sane wrappers or a simple string library for your own use, but the standard library's terrible string functions are an unending source of bugs.
- gpderetta 4y agoI don't see any mention or insinuations of arrays-are-pointers anywhere in the article. Am I missing something?
- coldpie 4y agoThis bit: But you might be asking. “Why can’t I just assign the source variable directly to the destination variable?” int main() { char source[] = "Hello, world!"; char* destination = source; strcpy(destination, source); // Copy the source string to the destination string printf("Source: %s\n", source); printf("Destination: %s\n", destination); return 0; } You can. It’s just that destination now becomes a char* and exists as a pointer to the source character array. If that isn’t what you want them this will almost certainly cause issues.
- unwind 4y agoThis is almost a cliche among many C language lawyers and/or Stack Overflow answer-rich people and I know you mean well, but: arrays are not pointers. In some contexts, the name of an array decays to a pointer to its first element. That is a better way of putting it, and it's a (much) weaker statement. Edit: if they were the same, this code: int foo[] = {1, 2, 3}; int *bar = foo; printf("%zu and %zu\n", sizeof foo, sizeof bar); Would print the same valde twice, but it doesn't. On Ideone [1] I got 12 and 8. [1]: https://ideone.com/CP7WTu https://ideone.com/CP7WTu
- int_19h 4y agoThis also makes a big difference once we start talking about pointers to arrays. int a[] = {1, 2, 3} int (*p1)[3] = &a; // ok int (*p2)[3] = &a[0]; // not ok int *p3 = &a; // not ok (It should be noted that these will compile with warnings in C due to implicit conversions via void*, but you're still risking UB if you actually use the resulting value. They are all errors in C++ because it doesn't have implicit conversion from void*.)
- bluetomcat 4y agoIn well-written C, you don't work with strings the way you do in other HLLs. For example, extracting and copying substrings is something unnecessary, unless you want to modify the parent string. Otherwise, a substring is represented by a pointer and a size_t length, and can easily be printed that way via the "%.*s" printf specifier: const char *s = "Hello World!"; const char *world = s + 6; size_t world_len = 5; printf("%.*s\n", world_len, world);
- gpderetta 4y agoOn other HLLs it is easy to have subviews on other strings. C makes is needlessly hard by requiring null termination in half the APIs.
- tom_ 4y ago* consumes an int, not a size_t: https://port70.net/~nsz/c/c11/n1570.html#7.21.6.1p5 https://port70.net/~nsz/c/c11/n1570.html#7.21.6.1p5
- TremendousJudge 4y agoI love how in every C code snippet on every comment on this thread, somebody got something wrong. I take it as a sign that it's probably best to avoid C as much as possible.
- giantrobot 4y ago! multithreaded it's not Hey, at least code
- avar 4y agoAny C compiler that isn't a trivial toy implementation will warn about that, so it's hardly a C gotcha.
- tom_ 4y agoclang and VC++ for x64 warn about this out of the box, but gcc seems to need -Wall.
- marcodiego 4y ago> If we try to print out some Japanese characters… [] The output isn’t what we expect. Yes it is. And I bet on a modern windows version it is too. The terminal has been (probably intentionally) neglected by ms for a long time, but as far as I know this has mostly been fixed on modern windows versions. EDIT: Author admits it later in the text "will be fixed in Windows 11 and Windows Server 2022" Also it says "strlen("有り難う")); [...] and the output is… The length of the string is 12 characters". But according to "man strlen": "RETURN VALUE: The strlen() function returns the number of bytes in the string pointed to by s.". It says nothing about "number of characters".
- kevingadd 4y agoIt makes sense to point it out even if it's fixed in win11, lots of people (myself included) are still on 10.
- dspillett 4y ago> The terminal has been (probably intentionally) neglected by ms for a long time, I don't think it is an intentional lack of care, just a lack of care. Internally MS devs affected by the appalling state of the console just did what the rest of us did and installed an alternative. > but as far as I know this has mostly been fixed on modern windows versions. Ish. The default console for powershell is better, but a lot of improvements you might be thinking are in there are in fact only in Windows Terminal (https://en.wikipedia.org/wiki/Windows_Terminal https://en.wikipedia.org/wiki/Windows_Terminal) which is not currently included by default.
- cmovq 4y ago> you might be thinking are in there are in fact only in Windows Terminal A lot of those changes are in ConsoleHost, so Windows 10 and 11 get those improvements (like VT100 sequences) in cmd.exe as well
- nordsieck 4y ago> Also it says "strlen("有り難う")); [...] and the output is… The length of the string is 12 characters". But according to "man strlen": "RETURN VALUE: The strlen() function returns the number of bytes in the string pointed to by s.". It says nothing about "number of characters". Yeah - when dealing with Unicode, you have to be very clear about whether you're dealing with bytes, runes or glyphs.
- _benj 4y agoWith the woes of string.h being known, why not just use an alternative like https://github.com/antirez/sds https://github.com/antirez/sds ? I’ve also been having a blast with C because writing C feels like being a god! But the biggest thing that I like about C is that the world is sort of written on it! Just yesterday I needed to parse a JSON… found a bunch of libraries that do that and just picked one that I liked the API.
- felipemnoa 4y ago>>I’ve also been having a blast with C because writing C feels like being a god Not trying to be a troll but as someone who has also written a lot of C in the past why do you feel like this?
- t43562 4y agoIt's not doing as many things behind your back as the dynamic languages and C++ do. More things are your responsibility.
- _benj 4y agoIt’s the access and control that it gives me! As when I’d pick Go because I was doing some concurrency, I can now explore a bunch of concurrency libraries, including some implementations that look a lot like Go channels. Want to watch a file for changes? I can do that all the way from taking to the kernel to picking a multi-platform library. I guess, I haven’t really found anything that I can’t do in C, and if I’m lazy, I can just embed and scripting language to handle things in a higher level. Macros are also very powerful! I’ve been writing code that writes code for me, export the thing to a .h file and import it using #include. pkg-config —list-all has become my friend and I keep discovering that the world is written in C and the access to libraries is huge! Also, idiomatic C is whatever you make it. There are a bunch of ways to skin a cat. Want a different platform? Build tool? Compiler? Debugger? Wanna write you own debugger? C is chill with that. It’s also such a simple language that without much effort you can know everything about it (I don’t care that much about anything over C99). I don’t know the whole ecosystem, or standard libraries, or data structures and algorithms and whatnot, but the language itself is quite trivial. With that said, I’m not using C in project teams. In that setting some strong conventions would likely be necessary or even better, something enforced by tools or the compiler (like Go), but yeah, I’ve been quite enjoying working with C and being kind of annoyed at other langs that I need to use for work because they keep doing all this stuff behind my back that is supposed to help me, but instead is a pain trying to debug and understand what is actually happening
- jmclnx 4y agoYes this is something to get use to. The BSDs created strlcpy(3) and wcslcpy(3) https://man.openbsd.org/strlcpy.3 https://man.openbsd.org/strlcpy.3 https://man.openbsd.org/wcslcpy.3 https://man.openbsd.org/wcslcpy.3 which to me will help with some of these issues. Too bad other Operating Systems do not have these. On Linux there is libbsd to get these, but I would like to see these to be added to the stdc. Instead the c23 standard is messing with realloc(3) which could break some old programs. I have not looked at that in detail yet, so maybe it is a non-issue :)
- deleted 4y ago[deleted]
- torstenvl 4y agoI don't know of a compiler that forces you to use the newest version of the standard, which is why I've always kind of thought "don't break old code" was treated too much like dogma. So from that perspective, a non issue. However, there is a problem that has nothing to do with old code: they increased the number of situations that constitute undefined behavior, with no public discussion and no justification. It's frankly dangerous behavior.
- alecco 4y agoUshering out strlcpy() https://lwn.net/Articles/905777/ https://lwn.net/Articles/905777/
- GabrielTFS 4y agoThese functions are in the current POSIX draft - though not published, it's quite unlikely to be removed (someone actually specifically filed an issue against POSIX to try and get it removed, basically on the basis of "it's not perfect so it should be removed", and the issue got rejected on the basis that there's no consensus for removal, and it seems unlikely this will change), and as a result, the functions are getting added to glibc: https://sourceware.org/pipermail/libc-alpha/2023-April/146967.html https://sourceware.org/pipermail/libc-alpha/2023-April/14696...
- anthomtb 4y agostrlcpy is nice due to the guaranteed NUL termination. strlcpy is not so nice due to the strange (IMO) return value of the number of characters in the source string. Which could be the number of characters copied or much, much larger than the number of characters copied. snprintf does the same thing. So using strlcpy is safe (by C's low bar) but using the return value may be highly unsafe.
- userbinator 4y agoWell-written C tends to minimise string usage in general, preferring to convert to another format as soon as possible. Allocating, copying, and passing around strings in large quantities is not a good idea for efficiency, but of course some people coming from other HLLs seem to try to do it anyway, which causes many other problems.
- bell-cot 4y agoTHIS. And programming, engineering, and life in general have so, SO many other situations where "X is not very good at doing Y". Yet (my experience) guys seem extremely resistant to the common-sense strategy of "then try to minimize how much Y you do with X".
- agumonkey 4y agoto the point I often wonder if strings should exist.. buffers -> symbols | structs.
- gaws 4y ago> preferring to convert to another format as soon as possible. Like what?
- userbinator 4y agoA structure of binary fields.
- bobajeff 4y agoI've been wondering lately why many people write c in c++ rather than just c. I think this might be the reason.
- qsort 4y agoPeople write C in C++ because they don't actually know C++ and think it's "basically C with classes and strings". There are legitimate reasons why someone would rather write C, but "I don't understand RAII" is not one of them.
- owlbite 4y agoC with basic templates is normally what I want. Occasionally other C++-isms drift in, but normally to the harm of the code quality.
- wruza 4y agoC++ doesn’t require you to commit to all of its features and/or paradigms. Using it as you see fit is valid. Just don’t advertise yourself as a C++ programmer to the job market, as it’s not what most people expect. There’s nothing wrong with “C with classes and strings” idea by itself, if that is your choice or a consciously sufficient level of competence.
- qsort 4y agoIt doesn't require you to commit to all of its features, that's certainly correct. But it does require you to commit to its principles; if you're needlessly passing naked pointers around, you're really writing C code with a C++ compiler.
- wruza 4y agoI don’t see how it requires you to commit to any principles, if you can avoid those you don’t need and still successfully compile. That’s called “suggests” or “allows”, not “requires”. Yes, some people are writing C code with classes and strings in C++. That’s why we call this mode “C with classes and strings”. I believe that you are attached to these principles (see their benefits), and that is fine. But not everybody likes full-on C++.
- infradig 4y agoI stopped when I read strcmp returns 0 if two strings are equal and 1 if they aren't.
- tom_ 4y agoA much better description of strcmp's behaviour: https://en.cppreference.com/w/c/string/byte/strcmp https://en.cppreference.com/w/c/string/byte/strcmp
- t-3 4y agoIt's actually 0 if equal, positive if greater, negative if less than. > The strcmp() and strncmp() functions return an integer greater than, equal to, or less than 0, according to whether the string s1 is greater than, equal to, or less than the string s2. The comparison is done using unsigned characters, so that ‘\200’ is greater than ‘\0’.
- jstimpfle 4y agocmp stands for compare, so the behaviour (returns <0, 0, or >0) is completely reasonable. With three possible outcomes, the function is suitable to be used for sorting.
- deleted 4y ago[deleted]
- jhatemyjob 4y agoIt's sad how often you find the truth at the very bottom of a HN thread these days
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- gavinhoward 4y agoOkay, I agree that by default, C strings are bad. But it doesn't have to stay that way. Someone else in the comments mentioned antirez's sds library for dynamic strings. This works, but you could also easily roll your own. All you need is an init function, and perhaps an assert or other check at the end of it that the string has a nul terminator. At that point, type checking will let you blindly pass those strings (or their char arrays) to any of those C functions without worry. Edit: I'll also add that I think a string library should have a difference between static strings and string builders (dynamic strings). It makes everything easier.
- magicalhippo 4y ago> by default, C strings are bad. C strings aren't bad. They can't be, because they don't exist. C doesn't have strings. And that is the issue. As you say, things get a lot better when you actually introduce strings as a concrete concept rather than a set of lose conventions.
- junon 4y agoThere is no "loose convention". A C string is a null terminated string of non-null bytes. That's the definition. Working with them in memory-unconstrained environments is unnecessarily hard.
- magicalhippo 4y agoThere is indeed just convention. The language defines string constants similar to what you say[1] (an array of characters, terminated by a null character), but in the language itself there's no way to declare that a function takes a string rather than a pointer to a character. Alternatively if you work with a fixed-sized character array, there's nothing separating it from "just" an array of characters that are not null terminated. So that strcmp expects a string rather than a pointer to a character is just convention. In languages which actually has strings as a concrete concept, like say Pascal and derivatives, you can actually differentiate between those two cases. [1]: https://www.gnu.org/software/gnu-c-manual/gnu-c-manual.html#String-Constants https://www.gnu.org/software/gnu-c-manual/gnu-c-manual.html#...
- FpUser 4y agoC strings are bad for sure. Consider those raw assembly. Instead of using it directly get some decent string library ASAP and use it exclusively.
- jstimpfle 4y agoThey are just arrays, like everything else in the language. If you don't want to manage plain arrays, better look for a different language.
- FpUser 4y agoYou are free to make good use of strings as arrays. I've written tons of code including firmware for MCU so I think I'll keep to my own practices.
- throwawaymaths 4y agothey are not "just" arrays, they are arrays with a null-terminated expectation. Most of the time. The lack of consistency and difficulty communicating the expectations is what is hair-pulling in C, and that's on top of the difficulty of communicating the difference between an array and a pointer in C.
- habibur 4y agoI don't use null terminated strings. ptr+len struct everywhere. And when I need to call an API, like fopen, I make a temporary copy of that string + the null termination, do my work and then free it. You can printf non-null terminated strings too. Check printf("%.*s", length, strptr).
- torstenvl 4y ago> You can printf non-null terminated strings too. Check printf("%.s", length, strptr).* I haven't checked yet, but I'm about 90% confident that's UB. Is printf() guaranteed not to read to the end of the string when you give it a length?
- nwellnhof 4y agoYes, it is. From the C11 spec: > Characters from the array are written up to (but not including) the terminating null character. If the precision is specified, no more than that many bytes are written. If the precision is not specified or is greater than the size of the array, the array shall contain a null character.
- torstenvl 4y agoThanks! I guess I didn't realize expressio unius est exclusio alterius applies in the C standard :D
- habibur 4y agoWon't read a single byte beyond the length you give it. And this is a standard practice in libraries for printing pre-determined lengthed string.
- t43562 4y agoA long time ago my solution was ptr+len but I allocated 1 more byte so that if a string had to be given to libc, I could terminate it at that time. No need for a copy then.
- 4y ago
- stephc_int13 4y agoIf you are using C and do some non-trivial work with strings you should either use a good library to handle strings or build your own. It is not that difficult in practice. The old C std lib is, in my opinion, outdated, obsolete and a very bad fit for complex string handling, especially on the memory management side. In my own framework, the string management module is using a dedicated memory allocator and a "high level" string API with full UTF8 support from the start. As a general rule, I think that the C std lib is the weakest part of the C language and it should only be used as a fallback.
- tyingq 4y agoFair, though there's always the crossover point where you need to interact with the OS, 3rd party libraries, protocols, etc. It's not difficult to miss a spot where your utf8 string gets mangled, truncated, etc.
- stephc_int13 4y agoYes, do not trust the OS, use the minimal API surface, and build your own toolkit, you won't have to do it often.
- maxloh 4y agoIs your framework open sourced?
- stephc_int13 4y agoNot 100% decided yet, but it is very probable that I will open source it.
- attractivechaos 4y ago> especially on the memory management side. Libc string functions don't manage memory. They can be used no matter where your strings are stored. It is more of a choice between generality vs convenience in common cases.
- nathell 4y ago> Our last function is strcmp. It looks at two strings and determines whether they are equal to each other or not. If they are it returns 0. If they aren’t it returns 1. No it doesn’t. RETURN VALUES The strcmp() and strncmp() functions return an integer greater than, equal to, or less than 0, according as the string s1 is greater than, equal to, or less than the string s2. The comparison is done using unsigned characters, so that ‘\200’ is greater than ‘\0’.
- Decabytes 4y agoI've added a footnote to my incorrect explanation and credited you. I'm still a C noob so thank you for pointing this out!
- andoma 4y agoReminds me of the incorrect cast of memcmp() return value that resulted in this bug: https://bugs.mysql.com/bug.php?id=64884 https://bugs.mysql.com/bug.php?id=64884
- cassepipe 4y agoI don't have MacOs to prove it but I believe `strcmp` on MacOs returns either 0, 1 or -1
- skywal_l 4y agohttps://developer.apple.com/library/archive/documentation/System/Conceptual/ManPages_iPhoneOS/man3/strcmp.3.html https://developer.apple.com/library/archive/documentation/Sy...
- cpeterso 4y agostrcmp's return value is loosely defined because some implementations return the difference between the characters in the string to avoid some conditional checks or jumps. Something like: int d = a[i] - b[i]; if (d == 0) return d;
- 4y ago
- mahoho 4y agoJust a pedantic comment, but 有り難う is arigatou or roughly "thanks", not "hello". Hello would usually be こんにちは or, more confusingly, 今日は
- teddyh 4y agoもしもし
- commandlinefan 4y agoSort of unfortunate, because there's really no good translation for "hello" into Japanese - you'd say こんにちわ in the morning, in the afternoon こんばんわ and もしもし when answering the phone...
- scrame 4y agoyou'd say ohayogozaimasu in the morning, konnichiwa in the afternoon and konbanwa in the evening. I don't know how to type japanese on my phone, but the first is literally "early" surrounded by honorifics. The characters for konnichiwa means "this/now", "day" and the "wa" at the end is an article making the previous phrase the subject of the sentence. Same with konbanwa, but for evening instead of day. no idea on the etymology of moshimoshi for answering the phone, though.
- syockit 4y agoAccording to an article I read, moshimoshi came from the telephone operators saying 申し申し (moushi moushi) to signify "I'm going to start speaking now". On a tangential note, 申し used to be the phrase when calling out to someone to ask something (similar to English "Excuse me").
- gpderetta 4y agoThe worst part of C strings is that they tend to show up in APIs (especially system calls). This make interoperability with other languages harder than it should m
- ziml77 4y agoThis is why I hate them too. You can use a custom length + pointer type for representing strings in your own code, but interfacing with other libraries and the OS almost always requires having a null-terminated string. It forces you to make copies just to tack on the null terminator.
- benmmurphy 4y ago`strlcpy` is the function you probably want. but again it is not standard. https://lwn.net/Articles/507319/ https://lwn.net/Articles/507319/ I think the reason people don't want to standardise this kind of function is it often gives wrong behaviour. for example if you are trying to copy a string into a fixed buffer and its too long then often it is an error or potentially even a security bug to truncate it. so these functions generally do the 'wrong' thing even though they are 'safer'. if you are dealing with static buffers then I think you should be explicitly checking the source fits in the target and then handling the error case. you could even have a function like `strlcpy` that does `strlen` then checks if it fits, then does the copy or return an error code. alternatively, if the string should always fit and you don't want to handle the error case then the safe thing to do is check at runtime that it fits then abort the program if it doesn't fit.
- Night_Thastus 4y agostrlcpy is not needed, strcpy_s (not strncpy_s) is safe and is part of the C11 standard.| In fact, strlcpy is worse: * strlcpy truncates the source string to fit in the destination (which is a security risk) * strlcpy does not perform all the runtime checks that strcpy_s does * strlcpy does not make failures obvious by setting the destination to a null string or calling a handler if the call fails.
- zabzonk 4y ago> strcpy_s is part of the C11 standard an optional part, which makes it pretty worthless, if it were not so already.
- kelnos 4y agoOn systems that aren't memory constrained, we just shouldn't be using static buffers at all. Just always use something like asprintf() and free() the result when you're done. No, it's not in the C or POSIX standards, and that's a shame, but it's at least available on Linux and the BSDs. I end up working on a lot of code that uses Glib, so I tend to use g_strdup_printf() a lot, which works the same as asprintf(). Ultimately the cost of allocations is usually not a big deal, and you gain a lot of safety. Sure, you then have to remember to free(), but I'll take a memory leak over a segfault (and its possible security consequences) any day. And if allocation cost is a problem, you can always go back and optimize with static buffers later. That shouldn't be the default that people reach for, though.
- photochemsyn 4y agoFor initial string input, i.e. from a network/file/terminal stream, using fgetc and/or fgets plus code to verify and sanitize makes the most sense IMO. This does mean you have to write a lot of C code for what would be simple tasks in other languages, e.g. a correct file open, read-to-dynamically-allocated-memory, and file close with good error checking is a full page (at least) of dense code in C and just two lines in Python. If you've done a good job sanitizing and verifying all the input to your program, only then does it becomes relatively safe to use the standard string functions, with caveats for multithreading. Asking ChatGPT to compare and contrast fgetc and fgets is a good place to start, and then ask how to use fgets to handle errors during stream I/O, and what can go wrong with multithreading etc. Then take a look at the sqlite source code for in-house C-string handling, here's the take-away comment: "Because there is no consistency, we will define our own." https://github.com/sqlite/sqlite/blob/master/src/util.c https://github.com/sqlite/sqlite/blob/master/src/util.c
- Night_Thastus 4y agoIt's important to mention that strncpy (and also strncpy_s) are really not a strcpy replacement, it's not intended for the same usages. The name is a total misnomer. Do not use strncpy that way! In any case, strcpy_s (which is a good replacement for strcpy) is part of the C11 standard. I'm confused how that isn't considered portable.
- zabzonk 4y agoit's an _optional_ part of the standard, and so can't be relied on. also, the idea behind it is pretty poor.
- Night_Thastus 4y agoDon't GCC, Clang and MSVC provide it? It may be optional in practice but if the major compilers support it, it's not really an issue. The idea may not be perfect, but for C which is intended to be low overhead, strcpy_s is about as good as it gets. If you want something more user friendly, that is what C++ is for with std::string, or library implementations like Boost or QT string.
- zabzonk 4y agoMSVC supports it, as it is an MS invention, designed by some intern, i guess. other compilers may or may not, by switches/#defines. it is worthless in any case.
- Night_Thastus 4y ago>as it is an MS invention Source, and why does it matter who made it? >designed by some intern You don't know that, nor is that how standards work. >other compilers may or may not So you've said basically nothing. >it is worthless in any case. It offers a low-overhead, safer alternative to strcpy. Is it perfect? No. But it's one of the better C options for those limited to the standard library.
- 4y ago
- js2 4y agoPrograms which have to deal with C strings beyond the bare minimum that libc provides will generally have a set of routines for making it more ergonomic. e.g.: https://github.com/git/git/blob/master/strbuf.h https://github.com/git/git/blob/master/strbuf.h
- flohofwoe 4y agoThis is from a C fan: If you are going to do any string heavy work, please use anything else than C (Python is pretty nice for this sort of stuff for instance). And if you need to use C anyway, then please use anything else than the string functions from the standard library. The C stdlib is (mostly) a leftover from the K&R era when opinions about what makes a good API were very different from today, and C was a much 'harsher' language. C is pretty nice for a lot of things, but working with strings definitely isn't one of them.
- cozzyd 4y agoI would say, unless there's a performance reason not to, always use asprintf for every string operation.
- axilmar 4y agoFor string heavy workload, C is ideal, provided that you don't use C string functions. You can always allocate a very large buffer and do your string operations there, using memncpy and the assorted functions which can be inlined in many architectures and be really fast. Then you can dispose of the buffer really quickly with one call or reuse it for later operations by simply setting a few pointers to initial status...
- flohofwoe 4y agoIf you have large data sets and need maximum performance I agree, but a lot of day to day string processing works on small data sets and isn't very performance-sensitive.
- BananaaRepublik 4y agoAs a newcomer to C, why is it that the C standard library doesn't get updated? Newer languages seem to place a lot of emphasis on getting their standard libraries as useful as possible. It's odd to be told not to use the standard library functions but to write my own instead. I'm really doubting I can just sit down and hammer out string functions superior to string.h.
- pixelbeat__ 4y agoString handling in C has many gotchas indeed. Here are some of my notes on the subtleties: https://www.pixelbeat.org/programming/gcc/string_buffers.html https://www.pixelbeat.org/programming/gcc/string_buffers.htm...
- teddyh 4y ago> strcmp takes two strings and returns 0 when they are true. ITYM “equal”, not “true”.
- draw_down 4y ago[dead]
- Tarragon 4y ago> So how can we handle this case safely? There are a few ways I can think of. strdup: > The strdup() function returns a pointer to a new string which is a duplicate of the string s. Memory for the new string is obtained with malloc(3), and can be freed with free(3).
- kevin_thibedeau 4y agowchar_t is a massive landmine that should never be used since its size varies by platform. The locale of the compiler has to match the end user for L prefixed strings to work correctly. Likewise char16_t and char32_t are just swimming against the easy path at this point. You're much better off sticking to UTF-8 and using the C11 u8 prefix on literals so you can use the regular string API and never have to worry about locale settings.
- Decabytes 4y agoThis is great advice! I wasn't aware of this and I will keep that in mind. When I first came across Unicode literals I was unsure when exactly you would use them over wchar_t
- zh3 4y agoI once got called in to fix an SS7 stack suffering from poor performance. Pretty well written, and not obvious at first sight why it was going slow. Most of it was low-level bit fiddling, and some small strncpy's() - generally about 8 chars or so. Didn't take that long to profile (well, printf's as no profiling available) and figure out it was the strncpy's causing the problem, but why? Well, there was a handy 8 megabyte buffer used for working memory that the strings were being copied into that for modification. From the strncpy() man page:- >If the length of src is less than n, strncpy() pads the remainder of dest with null bytes. Ah. So every little strncpy was essentially copying the string then zeroing out 7,999,992 bytes. And there were lots of little strncpy's...
- PaulDavisThe1st 4y agoThe appearance of strncpy() in any source code is an immediate panic attack for me. It should never be used, and if it is used, it should be removed. Similar rule for sprintf(), all instances of which should be replaced by snprintf().
- bluGill 4y agoUnfortunately there often isn't a better replacement in your standard library (embedded systems are weird). I ended up using strncpy followed by automatically setting the last byte of the string to null.
- PaulDavisThe1st 4y agostrncpy() in particular is so bad that you're better off (for a rare exception to the rule) just writing your own, that does what most people think strncpy() does (or should do) rather than what it actually does.
- saagarjha 4y agoWrap memccpy.
- saagarjha 4y agosnprintf has very similar performance pitfalls.
- russellbeattie 4y agoLiterally 25 years ago I was a beginner programmer and tried writing a .dll for Microsoft's Internet Information Server, which was relatively new at the time. (I hadn't so much as seen a Unix-based OS at the time, let alone understood CGI). C strings were mind boggling and frustrated me so much I simply gave up. Happily around the same time, MS introduced Active Server Pages and I was able to use that and never messed with C again. It's amazing the same issues still exist decades later.
- chihuahua 4y agoThat is the most mind-boggling part of this saga to me. People have been using C since the 1970s. It's now 2023, and there still isn't an obvious solution to this other than suggestions that every team should write their own string library from scratch. And apparently it all started with some genius deciding that using a single 0-byte at the end is so deliciously efficient and therefore obviously the way to go. We can't waste 4 bytes for the string length, that's out of the question. I think only the Pascal solution of having a single byte for the string length is worse.
- kloch 4y ago> At first this looks great, but there is a problem. What happens when the source string minus the null terminator is as long as the size of the destination string? The answer is that the destination gets filled with all the characters of the source string with no room left for the null terminator. The 'n' in strncpy is mainly there to help you avoid overrunning the destination, it does not guarantee whatever makes in there is null-terminated. This is why you should always explicitly set the last byte to zero after using strncpy (and never ever use strcpy). char dest[16]; strncpy(dest, src, 15); dest[15]=0;
- djha-skin 4y agoRelated and good read about strcpy in the kernel: https://lwn.net/Articles/905777 https://lwn.net/Articles/905777
- Dwedit 4y ago> "But for real if anyone knows how to get this to work on Windows 10 let me know!" Since the May 2019 update, Windows 10 has supported declaring the code page in a manifest file. In Visual Studio, you must add "/utf-8" to the compiler command line, this makes it parse the source code as a UTF-8 file, and makes it output UTF-8 string literals. To make console output work, call the Win32 function "SetConsoleOutputCP(65001);" To get support for opening files with names that aren't in your system codepage: * Create a manifest file as shown in https://learn.microsoft.com/en-us/windows/apps/design https://learn.microsoft.com/en-us/windows/apps/design /globalizing/use-utf8-code-page * Add this as an "Additional Manifest File" in Visual Studio project settings for the manifest tool Additionally, there is an undocumented NTDLL function "RtlInitNlsTables" that sets the code page for the process. It is difficult to use without a lot of example code, but some app locale type tools (used to change locale for a process) make use of this function.
- jeffrallen 4y agoProgramming in Go made me a better C programmer, because now I no longer use C strings, only a buffer/length/capacity struct.
- bruce343434 4y agoMeh. The w_char stuff is barely C's fault. You use wide (constant width) characters then set the terminal encoding to utf8 (variable length encoding). What did you expect? It's a windows issue. I can copy paste all sorts of utf8 in "normal" string literals, printf and puts them, and it just works in my terminal. RE counting characters: this is a whole can of worms. Do you want to count grapheme clusters? Code points? Anything other than just the amount of bytes? Use a unicode library. The latter part of this article is a bit like those articles that make fun of javascript for having floating point numbers behave like, gasp, floating point numbers.
- analog31 4y agoAs a cellist, I was about to sympathize when I read the title.
- simonblack 4y ago"We're not in Kansas any more, Toto" Or to paraphrase that "We're not in Python any more, and C is not Python". You know what sends me insane? Indentation and lack of fixed types in Python. But I don't have problems with C strings. Because I have grown to love and know C's string foibles just like the author will certainly not be driven insane by 'Python's shortcomings according to me'. The world is full of people who complain that something or other is different from what they know, so that 'other' is wrong. That's just being isolationist. Everything has its own advantages, its own disadvantages. Let's accept that and move on, instead of making mountains out of mole-hills.
- Sohcahtoa82 4y ago> Indentation and lack of fixed types in Python. Whenever I see someone complain about Python's indentation, my brain internally translates it to "I poorly format my code." If you code is properly formatted, then Python's indentation is never a problem. I praise Python's indentation-as-syntax because it prevents issues like a dangling else or a forgotten brace while also making proper formatting a requirement for your program to run.
- smackeyacky 4y agoLet's be clear about this: python indentation-as-syntax makes one particular style of formatting a requirement. Not everybody likes it. You might say I poorly format my code in other languages, but most of the time for readability you are somewhat stuck with the formatting of that code base which turns out to be a non-problem thanks to modern editors (sarcasm) like vi that can do brace matching and are older than python. Python's rigidity on formatting should solve that problem, but it really doesn't and over time it relaxed the rules somewhat which made it better.
- _a_a_a_ 4y ago> then Python's indentation is never a problem I'm sure you have some solid 3rd party research to cite to back up that absolutist claim. Or have some explanation that a it with my workmate's fault and not a git merge (which mixed up indenting between the two files) that introduced a bug that took hours to track down. Or why people like me loved the idea of meaningful indentation in python, then grew to not love it any more after experience. Or the bitch that it can be when you have to generate code and have to do more than just slap braces around it to delimit blocks. All in all, another one of those pure-opinion HN posts that are becoming too common.
- tragomaskhalos 4y agoK&R contains this beautiful koan-like string copy code: while (*t++ = *s++) ; Honestly the elegance of this thing was one of the hooks that made me fall in love with C. But this was from a now-forgotten age of innocence, as there are so many "nopes" around this line-and-a-half that one would, rightly, be tarred and feathered for ever putting it in a program today.
- kovac 4y agoCould you explain why this line should be discouraged? I'm a beginner in C, so I really don't know. That's why I'm asking.
- fatcow 4y agowhile (t++ = s++) ; You're assigning a char to another, relying on the return value being 0 to detect end of string. You're performing the copy while also increasing the pointers with ++ in the same expression. You're using the cryptic ; empty statement to signify nop, thereby confusing newbies. Etc
- commandersaki 4y agoEquivalent of a strcpy() there's no bounds checking.
- tragomaskhalos 4y agoAnother reason, in addition to other replies and separate from safety concerns, is that strcpy, memcpy etc are nowadays usually implemented via more efficient compiler intrinsics rather than an explicit loop.
- ranger_danger 4y ago> By default, Windows PowerShell .lnk shortcut is hardcoded to use the "Consolas" font Surely this is not the case for Japanese versions of Windows (or users with Japanese set as their display language?)
- lifthrasiir 4y agoYes, you can actual look at the full list from `HLKM\Software\Microsoft\Windows NT\CurrentVersion\Console\TrueTypeFont` (taken from my copy of Windows 10, bracketed comments mine): 0 Lucida Console 00 Consolas 932 *MS ゴシック [MS Gothic for Japanese] 936 *新宋体 [Simsun for simplified Chinese] 949 *굴림체 [Gulimche for Korean] 950 *細明體 [Windows MingLiU for traditional Chinese] Note that there are actually two global defaults which only are differentiated by leading zeros. This is intentional and can be used to enable additional fonts; it is a common tweak for Korean (and probably other CJK) users to add a preferred font with a name 0949 or 00949 etc.
- ranger_danger 3y agoJapanese versions use CP932 though, which according to your list would use a font that's not Consolas (MS Gothic in this case).
- lifthrasiir 3y agoI should have said "no" in place of "yes" (being a non-native speaker, I completely overlooked "not" in your original reply), otherwise I believe I didn't contradict you.
- TheRealPomax 4y agoHow to make C safe: by putting it back in the box and putting the box back on the shelf and then closing the door to the garage. Then using a safe-by-design language.
- kens 4y agoIt's strange that computer programmers think of themselves as being on the cutting edge of technology, but then we use a language that is over 50 years old. Of course there are going to be lots of problems with C strings since they were designed for a totally different world (no Unicode, no security issues, memory was precious, etc). The hardware is a million times as powerful but the software environment improves at a glacial pace.
- nneonneo 4y agoPop quiz, which of these is safe, given "char buf[80]" and arbitrary user input in argv[1]? gets(buf); scanf("%s", buf); strcpy(buf, argv[1]); scanf("%80s", buf); strncpy(buf, argv[1], 80); snprintf(buf, 80, argv[1]); ---- The delightful answer is none of them. The first three have no bounds checking at all, meaning that they will happily overflow the buffer to an arbitrary extent (gets, at least, will usually trigger a warning on modern compilers). The next two have off-by-one errors: scanf will write a NUL byte out of bounds (and that's exploitable! https://googleprojectzero.blogspot.com/2014/08/the-poisoned-nul-byte-2014-edition.html https://googleprojectzero.blogspot.com/2014/08/the-poisoned-...) while strncpy will fail to NUL-terminate the string. The last one uses the right buffer length, but treats user input as a format string and can leak memory contents or produce arbitrary memory corruption with the %n format specifier. C string handling practically invites off-by-one errors and horrible security practices out-of-the-box.
- cdelsolar 4y agoSo then what do you do?
- myrmidon 4y agoAre you sure that strncpy does an out-of-bound write here? I believe it doesnt, but would give you an unterminated string in buf which is also... less than ideal (if the input is 80 non-null characters or longer).
- nneonneo 4y agoUgh, I had that in my comment before I "refactored" it. Fixed now, thanks for pointing it out.
- Lt_Riza_Hawkeye 4y agoYeah, and I wouldn't say it's definitely unsafe. You can memchr a '\0' out of it (or not) to determine if a null terminator got in there or not.
- 4y ago
- dsvr 4y agoThere's a framework for C now at https://vely.dev https://vely.dev which may help with C strings safety and memory management, among other things.
- mantratome 4y agoI found Vely on dev.to and did a test project with it. Pretty neat and solid.
- pipeline_peak 4y agoHandling strings in C was enough for me to choose C++…
- hollowturtle 4y agoAny good resource around with all the common C pitfalls and relative solutions?
- tmsln 4y agoBeej's guide to C programming is very helpful: https://beej.us/guide/bgc/html/split/unicode-wide-characters-and-all-that.html#unicode-wide-characters-and-all-that https://beej.us/guide/bgc/html/split/unicode-wide-characters...