34 ms·
Efficient string copying and concatenation in C
- GuB-42 7y agoThis is one of the reason I tend to avoid str* functions in the first place, except for one strlen() per string. The way I copy and concatenate strings typically looks like: int len1 = strlen(str1); int len2 = strlen(str2); char *buf = malloc(len1 + len2 + 1); if (buf) { memcpy(buf, str1, len1); memcpy(buf + len1, str2, len2); buf[len1 + len2] = 0; } Of course, not memcpy()ing data around is even better if I can avoid it.
- js2 7y agos/int/size_t/ and check for overflow.
- GuB-42 7y agoTotally right about size_t, my bad, hopefully, the compiler will raise a warning. As for integer overflow, I don't actually know how to handle it properly. In normal conditions, it is unlikely to be a problem. If the two strings can fit in memory, the sum of their size should fit in a size_t, but I agree that making such assumptions can be a bad idea. Maybe the best way is to limit the size of the input strings to a reasonable value. That would prevent many out of memory situations too and potential DoS too.
- js2 7y ago> As for integer overflow, I don't actually know how to handle it properly. Something like: if (str1 && str2) { size_t len1 = strlen(str1); size_t len2 = strlen(str2); size_t buf_len = len1 + len2 + 1; if (len1 < buf_len && len2 < buf_len) { char *buf = malloc(buf_len); if (buf) { memcpy(buf, str1, len1); memcpy(buf + len1, str2, buf_len - len1 - 1); buf[buf_len - 1] = '\0'; } } } (I probably made a mistake above.) As you suggest, you'll probably run out of memory before you'll overflow, so in reality, you want to check len1 and len2 are some sane value, but of course, library functions don't usually have that luxury. Take a hint from git: #define unsigned_add_overflows(a, b) \ ((b) > maximum_unsigned_value_of_type(a) - (a)) if (unsigned_add_overflows(extra, 1) || unsigned_add_overflows(sb->len, extra + 1)) die("you want to use way too much memory"); https://github.com/git/git/blob/6d5b26420848ec3bc7eae46a7ffa54f20276249d/git-compat-util.h https://github.com/git/git/blob/6d5b26420848ec3bc7eae46a7ffa... https://github.com/git/git/blob/9d418600f4d10dcbbfb0b5fdbc71d509e03ba719/strbuf.c#L90 https://github.com/git/git/blob/9d418600f4d10dcbbfb0b5fdbc71... > making such assumptions can be a bad idea It's always a bad idea, especially in an unsafe language. Never trust user input.
- dwheeler 7y agomemccpy definitely has some advantages, so it definitely should be in the ISO standard. But memccpy has its own problems. In particular, when concatenating you have to constantly recalculate the "space remaining"; that is just asking for an off-by-one error that leads to a buffer overflow, and makes it more complicated to use. The discussion here doesn't detect attempted overflows, and that's a mistake; you often need to not just prevent an overflow, but you also need to detect an attempted overflow and do something different. You also have to pass \0, which makes the function call more complex (and perhaps under-optimized) since \0 would in nearly all cases be the parameter passed. So I'm glad this is being added, but it's at most a small step to improving simple string copying and concatenation in C.
- komali2 7y agoObligatory sidenote on website itself rather than content: please don't set such a low contrast between font color, choice, and background color. I had to copy/paste this article to read it.
- segfaultbuserr 7y agoUse the developer tool to inspect CSS. 1. Uncheck article .entry-content { color: #646464; } And a new CSS color rules appears. 2. Uncheck body { color: #333; } Done! I used to think it's ridiculous to manipulate a webpage like this manually, but now I believe: if it's helpful for your for an one-time browsing on a broken webpage, why not?
- carapace 7y agoyah, thin light grey sans-serif body font... "Why do you hate my eyes?"
- lota-putty 7y agoOT: Is there any API out there that implements say a length* prefix in all C strings? * 2 or 4 octets size
- ninkendo 7y agoWhat you're talking about typically goes by the name of "pascal strings", and while they're possible to do, C's string literals are not compatible with them, so nobody does it.
- tyingq 7y agosds (simple dynamic strings) is a decent compromise: https://github.com/antirez/sds https://github.com/antirez/sds (used in redis)
- childintime 7y agoIt is certainly possible to declare Pascal string literals with not much hassle: https://stackoverflow.com/questions/7648947/declaring-pascal-style-strings-in-c https://stackoverflow.com/questions/7648947/declaring-pascal... One of the answers states that GCC and Clang do have support for Pascal strings. Probably these strings do not work as (well) in #defines, i.e. they don't concat like regular literals.
- avar 7y agoYes, those are a dime a dozen. E.g. GString[1] is one example. Rather than strict "Pascal strings" like the sibling comments rightly points out would be stupid, they're a light struct with a char/size_t len (and usually also size_t alloc_len). This allows for skipping strlen() on them before e.g. copying, while using them as a C string by getting directly at the char. 1. https://developer.gnome.org/glib/stable/glib-Strings.html#GString https://developer.gnome.org/glib/stable/glib-Strings.html#GS...
- zokier 7y agosds comes to mind, but there are like gazillion different implementations of the same concept. https://github.com/antirez/sds https://github.com/antirez/sds
- wrs 7y agoThe Microsoft approach to this was to make a set of replacement functions (strsafe.h [1]) that are very explicit and not at all “clever”, as the strsplcasdfcpy functions seem to want to be. They return an error code so it’s obvious when the operation did what was expected, or ran out of space. [1] https://docs.microsoft.com/en-us/windows/win32/menurc/strsafe-ovw https://docs.microsoft.com/en-us/windows/win32/menurc/strsaf...
- dwheeler 7y agoBut most people want to use standard, portable calls. The C standard tried to add such functions in "Annex K", but unfortunately annex K hasn't received much of a pickup (for various reasons). So in many places the problem continues.
- codehero 7y agoEfficiency looks past current deficiency. We have the empty string: "\0" We have the null string: NULL There is no concept of an INVALID string, as float has NAN. This would be the result of trying to copy a string to a buffer that is too small. Or sprintf() into a small buffer. Or a raw string parsed as UTF-8 and is invalid. Correctness over efficiency.
- akersten 7y agoI wouldn't prefer one more special case to test against (empty string / null string / 'invalid' string). Why can't those operations just return error codes instead? How about memcpy if you try to memcpy into a buffer that's too small - it writes an 'invalid buffer' type instead?
- codehero 7y agoPropagating NAN is an elegant method in floating point and makes sense for well defined string encodings like UTF-8. memcpy and company are strictly for raw unencoded buffers.
- Someone 7y ago”This would be the result of trying to copy a string to a buffer that is too small.” C doesn’t have the notion of “size of buffer” (yes, arrays have a size that can be queried by sizeof, but only at compile time). You would have to fix that, first.
- charliesome 7y agoI'd argue that an invalid string concept would be neither correct nor efficient. Why should all code that deals with strings carry the burden of fallibility of a subset of string functions? You've mentioned NaN propagation in another comment and I think that's a perfect example of the problem with this approach. Sorting a vector of arbitrary floats is a notoriously thorny problem because any float could be NaN, and as NaN is incomparable to any other float, there is no total ordering of floats. There is no general solution to this problem that doesn't involve making assumptions that could be faulty for some applications.
- antirez 7y agoI can't see how incrementally improving the approach of the C standard library is the right design decision here. They need to just include a standard, well designed dynamic strings library to C, that stores the length alongside with the (binary safe) string itself, and that's all. Every serious C program uses it, null terminated strings are a joke. However because of this huge background with null terminated strings in C, such library should make sure to always automatically terminate strings with a null term, so that people can trivially print them, call strlen() against them and so forth when needed and when there is no binary data inside.
- mehrdadn 7y agoSounds an awful lot like C should just become C++ already...
- antirez 7y agoNo need, C dynamic string libraries totally look like plain C.
- doboyy 7y agoNo doubt, but they are all non standard which is irritating for people who expect that out of the box.
- flohofwoe 7y agostd::string has its own share of problems, but at least it's useful for C programmers as a warning on how not to implement a string library (for instance: a string shouldn't be important enough to give it its own allocation, it shouldn't be mutable by default, there should be string manipulation functions which are actually useful in real world situations instead of being an academic exercise, and so on...).
- saagarjha 7y agoI’m not sure I understand your first and third complaints and the second one is why const exists.
- ChrisSD 7y agoDoes every org have their own C string library or does it just feel like it?
- nitwit005 7y agoAnd their own logging, hash map, and list utilities. Admittedly, people seem to write string and logging libraries even in languages that do provide them.
- longcommonname 7y agoAnd their own serialization format
- camgunz 7y agoThe last job I had working in C didn't--we leaned heavily on libraries for stuff like strings, logging, hashtables, and serialization. Implementing that stuff yourself is either a big timesink, or just asking for bugs and security issues.
- deleted 7y ago[deleted]
- dwheeler 7y agoI think memccpy is an improvement, so I support it, but it's still complicated to use in practice. The article gives this example for a copy followed by concatenation: char *p = memccpy (d, s1, '\0', dsize); dsize -= (p - d - 1); memccpy (p - 1, s2, '\0', dsize); Notice that you have to recalculate dsize (correctly without one-off errors!), it assumes you have dsize itself, and this doesn't detect overruns (which in many cases you should do). So real-world code would look more like this: // dsize is the space *available* in d, including \0 size_t dsize = sizeof(d); // if d is an array // ... char *p = memccpy (d, s1, '\0', dsize); dsize -= (p - d - 1); if (dsize <= 0) goto overflow; // handle overflow char *q = memccpy (p - 1, s2, '\0', dsize); dsize -= (q - (p - 1) - 1); if (dsize <= 0) goto overflow; // handle overflow It's a little easier to understand than strncat/strncpy versions, it doesn't unnecessarily read its inputs past where they are needed like strlcat/strlcpy do, and it's more efficient than snprintf. So yes, it's an improvement and I support it. However, this is still rather complex; in particular, it's way harder to understand compared to code that uses snprintf, and certainly harder to understand than pretty much any other programming language higher level than assembly. So let's accept this improvement, and keep striving to do better.
- imglorp 7y agoIf you want simple, it's hard to beat the original K&R (1978!) while(*p++ = *q++); which is what sold a generation on the whole idiom. Of course its time is long gone but useful to learn.
- dwheeler 7y agoThis construct: while(*p++ = *q++); is simple. But I agree with you, its time has long gone, because in many programs, this code is also wrong. This idiom assumes that that the source can never be longer than the destination. There are now a legion of attackers who will exploit this code and harm its users. Modern C programs often have to work in the presence of attackers.
- camgunz 7y agoYeah I think strlcpy is way better exactly for this reason. I'm sure the idea behind returning the pointer was call chaining, but you shouldn't ever be doing that anyway, and with strlcpy you basically can't do an off-by-one.
- zokier 7y agoC-coders of HN, do you use plain vanilla C strings in your project (s)? I was under the impression that most (at least bigger ones) use some custom length carrying string type to avoid exactly these sort of problems
- saagarjha 7y agoIn the projects I work on (which are usually small) I use plain C strings.
- kstenerud 7y agoIt depends. If I'm writing a library, it's bare pointers. If I'm writing something not a library that's big enough, I'll use a struct {size_t length; char* string;} where the length is the string length, and string contains (length) characters + a nul byte. I might even mix in allocation data for the total allocated size of the buffer if it's important enough. Simple to implement and use (and also backwards compatible), provided you have a library of common functions for allocating, copying, etc. If I'm size constrained, I'll consider uint16_t for the length field. If I'm REALLY size constrained, I'll use a VLQ [1] for the length field and take the slight performance hit. [1] https://github.com/kstenerud/vlq/blob/master/vlq-specification.md https://github.com/kstenerud/vlq/blob/master/vlq-specificati...
- billziss 7y agoThese days my work is focused on kernel/systems programming, so plain C (char or wchar_t) strings outside the kernel, and whatever the kernel requires inside it (e.g. UNICODE_STRING on Windows). When I do apps I have my own length-prefixed variant.
- camgunz 7y agoNo. I would use GLib or sds. If I'm writing a library I generally try to avoid allocation, so that usually renders the question moot in those cases.
- dooglius 7y agoThe right thing to do here is generally to avoid dealing with strings at all: you only need it to parse input and print/log output, beyond that everything should be using integer identifiers and handles.
- papermachete 7y agoWhy are embedded developers unnerved by the concept of a featureful standard library? There are Linux distributions aimed at statically compiled, musl-based packages, hence you can very much choose what you need for your project. Look at C++'s std::string in GCC, Clang, and MSVC's standard libraries and respectively its development history. Of course you can make a minimalistic standard string and also eliminate nullpointer checks, trailing \0 checks (everyone passes size_t len anyways), and allocation issues in runtime. The only standard thing about C strings are vulnerabilities.
- umvi 7y agoI used to work on routers that essentially ran embedded linux inside and our new projects started all being c++ after a while. It's amazing after the switch to c++, we basically stopped ever seeing string/array related segmentation faults now that developers use std string/vector/etc by default. Yeah, it's more resource expensive, but we have plenty of RAM now on these boards and it's totally worth not having a maintenence nightmare like our legacy C projects which I swear get a new bug report every other month about a newly discovered string/array-related segmentation fault (I'm not saying C++ can't ever be a maintenance nightmare, to be clear, but as far as memory related issues go, they disappeared the moment we started using std library).
- papermachete 7y agoInteresting, did you make your own allocators? Can you tell me the company?
- legulere 7y agoWhy not abandon \0-terminated string and pass the length of strings as an additional parameter?
- carapace 7y agohttps://stackoverflow.com/questions/25068903/what-are-pascal-strings https://stackoverflow.com/questions/25068903/what-are-pascal...
- legulere 7y agoPascal strings prefix the string with the length. What I'm arguing for is more like fat pointers, where the length is stored at the same location as the pointer. Something that is already standard to do in C for binary data.
- papermachete 7y agoBecause floating variables are ambiguous and you can lie about length (unintentionally too).
- wkz 7y agoI feel like one alternative is missing from the list: fp = fmemopen(buf, sizeof(buf), "w"); fputs("one", fp); fputs("two", fp); fputs("three", fp); fclose(fp); Not sure how it performs, but it reads pretty well IMHO.
- SignalsFromBob 7y agoThat web page's color choices have made it very difficult to read. I don't know who thought putting light grey text on white was a good idea. I had to copy and paste the text to a text editor in order to read it.
- dmortin 7y agoI usually just press Ctrl+A to select all in these situations to make the text readable.
- johnisgood 7y ago> Of the solutions described above, the memccpy function is the most general, optimally efficient [...] This does not seem to be the case for me AT ALL. strcpy for example, is a lot faster than memccpy. Here are my results: $ gcc -O0 bench.c && ./a.out memccpy: 0.008405 strcpy: 0.002913 $ gcc -O3 bench.c && ./a.out memccpy: 0.007933 strcpy: 0.002590 $ clang -O0 bench.c && ./a.out memccpy: 0.008771 strcpy: 0.003225 $ clang -O3 bench.c && ./a.out memccpy: 0.007966 strcpy: 0.000383 $ musl-gcc -O0 -static bench.c && ./a.out memccpy: 0.007849 strcpy: 0.005647 $ musl-gcc -O3 -static bench.c && ./a.out memccpy: 0.005754 strcpy: 0.005625 $ tcc bench.c && ./a.out memccpy: 0.014252 strcpy: 0.004045 Source code can be found here: https://slexy.org/view/s2EHngPvDh https://slexy.org/view/s2EHngPvDh --- The differences seem to be quite interesting. Did I mess up the code? Compare gcc -O3's strcpy and clang -O3's strcpy: 0.002590 vs 0.000383! musl-gcc on the other hand has much more similar results. --- $ gcc --version gcc (GCC) 9.1.0 Copyright (C) 2019 Free Software Foundation, Inc. This is free software; see the source for copying conditions. There is NO warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. $ clang --version clang version 8.0.1 (tags/RELEASE_801/final) Target: x86_64-pc-linux-gnu Thread model: posix InstalledDir: /usr/bin $ tcc -v tcc version 0.9.27 (x86_64 Linux)
- ryacko 7y agoLooks like something was optimized into being removed.
- johnisgood 7y agoIt could be, yeah, I did not have the time to check out the assembler code. I will do it tomorrow, but someone else may do it before me and elaborate on the reasons. :)
- eMSF 7y ago>Did I mess up the code? You certainly did; both your benchmark functions overflow their internal buffers. However, it's not all that improbable for memccpy to be slower than strcpy in this case, even if it was used correctly; after all, it does more, and by doing so would prevent the buffer overflow if supplied with correct arguments (specifically, the last one). As to how much slower, I cannot tell. Also, you're skipping over half the point of the article by using a (fixed) source string, the size of which is known in advance. In general, I would not benchmark library functions on local variables (buffers) without observing their results. It's far too easy for the compiler to remove the call altogether when said removal doesn't make any difference on the output.
- loeg 7y agoThe refusal of glibc (and I suppose, POSIX) to adopt strlcat/cpy is continually obnoxious. That said, if you actually need efficient string operations, you probably want a Rope data structure rather than any libc primitive.
- Gibbon1 7y agoI've never used a rope library, but having read a few essaying on those I keep coming back to it. I also keep coming back to hiding implementation details behind closures. I'm likely not smart enough to understand why that's a bad idea tho.
- ncmncm 7y agostrlcat and strlcpy are not adopted because they are a really bad design. Using them correctly takes more code than not using them. In practice they are never used correctly, making them, an "attractive nuisance", a feature that causes more trouble than its absence.
- loeg 7y agoThis argument doesn't hold water. They're absolutely no worse than strcat/cpy and strncat/cpy, which glibc implements. I totally disagree with your premise that truncation is an incorrect use. In reality, the alternative to strlcpy/cat isn't "force programmers to write correct code," it's "programmers will just use the crappier available functions with even worse behavior on overrun."
- ncmncm 7y agoPerhaps you have some better reason why Posix has rejected it, again and again? Some sort of conspiracy is conceivable, but in service of what? I have seen much, much better designs, that take into account that these functions are rarely called in isolation. In those, calls cooperate with previous and subsequent calls to share the burdens of maintaining correctness and safety.
- 7y ago
- pksadiq 7y ago> The committee chose to adopt memccpy but rejected the remaining proposals. Is that the case? Reading the updated standard draft[0] they also included strdup and strndup. May be they rejected first, then chose to add later. [0] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2385.pdf http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2385.pdf
- yyyk 7y ago"The strlcpy and strlcat functions are available on other systems besides OpenBSD, including Solaris and Linux (in the BSD compatibility library) but because they are not specified by POSIX, they are not nearly ubiquitous." This ignores how often they are (re)implemented in userland. glib, X and even the linux kernel have implementations. Perhaps we could just standardize what programmers chose rather than allow glibc an unjustified veto?