4 ms·
Documented behaviour is a little different to "effectively broken". The difference between strncpy and strlcpy is that strlcpy will NUL terminate the last byte
by jmts 8y ago
Documented behaviour is a little different to "effectively broken". The difference between strncpy and strlcpy is that strlcpy will NUL terminate the last byte for you always. There is nothing stopping you from doing the same thing yourself when you use strncpy. If you care enough, write your own strlcpy - it's only one extra line.
- simias 8y agoI'm not saying it's hard to work around but I maintain it's broken. You have a function that deals with C-string that in some conditions returns something that's not a C-string but can't be trivially distinguished from one and will trigger undefined behavior if used like one. It's terrible ergonomics and almost certainly not what you want to do in any situation. You can argue that truncation is an error condition but then it ought to notify it somehow, for instance by returning NULL in such a case. And even then it's incoherent with snprintf which doesn't have the same behaviour and does always terminate with '\0' even in case of truncation (assuming non-0 buffer length, of course). It's just an unnecessary footgun that serves no practical purpose. It would be like a date function that gives you the today's date except on the 4th of December where it replies that it's the 31st of February. Not hard to work around but still broken.
- caf 8y agoIt's not broken, but it is misnamed. This is because it is not intended to work with the same kind of string that the other str* functions work with (ie. an ordinary null terminated string). Instead it's supposed to work with fixed-width string fields that pad out values shorter than the field width with nulls. This is how original UNIX directory entries were stored. See how the name is copied into u.u_dbuf here: https://github.com/hephaex/unix-v6/blob/daa355109625a50e6b1080184dee30c9136549d1/ken/nami.c#L72 https://github.com/hephaex/unix-v6/blob/daa355109625a50e6b10...
- simias 8y agoI see your point but at this point I think it's just a matter of taste. I don't really see how having a function meant to deal with a special case of character buffers disguised as a general purpose string manipulation routine in the stdlib could be considered reasonable. I understand why it's here, I understand the history, I understand why it made sense at some point to have such a function but you won't be able to convince me that it's not broken or that it shouldn't be deprecated in favor of strlcpy (ditto for strncat/strlcat). After all it is in <string.h>, not <fixed-width-string.h>, it's pretty heinous that it fails at the very low bar of actually producing a valid C string every time (especially given the very high prejudice of having rogue unterminated "strings" in a C program).
- caf 8y agoIt seems fairly unlikely that the C standard would add strlcpy() and strlcat() when it already has strcpy_s() and strcat_s() in Annex K.
- heisenbit 8y agoA misname that may have cost cost billions in bugs and security issues.
- burfog 8y agoIt has also avoided security issues, some of which get created when people get the idea that every strncpy must be replaced by strlcpy. The strncpy function writes to the entire buffer. This is important if you will be passing the buffer across a security boundary, for example in a network packet or as a struct copied into a publicly visible file. If trailing bytes are not cleared, then secret data (which happened to be sitting in memory) can get leaked.
- dxhdr 8y agoI agree it's terribly named, however discussing the behavior of strncpy is important because the casual reader or new C programmer will see "don't use strncpy, it's broken" and then come away with the wrong idea. strncpy is not inherently broken but it's most likely not the correct function to use. As a C programmer it's important to understand why and what the alternatives are (and there are several).
- wruza 8y agoIf you had to name it according to unix traditions, how would you?
- buckminster 8y agoMaybe fldcpy and fldcat to make clear than a fld (field) is entirely distinct from a str. A better choice if you have a time machine would be to remove these from standard C altogether.
- jschwartzi 8y agoI can use a sledgehammer to break my leg, but that doesn't mean the hammer is broken by design. I just have to be careful where I swing the hammer and what I hit with it.
- ehaliewicz2 8y agoThat's a poor analogy. It calls itself a string function but isn't. So it's more like a airsoft gun that actually shoots bullets, but you don't know until you decide to shoot yourself in the leg to see what it feels like.
- simias 8y agoIf the sledgehammer defaults to "leg breaking mode" then it's broken. How often do you actually use strncpy actually intending it not to return a non-NUL terminated string on truncation? I hate this mentality in a lot of C circles that boils down to "there's no bad language, just bad programmers, man up pussy". I like C, I use it a lot, it's one of the first languages I learned and it's been my main "professional" language for more than a decade. Yet I can also see that it has many unnecessary sore points. Having switch not break by default, gets(), array shenanigans, some aliasing rules, the hundreds of completely different meanings for "static", macro hygiene and I could go on... You can say "it's not a big deal and it's not going to change at that point anyway" and sure, I'm not arguing for a revolution, but let's not act that it's an absolutely perfect language and I'm an idiot that doesn't get it for pointing out these issues.
- int0x80 8y agoAbsolutely. strncpy is a very badly designed interface. Is it documented? Yes. Does that make it any good? No. Principle of least surprise. Idiot proof. Semantics that match common usage. Call it what you like.
- jschwartzi 8y agoI'm not arguing that it's a completely perfect language. But broken is an incredibly strong word for something that works when used as described in the documentation. I think the API designer was trying to be parsimonious and not assume too much about what you want to do with the data in your destination buffer. The requirement was to provide buffer overrun protection when the source string is too large, and this function provides that. Beyond that, the requirement that the character array be null-terminated is your decision. In fact, if you determine that it isn't null-terminated you can do other things besides null terminate it yourself. You might want to provide a helpful error message to consumers of your API rather than truncating their input data silently. Or you may try reallocating your buffer until the string fits. There's actual error handling that you can do with this function. In contrast, automatically null terminating the string makes that more difficult. The other issues that you have with C seem like preferences. There's nothing wrong with not breaking by default in a switch statement as long as you know that that's what happens.
- pedasmith 8y agoA documented behavior which causes many security and crash issues is, almost be definition, effectively broken. Can you come up with an example of an API which you would consider to be effectively broken and yet is not actually broken? Presumably, it would be an API that's easier to misuse than strncpy.