10 ms·
Linux eliminates the strncpy API after six years of work, 360 patches
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1a3746ccbb0a97bed3c06ccde6b880013b1dddc1 https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
- mrlonglong 4mo agothe zero terminated string is I think is computing's biggest mistake. Pascal style strings were much safer.
- jackbucks 4mo agoIt was definitely an interesting way to allocate pointers. I did once have a very large project where devs didnt understand this and resolved hundreds or more off by one and memory overwrites in C due to this feature. But at the same time, I think blaming the software was kind of a cop out. Devs were in a hurry and simply didnt respect the rules. Given todays software engineer at large. Nerfing programming languages so they cant destroy things might not be a bad idea. But AI will nerf everything.
- fragmede 4mo agowhy is AI gonna nerf everything? sure it could be used as the easy button, but I just spent two hours this morning learning about the neuroscience of how memory works in the brain that I didn't mean to and now I want to run studies on how memory works. Why do you assume that AI is gonna nerf everything?
- AnimalMuppet 4mo agoAGI might. AI? No way. See, AI was trained on existing data - on all that existing C code out there (sure, and also on all the papers and articles saying what was wrong with that C code). Those bugs are in the training data, and often not marked as bugs. So when AI generates C code, is it going to avoid making the mistakes that human code made? No, it's going to generate the kind of code it was trained on. How could it be otherwise? That's not going to nerf anything.
- CamperBob2 4mo agoWhen's the last time you saw a decent coding model create a buffer-overflow bug while trying to use C strings? Serious question. Anyone else seen this happen in the last 12-18 months? If so, which model and version were you using?
- macintux 4mo agoWould you even know? Serious question. The volume of code the models can produce, the subtle ways these bugs can manifest (or even only manifest when under attack), it seems like they would be easy to overlook.
- CamperBob2 4mo agoI have a habit of getting GPT 5.5 to review everything Opus writes for me, and vice versa. The model in the reviewer role frequently finds things I overlooked myself. Occasionally in parts of the code I wrote. No modern LLM has found any buffer overflow bugs in parts of my code that originated from another LLM. Again, though, they have found one or two that were my fault.
- bigstrat2003 4mo agoHaving one clanker verify the output of another has minimal, almost non-existent, persuasive value as evidence.
- CamperBob2 4mo agoHopefully you're close enough to retirement age that it's a moot point for you. If not, you're headed for a bad time, a major attitude adjustment, or (most likely) both.
- smackeyacky 4mo agoI had Claude write a bit of stupid C# the other day that had an off by one string truncate. Surprised the hell out of me.
- dietr1ch 4mo agoI think it was NULL itself. It was a long way until we realised we don't want invalid values and could use the type system to help us use special values safely.
- jkercher 4mo agoMeh, I think NULL is fine in C. It's an extra, valid state to represent pointers at no cost. Unlike the more hand holdy languages, it's quite rare for a pointer in C to have the ability to be NULL since, more often than not, it's pointing at something known. It's actually quite rare to see NULL checks unless it's API code or something like that. I can see this being more of a problem in a managed language where anything can be NULL at any time.
- kelnos 4mo ago> to represent pointers at no cost I wouldn't call "cause of bugs and security issues" "no cost". > it's quite rare for a pointer in C to have the ability to be NULL As a C programmer for more than 25 years, that is the exact opposite of my experience.
- UqWBcuFx6NV4r 4mo ago[flagged]
- IgorPartola 4mo agoIs None OK in Python? NULL in C just doesn’t belong at the end of a string. But IMO having a “there is no value here” designation is not a bad thing.
- none_to_remain 4mo agoI think you're mixing up the NULL pointer and the NULL (sometimes NUL) character.
- jibal 4mo ago
- themafia 4mo ago> Pascal style strings were much safer. The limitations were brutal. Initially you could only have 255 bytes in a string. The length of a string and the size of the allocation are now separate and you may need to think about that unused memory in your design. The problem now doubles with the introduction of UTF-8. Your string size is in bytes and you need to track characters separately. If you want to create an array of strings you either need to specify the length of all strings and accept the memory overhead or have an array of pointers to strings. If you use an array of pointers you may end up choosing to use the 'nil' value as a sentinel that means "end of list." So we're right back where we started. -- Because someone decided to downvote this HN has limited the speed at which I can reply. This site is tragic and I'm fully done with it now. You can spread propaganda and poorly sourced zeitgeist and be among friends but if you try to have a genuine conversation about programming languages you are made to be unwelcome immediately. Screw this. -- > No other data structure works like this. The linked list. > You can't mess this up in an array C happily decomposes arrays into pointers. You can erase your length information from the type. This was an intentional decision. > Strings are the only data structure that assume there will be a NULL at end. Which is why almost every string API has a version that allows you to specify the maximum length. The fact that you can use a NUL doesn't mean you have to. Which is why the concept of "sentinel values" is broadly used in many types of applications you haven't considered here.
- AlienRobot 4mo ago>The problem now doubles with the introduction of UTF-8. Your string size is in bytes and you need to track characters separately. That isn't really a problem. The problem with null-terminated strings is specifically what happens when you reach the end of the allocated array and there ISN'T a NULL character. Every string function is designed to keep going until it finds the NULL character, so if a hacker gets rid of the NULL character, he can exploit pretty much any standard string manipulation function being used elsewhere in the program to manipulate whatever memory comes AFTER the string data structure. No other data structure works like this. You can't mess this up in an array, because no function that manipulates arrays is just going to keep going until there is a null. That would be stupid because it would require users of the function to add a NULL to the end of their arrays before passing it to the function, so instead we just pass the size of the array to everything. Strings are the only data structure that assume there will be a NULL at end. By the way, I read once that if you use UTF-32 every code point will be 4 bytes, constantly, but even then a single code point isn't necessarily a single character. Text is just complicated.
- msla 4mo agoIn addition to having to pick a size for the length counter and then, later, having to differentiate between lengths in bytes, codepoints, and glyphs, you can't subdivide a Pascal string using pointer arithmetic. To pass just the end of a string into a function, you have to either copy the tail of one Pascal-style string to another with a smaller size value, or your string has to be a struct with an integer and a pointer to the actual data instead of just an integer stuck on the beginning of the string. The first is a lot of copying in some cases, the second raises the specter of structs with invalid pointers. That's not to mention the potential problems that would cause with caches.
- estebank 4mo agoThe third option is to have a variable width length: the top most bit signals whether the next byte corresponds to the length or to the start of the string.
- cornholio 4mo agoYou can have a universal variable length field, for example 2 bytes for strings < 32768, then four bytes, 8 bytes etc. On the critical short string path, it costs just a single bit test. The glyph vs byte issues need to be dealt with in both formats. The subdivision issue is a good perspective, but i would argue the performance impact of cloning substrings is dwarfed by the redundant full string reads to find length.
- lelanthran 4mo ago> You can have a universal variable length field, for example 2 bytes for strings < 32768, then four bytes, 8 bytes etc. To hold the length of a string, I'd do something similar to unicode: 7-bits for size + 1-bit for continuation, then 15 bits for size + 1 bit for continuation, then 23-bits for size + 1 bit for continuation, etc. Or maybe even do it exactly the same as unicode: 0XXX XXXX -> length of string is in those 7 bits 1XXX XXXX XXXX XXXX -> length of string is in those 7+8 bits 11XX XXXX XXXX XXXX XXXX XXXX-> length of string is in those 6+8+8 bits ... > On the critical short string path, it costs just a single bit test. A few more clock cycles compared to NULL-termination, although my alternatives above require even more clock cycles. If the hardware had instructions for sentinel values, things would be easier (Like how DOS calls used '$' termination for strings) and safer. Load a sentinel byte into a register and have dedicated copy and compare instructions that take each two addresses (src and dst) and copies (or compares) src/dst until the terminator is reached (with copy copying the sentinel as well). Considering that sentinel values are needed so often, and are so useful, it's surprising that this is not in any ISA. What we have now is kludgy workarounds in the HLL for this. It's hard to blame the HLL, because some workaround has to be implemented.
- bsder 4mo agoZero terminated string is a special case of sentinel value termination. And sentinel value terminations make a lot of sense when you have punch cards and fixed length records that you need to carve into pieces. Nobody expected any decisions they were making in the 1960s and 1970s to have any bearing on computing a half-century later. They all expected to have their mistakes long papered over by smarter people at some point. But we ALL make the mistake of underestimating inertia.
- fragmede 4mo agocompared to Von Newman versus Harvard architecture for LLMs? I think that's a far bigger mistake.
- pjc50 4mo agoNeumann, and .. what? In what way?
- fragmede 4mo agoPrompt injection only works because there isn't two streams of input to give to the LLM. Von Neumann being the architecture with a single shared memory for both data and instructions. If there were a clean way for the LLM model to distinguish between system messages vs user messages, we wouldn't have that problem.
- amomchilov 4mo agoI don’t know how you could keep the two isolated, without drastically dropping up the utility of LLMs. Part of their wonder is how they can behave differently depending on the data they’re working with. We like that feature when the data is the “good stuff” (docs, compiler messages, etc.), but how you tell that apart from “bad stuff” (prompt injection on official-seeming pages). We basically expose LLMs to the same social-engineering vulnerabilities that humans have.
- layer8 4mo agoAlmost as bad as newline-terminated lines. ;)
- sourcegrift 4mo agoWhat's bas about them and what are the alternatives, genuinely curious since never seen them spoken of
- smackeyacky 4mo agoZero terminated strings were the basis for an awful lot of useful software. Calling them the biggest mistake in computing is a bit OTT. I haven’t programmed anything Pascal related for 30+ years but I dimly remember thinking at the time that I wished the string system wasn’t so hard to use.
- asdfasgasdgasdg 4mo agoThat useful software would not have been less useful if the strings in it were represented as size + buf.
- crackez 4mo agoOh really? Have you tried to rewrite anything to put your theory to the test? I don't think it's as straight forward as you think it is...
- smackeyacky 4mo agoExactly. The pascal I used had no way to dynamically allocate a string they were all fixed at compile time. That really sucked.
- ComputerGuru 4mo agoThat argument isn’t valid. The argument would be “this string design enabled a whole lot of useful software” but that’s a different matter. (And it could very well be the case.)
- zzrrt 4mo agoLead was the basis for an awful lot of useful gasoline. Doesn't mean it was the only solution or the best one.
- JdeBP 4mo agoA more accurate re-phrased version of the original is that they are the biggest mistake in the C language. * https://news.ycombinator.com/item?id=48614913 https://news.ycombinator.com/item?id=48614913 * https://news.ycombinator.com/item?id=24454369 https://news.ycombinator.com/item?id=24454369 * https://news.ycombinator.com/item?id=1014533 https://news.ycombinator.com/item?id=1014533
- dmazzoni 4mo ago255 characters ought to be enough for everybody, right?
- RetroTechie 4mo agoYou mean bytes.. we have multi-byte characters now.
- BobbyTables2 4mo agoPartly agree but there would have been squabbling on the data type of the size, unless it was variable length. The latter would have had other issues too. For a while, 16bit would probably have seemed too extravagant. Now 32bit would probably seem too small. For a “strongly typed” language, C is pretty damn loose where would have mattered.
- poly2it 4mo agoNo, there would not have been and this is most likely not the reason. size_t exists for precisely this use case. It has existed since C89.
- DarkUranium 4mo agoI like the D approach where arrays are just `struct { size_t length; T* ptr; }` internally --- and strings are just arrays of `immutable(char)`. It has a big advantage over the Pascal approach in that you can do zero-copy slicing, since the length is separate from the actual data. And `size_t` makes perfect sense for the length here. If your strings are longer than the address space (which `size_t` technically isn't, but is practically very strongly correlated to it), then you're going to have a problem regardless of the number of bits for the length anyway.
- astrobe_ 4mo agoThis only makes a difference in terms of memory size, not in terms of speed, because for decades processors and compilers have been optimized for moving bytes around. But one would note that in order to gain memory for this particular case of slicing, one introduces 2 extra words (size and pointer) for every other cases. Like perhaps the second most common string operation, concatenation. In those other cases, the benefit is slightly negative. I've had extensive experience with "counted strings" because I implemented a bunch of Forth interpreters which also uses this scheme. Including the common trick of using counted and zero-terminated strings, which is the worst of both worlds in the end. Forth is the kind of language that quickly show you how bad your choices are. I eventually dropped all that and adopted ASCIIZ strings because they are generally more efficient (if you pay attention to the strlen() performance pitfalls) and having a dead simple interface with the rest of the world (OS, libraries) is more valuable.
- mikewarot 4mo ago[dead]
- Conscat 4mo agoClang and GCC both let you use Pascal strings in C if you would like (with `\p`). But Pascal strings aren't that useful today because the maximum length is too short.
- jxbdbd 4mo agoWhy would a pascal string be any shorter than a C string? A C string is one pointer reaching all of memory, a Pascal string is two pointers reaching all of memory
- bc_programming 4mo agoA pascal string is a single byte with the length, followed by the data. Some implementations use more bytes for the length data, such as Delphi which changed over to a 4 byte prefix length, though those aren't technically Pascal strings anymore. I can't find anything about a Pascal string being two pointers?
- jll29 4mo agoIt is conceivable, for both Pascal and C, to have more than one string implementation side by side, so the developer can choose to use the best-fitting one. In C++23, variant<> permits to do what Rust's typed enums introduced (e.g. Result sum type that is either a "real" result - with result type - or an error - with error type -, each strongly typed). If you do that, a definition like class IString { /* basic string functions */ }; class MiniString : public IString {}; class CZeroTerminatedString : public IString {}; class PascalString : public IString {}; class CppString : public IString {}; use String = std::variant<MiniString, CZeroTerminatedString, PascalString, CppString>; // define one type for all impl. permits to define string functions that operate over the sum type String, and which use the methods defined in the interface IString, and which then work for all string implementations. The developer can then pick the most suitable implementation, i.e. CMiniString for very, very short strings (that fit into 64 bits, so approx. <= 8 UTF-8 characters), CZeroTerminatedString (for char *co = "test\n"; zero-terminated old style C strings), CPascalStrings for strings that carry a length in s[0] or as a struct member or class field, and CppString as a wrapper for the C++ std::string that implements IString. Sum types are a type-safe and memory-preserving way to do what in the older days was sometimes implemented using a "union {}" (which was not type-safe).
- lelanthran 4mo ago> the zero terminated string is I think is computing's biggest mistake. No. They had trade-offs to make, and sentinel-based sequences are a needed thing, even outside of strings. The mistake was that ISAs never looked at what HLL needed, then add the necessary instructions (I posted more about this below). Even NULL is not a big mistake, when looked at in context of the time in which it was developed.
- layer8 4mo agoThere is a middle ground that Visual Basic (and then COM) took, with the BSTR type: It’s still a pointer to a zero-terminated char array, but there is a length field immediately preceding the first pointed-to byte. This is still compatible with a C string (assuming no embedded null characters), but BSTR-typed functions can take advantage of the length value.
- tremon 4mo ago> This is still compatible with a C string Strictly speaking, it's not alignment compatible from CString to BSTR unless you declare all strings to be at most 255 characters or the cpu architecture doesn't require aligned access for multi-byte words (like x86). The BSTR alignment must match the alignment of the length word, meaning you can't convert a randomly-aligned C string to BSTR by simply attaching a prefix in-place. Also, having the length embedded in the value rather than in the pointer makes it impossible to create BSTR (sub)slices without performing a memcpy. Fat pointers do not have this restriction.
- layer8 4mo agoA BSTR object is compatible with functions expecting a C string. The other direction obviously never holds, unless the C string is a BSTR to start with. Yes, there is a trade-off between slices using the same format and having compatibility with C strings. Hence “middle ground”. You can still use a string-slice type on top of BSTR, it just would be a separate additional type. Note that languages like Java also don’t have a singular type for strings and string slices.
- jiggawatts 4mo agoThese are great for "data smuggling" attacks where one layer of code assumes the length is 'x' and another layer assumes it is 'y'. It makes hybrids like this very dangerous for anything even remotely security-adjacent, such as roles, tokens, etc. This kind of thing caused the CVE-2009-2408 and CVE-2009-2510 "Null Truncation in X.509 Common Name Vulnerability."
- badsectoracula 4mo ago
- sourcegrift 4mo agoRust has "pascal style strings" (quotes because the concept is slightly different) so it's not a done deal
- larodi 4mo agoWonder when is someone going to brave and fork the linux kernel and try to ffwd it with automatic programming.
- fragmede 4mo agowhy would you start there instead of creating something from scratch ?if you can port drivers just as easily meaning you don't especially give a shit about hardware you're running on in the first place, why even deal with linux? The battle tested LRU cache system?
- literalAardvark 4mo agoIt's much easier to use something with all the edge cases already handled as a starting point.
- convolvatron 4mo agoI've seen several workalike kernels in various stages of completion. at least one of them was able to run some pretty substantial applications (Postgres, nginx, that kind of thing), and that is still I guess around 250kloc. but it only really has drivers to support hypervisor devices. unfortunately as time goes by, the linux api surface gets larger and more convoluted. so there's going to be some coverage you're just never going to get. but in the abstract, definitely. linux is so bloated at this point that its not clear that it can ever be 'made safe'.
- larodi 4mo agoSome if not most coverage will be off, indeed, but then the important stuff can get you lots of benefits. This makes sense even today for selectively patching the kernel. I’m sure many people been odd by the complexity of it while now it is doable albeit with agents…
- larodi 4mo agoWell in reality if you want a custom OS perhaps scavenging parts is a thing to do indeed. I just speculated whether Linux can be further improved by automatic programming and still keep the handmade parts.
- PlunderBunny 4mo agoI worked on a Win32 app that used space-padded strings, i.e. the destination string was padded with spaces, but there was still a null on the last byte. You had to use special versions of the string functions for length, copy etc. I’m not sure why this was - the source base was so old it might have had its origins in Pascal struct behaviour.
- egorfine 4mo agoI think this behavior has its roots in COBOL, not pascal.
- kps 4mo agoWhich has its roots in punch cards, where pre-computer hardware operated on fixed-sized fields and an unpunched column is equivalent to a space.
- bebe83939 4mo agoPerhaps prevent realocation when string size changes? Or aligning cpu cache lines?
- jkfkfkj 4mo agoIt can perhaps be due to the string originating from a sql database ”char” field, I.e. not ”varchar”. Char fields in databases are space padded.
- naturalmovement 4mo agoA reminder that we've had strlcpy[1] for ~ 30 years but it was never accepted into the Linux world because of typical petty open source bullshit. This is why we can't have nice things. [1] https://man.openbsd.org/strlcpy https://man.openbsd.org/strlcpy
- BoingBoomTschak 4mo agoActually, glibc 2.38 has it.
- naturalmovement 4mo agoWow it only took them 26 years to import a 30 line C function, a third of which is comments? I should have sent them a nice fruit basket to commemorate the occasion.
- ericbarrett 4mo agoThe Linux kernel had strlcpy over 20 years ago. It was removed in favor of strscpy because the latter was judged a better interface. Here's a 2022 article: https://lwn.net/Articles/905777/ https://lwn.net/Articles/905777/
- avadodin 4mo agoReturning an error is better but you're using ssize_t which is a tradeoff. The race conditions appear to be a result of the Linux kernel implementation but UNIX style syscalls introduce these races by default. It is not an inherent flaw of the API or even the implementation Linux was using. The only useable C string API has always been memcpy anyways.
- senfiaj 4mo agoI wonder, why not use a string buffer paired with its length? For example, maybe use struct that has char pointer, and 2 ints (occupied length + total buffer length). Almost like c++'s std::string. This null terminator thing really sucks, it's potentially insecure and often unperformant.
- GalaxyNova 4mo agoYes I have seen it happen a few times with `strlen` being called in a loop silently causing O(N) to turn to O(N^2)
- senfiaj 4mo agoExactly, you can't write clean concise code when working with c strings. Almost every c string manipulation requires cognitive load: "Is the buffer size enough (including null terminator), should I reallocate it?", "I need to have the offset from the last concat, to make next concats performant", "Umm, shold I put null terminator at i or i + 1?"... It really sucks, it's akin to death by thousands of cuts.
- jkrejcha 4mo agoReminds me of an article[1] that described how he cut GTA Online loading times by 70% because strlen was getting called for effectively every character in a string [1]: https://nee.lv/2021/02/28/How-I-cut-GTA-Online-loading-times-by-70/ https://nee.lv/2021/02/28/How-I-cut-GTA-Online-loading-times...
- sweetjuly 4mo agoI remember reading this blog post when it was first published, but the subsequent updates are better than I would've ever expected this to turn out. Worth checking it out again if you've seen it before :)
- sgerenser 4mo agoJoel Spolsky coined the term “Shlemiel the Painter’s Algorithm” for this type of thing back in 2001: https://www.joelonsoftware.com/2001/12/11/back-to-basics/ https://www.joelonsoftware.com/2001/12/11/back-to-basics/
- D-Coder 4mo agoNote that "360 Patches" is 360 uses of strncpy that have been removed, not necessarily bugs.
- dpark 4mo agoI would imagine 360 patches removed way more than 360 uses of strncpy. But yeah, it’s not a given that each of these patches addressed a bug. (Also not a given that there were only 360 bugs fixed.)
- lambdaone 4mo agoThis sort of boring grind is where the real work of systems engineering is done. Big infrastructure projects like this work on making the Linux kernel more reliable while still keeping it workable throughout the process move on the scale of decades, not months.
- appplication 4mo agoOn one hand I understand why it’s decade scale (the long tail of users/dependencies is really, really long) but on the other hand it doesn’t feel like a tenable pace at which we can make meaningful long term progress. Less of a gripe and I guess more a paradox of critical infrastructure.
- mvdtnz 4mo ago[flagged]
- devsda 4mo agoDid anybody else misunderstand the title as removing strncpy func for linux users ? For a moment, I misunderstood it as (g)libc removing strncpy and was worried about the trouble its going to cause.
- WalterBright 4mo ago"The strncpy function within the Linux kernel has been a "persistent source of bugs" for years due to counter-intuitive semantics and behavior around NUL termination along with performance issues due to redundant zero-filling of the destination." Huh. Whenever I've been asked to review C code, I always looked for strncpy and always found a bug with it.
- qarl 4mo agoAm I going to be the first person to ask this after five hours? Really? Wouldn't this work be extremely easy to implement with an LLM coder?
- deleted 4mo ago[deleted]
- qustio 4mo agoI don't think the bottleneck was that it took six years to Ctrl-F strncpy and type in new code for each file.
- qarl 4mo agoIt's a shame you're misrepresenting what is actually going on. In another comment here I explained that I have run a test: asking Claude Code to add a substantial feature to 270 different C programs. Despite your beliefs - it went extremely well.
- qustio 4mo agoHuh, are you confusing me with someone else? I don't doubt Claude Code did that, I do the same for refactors all the time. But xscreensaver theme tweaks for personal use have a much lower standard for quality control, regression testing, side effects, etc than a kernel used by billions of devices with thousands of interconnected drivers and subsystems. Not to mention the coordination problem to get every maintainer on board and patches approved for each specific area when working on a project of that scale, even for a relatively narrow change. Claude Code doesn't really help with that so don't see why the expectation would be a significant speed up (and doing it all in a single patch would definitely be rejected).
- qarl 4mo agoYes, I understand the difference in rigor. I refuse to believe the six year delay here was getting people to test a patch. Which, actually, Claude Code will also do quite well.
- Animats 4mo agoThis is a job for Claude! What happens if you turn a job like that over to Claude Code? A mess? Good results? Code bloat? Worth trying on existing C programs.
- qarl 4mo agoI ran a test where I added a "light" mode to xscreensaver: unique changes to over 270 different C programs. It mostly did an amazing job in a short period of time. EDIT: Of course I get downvoted for saying this. HN isn't interested in reality any more.
- ninjin 4mo ago> Of course I get downvoted for saying this. HN isn't interested in reality any more. I suspect that rather many of us are simply just tired of Claude and friends getting shoehorned into any conversation about programming at this point. It is about as fun as the Rust Brigade entering any discussion about C. It adds nothing new to the discussion and it is frankly tiring since we pretty much at any time have a handful of conversations on the front page already covering "AI" topics anyway (counting four at the time of writing this).
- qarl 4mo agoWell - except in this conversation it's incredibly relevant. It took six years to do this work when the work is likely mostly mechanical and could have been done much more quickly and safely with an automated system. I thought automation would be interesting to HN - given the context and the fact it was not used.
- krupan 4mo agoAn LLM is not a mechanical automated system. A deterministic search and replace would be a mechanical automated system. Clearly it wasn't that simple of a problem though.
- 4mo ago
- jibal 4mo agoThe purpose of strncpy, which was originally part of the UNIX kernel code, was to copy file names to and from directory entries that consisted of a 2 byte inode number and a 14 byte zero-padded but not zero-terminated name field. I started warning my colleagues against using it the moment I saw it for the first time about 50 years ago.
- dare944 4mo agostrncpy appears somewhere around the Unix v7 time frame, however only as function in the standard C library. It is not used in the v7 kernel itself.
- jibal 4mo agoThe code for strncpy was in the UNIX kernel since at least V6. It was eventually added to the C library under the name strncpy. Sometimes those entries were processed in userland, e.g., by fsck. The utility of strncpy is noted in the C89 rationale (FWIW I was once a member of X3J11, the C89 standards committee): "strncpy was initially introduced into the C library to deal with fixed-length name fields in structures such as directory entries. Such fields are not used in the same way as strings: the trailing null is unnecessary for a maximum-length field, and setting trailing bytes for shorter names to null assures efficient field-wise comparisons. strncpy is not by origin a "bounded strcpy," and the Committee has preferred to recognize existing practice rather than alter the function to better suit it to such use." And I just found this comment from John Mashey (I never met John but he and I both worked under Ted Dolotta, John at Bell Labs and me at ISC in Santa Monica): https://softwareengineering.stackexchange.com/questions/438025/what-was-the-original-purpose-of-c-strncpy-function https://softwareengineering.stackexchange.com/questions/4380... "I can answer definitively, since I wrote the originals ~1977, having moved from BTL Piscataway to Murray Hill. They were first named str*n, but were later renamed strn*, as there was some system in BTL that needed first 6 letters of external names to be unique. I was working on kernel & user code that supported rudimentary per-process accounting, which started with someone else, but needed extensions due to big increase in UNIX systems in computer centers, who wanted more performance analysis. I.e. this was supported by commands like accton(1), acctcms(1),acctcom(1), acctmerge(1) (all in UNIX/TS 1.0, Nov 1978, which was ~Research V7 with first steps of PWB/UNIX influence. Think of that as 1.0, then PWB/UNIX 2.0, then UNIX System III... The records described in acct(5) held the last 8 characters of the command pathname,truncated if necessary and thus possibly not null-terminated. I found multiple instances of inline code to manipulate these, which seemed a bad idea, so I wrote the str*n functions and replaced the inline code, and also used them in the various commands. I also thought it was a good idea for better code safety.:-) Sigh."
- twothreeone 4mo agowow, very humbling. I'm actually amazed how many people contributed to this. It's easy to get attribution for "cool new features", but arguable removing bad features is even more important for something as fundamental as the kernel. Cudos! I'm sure these are the sorts of things that will go down as folklore from the "founding ages", when everyone will have forgotten how to understand source code in 50 years and the Claude/Codex cruft just silently keeps piling on and burning the majority of our planets energy.
- skywal_l 4mo agoReminds me of Deepness in the Sky (Vernor Vinge) where a guy maintains a ship by doing software archeology. He is the only guy who knows what the Unix epoch is.
- needusername 3mo agoDoesn‘t he hack the ship and confuse the Unix epoch with the moon landing?
- nuc1e0n 4mo agoI'm of the opinion AI slop code will become untenable way before that.
- bigstrat2003 4mo agoI'm of the opinion it already is. But it seems that many people have yet to reach the point of being fed up with it.
- zaik 4mo ago> everyone will have forgotten how to understand source code in 50 years I don't think this will happen. Human desire to understand how things work will still be around in 50 years.
- DerSaidin 4mo agostrtomem_pad seems redundant with memcpy_and_pad, and also it requires the preprocessor: https://github.com/torvalds/linux/blob/1a3746ccbb0a97bed3c06ccde6b880013b1dddc1/include/linux/string.h#L407 https://github.com/torvalds/linux/blob/1a3746ccbb0a97bed3c06... I was curious: Why have it, instead of just using memcpy_and_pad? AI's answer (paraphrased) was * Avoid possible bugs from manually write sizeof(dest) * Enforces the __nonstring Attribute * signals: "I am converting an actual C-string into a fixed-width legacy memory field." vs copy binary data & pad it. Interesting to learn about the __nonstring attribute: https://github.com/torvalds/linux/blob/1a3746ccbb0a97bed3c06ccde6b880013b1dddc1/include/linux/compiler.h#L225 https://github.com/torvalds/linux/blob/1a3746ccbb0a97bed3c06... https://github.com/search?q=repo%3Atorvalds%2Flinux+__nonstring%3B&type=code https://github.com/search?q=repo%3Atorvalds%2Flinux+__nonstr...
- cm2187 4mo agoA lot of pain and suffering to avoid having a string datatype.
- edoceo 4mo agoWhat's a way they could get a strong data type here? Wouldn't that also require a large refactor of the code around strncpy to use the type and its functions?
- cm2187 4mo agoToday yes, but 40 years ago someone made the decision that a string was a char array and that every string manipulation going forward would require manipulating arrays. Talking about costly decisions. It’s actually interesting to compare the pain and suffering of switching to a string datatype in the 80s (refactoring the limited code base then) vs the next 40 years of unnecessary boiler plate syntax and bugs for not having this type in key APIs.
- tialaramex 4mo agoLinux doesn't exist in the 1980s, Linus started this work in the 1990s. But yes, the string slice type should have existed in C89 and it's very obvious from here that not having something of this sort - maybe what Rust would call &[u8] the reference to a slice of bytes - was a big problem for C. The correct way to represent this is what's called a "fat pointer". A pair of values, one is a conventional "thin" pointer to the start of the slice, and the other is a count. Your register pressure increases in the compiler backend but problems are significantly reduced because you have fewer bounds misses.
- mirsadm 4mo agoI'd be curious to see how much CPU time is wasted on looking for a null every time strlen is called. The extra length integer is probably insignificant compared to that.
- sirwhinesalot 4mo agoI have in the past made fun of the Linux kernel devs, supposedly some of the best C developers in the world, for not knowing how to make stringbuffer and stringview types, but to be fair to them we didn't have the consensus we have today on the topic. You know who did have the right idea though? Dennis Ritchie, who proposed a fat pointer type for C all the way back in 1990. Would have made for a perfect addition to C99. Imagine how different the world might have been had the committee added that in. We had a second chance with the release of the "C's greatest mistake" blog article from Walter Bright in 2007, essentially pushing for the same idea as Ritchie (slices/stringviews) but explained with much clearer language. Alas, didn't make it to C11. We're now in C23, still nothing. But we did get _Generic and VLAs! Party hard.
- pjmlp 4mo agoThat only goes to show where WG14 priorities are.
- anal_reactor 4mo ago> but to be fair to them we didn't have the consensus we have today on the topic. This is my pet peeve of teamwork. We can choose solutions A, B or C. Each has upsides and downsides. We debate for two weeks, then we choose nothing.
- david-gpu 4mo agoIsn't that due to a lack of leadership, rather than a problem with teamwork itself? Somebody has to be ultimately in charge and willing to put a stop to endless debate.
- gbin 4mo agoIt is more like a "design by committee" issue to me. The decisional structure needs to be built for some opinionated decision. When you have a committee there is not one clear thing they are solving for as everyone has an agenda to tug the language towards their own interests. The result of that can certainly be inaction because it is easier to say no than yes.
- pjmlp 4mo agoNow lets put that work into money, to assert what was the cost impact of replacing strncpy().
- rswail 4mo agoThings that have bugged me for 40 years... * NUL terminated strings (and now, non UTF-8 encoded strings on input/output) * Using LF or CR or CRLF as line terminators, and pipe/comma-delimited fields when there were other unambiguous ASCII characters that could have been used (eg, GS, FS, RS) that would have made the encoding/decoding of line termination an I/O thing keeping HT/VT/CR/LF/FF as literally print related codes.
- flohofwoe 4mo ago> non UTF-8 encoded strings on input/output UTF-8 on stdin/stdout works perfectly fine (unless you are on Windows of course, which is stuck in in the early 90s when it comes to international text encoding). > Using LF or CR or CRLF as line terminators This is also an operating system convention, and it would be better if programming languages wouldn't try to "guess" the correct line endings, since this causes more problems than it solves - but again, this is mostly a Windows specific problem, and it's Microsoft's job to finally bring Windows into the current century.
- rswail 4mo agoNo, it was an Apple, Unix, and Microsoft problem. Unix used LF, Apple used CR, Microsoft used CRLF. They are all ASCII carriage movement codes, which is about driving the paper feed and print head of an ASR-33 or equivalent. So they all made the "wrong" decision about what to store in a file. They just chose different wrong characters.
- flohofwoe 4mo ago> Apple used CR Apple hasn't been using CR since the release of OSX (26 years ago). Microsoft could have made the switch at any time too (just as they could have switched to UTF-8 as universal text encoding on Windows), they just choose not to. In the end it's not the job of programming languages to clean up Microsoft's mess ;)
- bobmcnamara 4mo ago
- rswail 4mo agoIn all the comments in this thread it's interesting how people confuse: * NUL: An ASCII non-printing character with the byte value of 0 * NULL: A pointer that does not point to usable memory with the value that compiles in C to be equal to ((void *) 0).
- layer8 4mo agoNUL was always just an abbreviation for null: https://www.rfc-editor.org/rfc/rfc20.html#section-4 https://www.rfc-editor.org/rfc/rfc20.html#section-4 I don’t think anyone in this thread is confusing the null character with the null pointer.
- rswail 4mo agoI've seen a lot of confusion, where people are talking about checking for a NULL at the end of a list of pointers which is very different to a NUL at the end of a string. Yes it was an abbreviation in ASCII, as are all the non-printable first 32 codes.
- kstenerud 4mo agostrncpy is 99.999% of the time NOT the correct function to call, so this is a huge win. It's just a shame that such a confusing name was chosen for such a niche use case (fixed width records that require null padding).
- stcg 4mo agoI wonder what is the difficulty in rewriting strncpy uses that makes it take six years? Was it widespread? Or was it more of a long going effort, where it was only changed if there were some changes in the same file? Or is there some other thing that makes it difficult?
- GTP 4mo agoI always thought that srncpy was the safe alternative to strcpy. Now that I think of it, I'm unsure if the NUL terminator is counted into strncpy's size or not, which would be a likely source of errors. But, could someone explain better what the problems were? And also, would have to pick the right function in the list of given alternatives much better?
- GabrielTFS 4mo agoThe issue with strncpy is that it doesn't actually necessarily terminate - in fact in any case where the source is larger than the destination it will just leave it unterminated (like, it will copy the last character it can from the source instead of terminating the destination string with a NUL)
- rurban 4mo agoNo, the safe alternatives end with _s. They do check matching buffer sizes, and enforce zero-termination. Unfortunately WG14 hates them also, because Microsoft. Microsoft did indeed break some of the, but you can use better alternatives, like my safeclib
- thiht 4mo ago> In place of strncpy, Linux kernel code should use strscpy() for NUL terminated destinations, strscpy_pad() for NUl-terminated destinations with zero-padding, strtomem_pad() for non-NUL-terminated fixed-width fields, memcpy_and_pad() for bounded copies with explicit padding, or memcpy() for known-length memory copies What a nightmare, does it have to be so convoluted?
- MarkMarine 4mo agoPerformance. A safe Swiss Army knife function that did most of this would be slow because of the internal branching you’d need to be safe, and because there is developer intent in the selection of these functions. I’d rather have the choice and clear dev intent when I see the function used when reading code.
- hahn-kev 4mo agoCouldn't they at least give them better names?
- bobmcnamara 4mo ago> does it have to be so convoluted? Getting strncpy right always has been
- henrypoydar 4mo agoNo code is faster than no code.
- Trialog 4mo ago[dead]