39 ms·
The case for a modern language
- tragomaskhalos 5y agoEither I'm going mad - in which case please set me straight - or the Rust example doesn't even compile: had to remove the odd-looking borrows on the method calls, and replace the type annotation in the final 'if let' with a turbofish on the call.
- tlamponi 5y agoNope, you're right, I need to apply the following diff (wrapped in a fn main() {}) to avoid rustc complaining: # diff -up test.rs test.fixed.rs --- test.rs 2022-01-22 16:03:57.302742242 +0100 +++ test.fixed.rs 2022-01-22 16:03:27.766250377 +0100 @@ -2,13 +2,13 @@ fn main() { // pretend that this was passed in on the command line let my_number_string = String::from("42"); // If we just want to bubble up errors - let my_number: u8 = &my_number_string.parse()?; + let my_number: u8 = my_number_string.parse()?; assert_eq!(my_number, 42); // If we might like to panic! - let my_number: u8 = &my_number_string.parse().unwrap(); + let my_number: u8 = my_number_string.parse().unwrap(); assert_eq!(my_number, 42); // If we're a good Rustacean and check for errors before trying to use the data - if let Ok(my_number: u8) = &my_number_string.parse() { + if let Ok(my_number) = my_number_string.parse::<u8>() { assert_eq!(my_number, 42); } }
- tialaramex 5y agoOne of the best things to happen to C++ discussions was people starting to write godbolt links for their code. Immediately the code being discussed becomes code somebody actually compiled and maybe tried running and not just "oh, ignore the fact it's syntactically invalid - you know what I meant". No we don't. You can obviously write a godbolt link for Rust too, but Rust's playground is also a reasonable choice. I think that code maybe makes more sense with a turbofish for each parse call and type inference, but maybe that's just a matter of taste. If the author had used godbolt or playground or whatever they'd have written code that compiles and we'd not be guessing.
- stncls 5y agoFrom the article: char *forty_two_bee = "42b"; char *end; errno = 0; // remember errno? long i = strtol(forty_two_bee, &end, 10); > This will return 0 No, this will return 42. strtol() parses greedily until a character cannot be parsed, but then it returns the conversion of what it did parse. I guess the fact the author got this wrong... kind of proves their point that strtol()'s API is not great? On the other hand, while the article purports to criticize a language, it then proceeds to only cover its standard library. Sure, C's stdlib is old-fashioned, but there are many things in C that are much worse than its standard library! (And I say that as someone who still likes the language.)
- EVa5I7bHFq9mnYK 5y agoThat case has existed for 40 years now, yet C still stands. Guess its the power of network effect.
- PaulDavisThe1st 5y agostrto*() is the wrong API to use if you care about errors. char* forty_two = 42; int i; if (sscanf (forty_two, "%d", &i) != 1) { /* error */ } Sometimes, there's more than one way to skin a cat, and one of them is more suited to the task at hand.
- JulianMorrison 5y agoNow what happens if you pass it a string without a null byte terminator?
- PaulDavisThe1st 5y agoI made no claim that using sscanf instead of strto* fixes all the issues with C string represenation.
- dcposch 5y agoAuthor mentions four increasingly obscure C replacements (first I've heard of Odin) without mentioning that the creators of the original C and Unix went on to make Go. Go does not have manual memory management. Despite (actually because of) that captures the spirit and design goal of the original C beautifully. It's a minimalist systems programming language. One of the amazing things about Go is the standard library-- the thing he complains about with C. The Go standard library is incredibly readable. It's night and day from C/C++ where opening glibc/STL etc is assault on the senses.
- dkbrk 5y agohttps://fasterthanli.me/articles/i-want-off-mr-golangs-wild-ride https://fasterthanli.me/articles/i-want-off-mr-golangs-wild-...
- deleted 5y ago[deleted]
- treeshateorcs 5y agothe rss feed is broken on that site, it outputs relative links (as opposed to absolute links)
- AtlasBarfed 5y agoI thought Zig didn't have unicode strings? https://www.reddit.com/r/Zig/comments/9q3or3/how_to_deal_with_strings_in_zig/ https://www.reddit.com/r/Zig/comments/9q3or3/how_to_deal_wit... If that's true, Zig is NOT a modern language. Modern languages use international strings, and are unicode aware with a good unicode aware string library. For crap's sake, the code example for comparing modern languages USES A STRING. The fact it is not unicode doesn't matter.
- anonymoushn 5y agoIs a 3-year-old reddit comments section really a better source for this than the Zig standard library? https://github.com/ziglang/zig/blob/master/lib/std/unicode.zig https://github.com/ziglang/zig/blob/master/lib/std/unicode.z...
- p0cc 5y ago
- dottedmag 5y agoSeriously, the biggest gripe about C is the design of standard library? Not the pervasive undefined behaviour and compilers that become more aggressive every release about breaking previously-working code? Not the reams of code that assume sizes of integers and signedness of char? Not the wild build process that makes it awfully hard to actually build anything that has any dependencies whatsoever. strtol. Damn, what a nuisance!
- dkersten 5y agoThe standard library is part of the language. Its also how the language is used, both because the language design encourages it and because people have a tendency to copy how the standard library does things, since you learn to do things how the libraries you use do them.
- hmry 5y agoWhere does the article say that those other things are not also big problems? In fact, it specifically says strtol is one example of the things wrong with C. This comment seems needlessly dismissive.
- wudangmonk 5y agoWhy complain about a functions though? you can just write your own that does exactly what you want it to do. I don't see how this is a flaw in the language. You could at least mention things like the loss of size information when passing arrays between scopes, that is an annoyance that can be considered as a problem in the language.
- dottedmag 5y agoRead the last section. None of the real problems of C are even mentioned.
- shadowofneptune 5y agoYou do not consider it a sign of C's problems in the modern world that so many of its core functions (atoi, atol, atoll, atof, gets, strcat, strcpy, sprintf, etc.) are unsafe and yet still are out there in production and teaching?
- kazinator 5y agoThis isn't the usual way this is coded: char *one = "one"; char *end; errno = 0; // remember errno? long i = strtol(one, &end, 10); if (errno != 0) { perror("Error parsing integer from string: "); } else if (i == 0 && end == one) { fprintf(stderr, "Error: invalid input: %s\n", one); } else if (i == 0 && *end != '\0') { f__kMeGently(with_a_chainsaw); } It's actually like this: errno = 0; long i = strtol(input, &end, 10); if (end == input) { // no digits were found } else if (*end != 0 && no_ignore_trailing_junk) { // unwanted trailing junk } else if ((i == LONG_MIN || i == LONG_MAX)) && errno != 0) { // overflow case } else { // good! } errno only needs to be checked in the LONG_MIN or LONG_MAX case. These cares are ambiguous: LONG_MIN and LONG_MAX are valid values of type long, and they are used for reporting an underflow or overflow. Therefore errno is reset to zero first. Otherwise what if errno contains a nonzero value, and LONG_MAX happens to be a valid, non-overflowing value out of the function? Anyway, you cannot get away from handling these cases no matter how you implement integer scanning; they are inherent to the problem. It's not strtol's fault that the string could be empty, or that it could have a valid number followed by junk. Overflows stem from the use of a fixed-width integer. But even if you use bignums, and parse them from a stream (e.g. network), you may need to set a cutoff: what if a malicious user feeds you an endless stream of digits? The bit with errno is a bit silly; given that the function's has enough parameters that it could have been dispensed with. We could write a function which is invoked exactly like strtoul, but which, in the overflow case, sets the *end pointer to NULL: // no assignment to errno before strtol int i = my_strtoul(input, &end, 10); if (end == 0) { // underflow or overflow, indicated by LONG_MIN or LONG_MAX value } else if (end == input) { // no digits were found } else if (*end != 0 && no_ignore_trailing_junk) { // unwanted trailing junk, but i is good } else { // no trailing junk, value in i } errno is a pig; under multiple threads, it has to access a thread local value. E.g #define errno (*__thread_specific_errno_location()) The designer of strtoul didn't do this likely because of the overriding requirement that the end pointer is advanced past whatever the function was able to recognize as a number, no matter what. This is lets the programmer write a tokenizer which can diagnose the overflow error, and then keep going with the next token.
- xigoi 5y agoSure, you can't get away from handling the cases, but as the article clearly demonstrates, there can be a much better interface for it.
- WalterBright 5y agoThe #1 problem with C is buffer overflows. The solution is pretty simple: https://www.digitalmars.com/articles/C-biggest-mistake.html https://www.digitalmars.com/articles/C-biggest-mistake.html and does not break existing code.
- thingsgoup 5y agoNeat idea. What does implementing new syntax in one of the established C compilers involve? Is it the kind of thing that could be reasonably tackled in a small patch just to play with?
- WalterBright 5y agoIt wouldn't be hard. The semantics are straightforward and don't interfere with the way a C compiler already works.
- bhaak 5y ago> C is probably the patriarch of the longest list of languages. Notable among these are C++, the D programming language, and most recently, Go. There are endless discussion threads on how to fix C, going back to the 80’s. Why is Java missing in that list?
- thingsgoup 5y agoI suspect it’s because manual memory management in Java isn’t built into the language. Is it even possible? I’m not a Java programmer and I don’t know. My understanding has been that the runtime doesn’t expose the memory model to you.
- kaba0 5y agoIn a way it does. Java just likes to push most features to methods on special objects, instead of exposing them as native functionality (to avoid backwards incompatible changes). So it would look something like MemorySegment.allocateNatice(100, someScope). This new API has a runtime ownership model, so by default only a single thread can access this memory address, and it can be freed at will.
- ErikCorry 5y agoIn practice it's _worse_ than that because you probably don't want a "long", you probably want a particular size like a 64 bit integer. So you have to add ifdefs to call either strtol or strtoll depending on the size of "long" and "long long". And if you are using base 16 then strtol will allow an optional "0x" prefix. So if you didn't want that you have to check for it manually. Strtol also accepts leading whitespace so if you didn't want that you have to test manually for it. Don't pass a zero base thinking it means base ten. This works almost all the time but misinterprets a leading zero to mean octal. Good luck!
- remorses 5y agoThis is crazy
- ngcc_hk 5y agoWelcome to the real world. At least it is understandable crazy. After all it is in a library call not part of the mental model of c. To compare, I can’t say about the mental baggage you carry with javascript the core language. Still, it sort is of works. World moved on. Good luck.
- draw_down 5y ago
- thesuperbigfrog 5y ago>> you probably don't want a "long", you probably want a particular size like a 64 bit integer. So you have to add ifdefs to call either strtol or strtoll depending on the size of "long" and "long long". stdint.h (https://en.cppreference.com/w/c/types/integer https://en.cppreference.com/w/c/types/integer) provides fixed-width integers in specific sizes. It became a standard in C99.
- ErikCorry 5y agoDoesn't provide strto32 and strto64 though so you still need an ifdef.
- futharkshill 5y agoIf a user wants to parse integers etc. from a string, the function snprintf and family is often applied. It is a neatly simple function. This article seems to invent a problem rather than an organic one.
- ogogmad 5y agoThe article argues that there is no easy way to detect whether the parsing finished successfully. As a consequence, the C standard library is unsafe when used normally. It's interesting how beginners are encouraged to use various string functions which are not safe to use with external input.
- creativemonkeys 5y agoOne of C's design principles is to be fast at the cost of safety, just like an F1 formula car. It will let you make fast mistakes. You drove a Corolla in college, then got a job and drove a cool BMW for several years and now you think you're hot shit, so you hope in an F1 car and not only does it take forever to learn how to drive it, it has to be driven on a special track and the gearbox is different, what a nuisance! "If only we could add 4 doors, automatic transmission, snow tires, and a trunk to put our stuff in, people won't keep getting into accidents with this car", you say. Right, but then it becomes a BMW. If you want real speed, you need to first go slow and master the car because otherwise you'll crash and burn. C is messy because real world hardware is very messy. You can't push bytes through the hardware at its speed limit without getting your hands dirty, and we all come out into the real world wearing "class Dog extends Animal" white gloves. To use C effectively, you should not be coding in C in your mind. You should be thinking in assembly, but your fingers should be typing C code. It's not safe, but if you want to reach 230MPH and accelerate at 60MPH in 2.6 seconds, you better know exactly what you're doing when you hop behind the wheel of that car. It's not for the weak.
- comex 5y agoThere is nothing about the design of strtol that makes it particularly fast. If anything, the extra checks and accesses to errno (which on modern systems is generally an implicit function call) that are required to use strtol correctly represent unnecessary overhead, though only a trivial amount of it. But mostly it’s just an awkward API design.
- adwn 5y agoThere's nothing fast about zero-terminated strings. In fact, many operations on them are much slower than sane alternatives, because they first have to scan the entire string to compute its length. You can't even create a temporary substring without either modifying or copying part of the original string. How lame is that? Zero-terminated strings are almost never the best solution, so why are they the language-supported default? > You should be thinking in assembly, [...] Well, then you shouldn't by typing in C, because Undefined Behavior coupled with modern C compilers will make sure that what you get is not what you thought. *cough* signed integer overflow * cough* > You can't push bytes through the hardware at its speed limit without getting your hands dirty Rust proves you wrong (maybe some other languages, too, but I don't know them as well)
- gengiskush 5y agoHow about leaving the old stuff you want to "replace" alone? People are using it.
- skywhopper 5y agoWhat a weird post. The examples from Rust and Zig don’t fail gracefully, so they can’t be considered complete. Panicking on bad user input is bad code, too. And the main complaint seems to be that the C stdlib could be improved. But where it has been improved, the author complains that it’s really just doing the ugly stuff under the hood. What does the author think the Rust stdlib function is doing exactly?
- prirun 5y agoPL/I, created in 1964, had strings. Real strings, where the compiler knows the length even when it gets passed around and is declared char(*) var in the receiving function. You can't have buffer overflows because the compiler and runtime know every string's current length and allocated length. This isn't a particularly hard problem. C just took a shitty shortcut to fake strings using byte arrays and the world glombed onto it. Now we're stuck with a crappy "standard" that people should have scoffed at when it first showed its ugly face.
- agumonkey 5y agoIt seems the world is a pile of shitty shortcuts.
- judahmeek 5y agoI don't know about the world, but our brains sure are! xD We rely on cognitive biases (shitty shortcuts) for as long as we can. In many cases, for longer than would be optimal.
- 41c2 5y ago
- gameswithgo 5y agoat the time the choice made sense for C, today it is a vanishingly rare case that you want a language that doesn’t know how long its own arrays and strings are, or that doesn’t know what might be null.
- grymoire1 5y agoPersonally, I'm tired of people bitching about C. At the time, the choice was C or assembly language for embedded/operating systems. There was no other choice in the 1970's. In fact, it wasn't even an option for most of the 1970's. If you worked at a company and wanted a team of people to develop on a multi-user system, and port it to a single-user stand-alone system, you were out of luck. Our company sold test equipment based on the Data General minicomputers, and while DG had multi-user systems and single-user systems, they had no common programing language besides FORTRAN. It was so frustrating. And then Digital came to us and wanted to buy a lot of systems, but it had to be running on a PDP-11. Trouble is, our test system was written in Data General assembly language. We had to re-write the system in a portable high-level language that could run on RSX-11 OS. But how? We searched for a suitable programming language we could buy support for, and ended up using PASCAL - which was a P-code interpreter. The P-Code was portable across operating systems. So I "ported" an assembly-based system to Pascal, and was able to have equivalent runtime performance, because the DEC system had RAM-based overlays and the DG had disk-based overlays. Otherwise, performance of Pascal over ASM would have made it unfeasible. A few years later, C was commercially available. Oh I wish it was a choice that was available then. The rule of thumb was that C would run with 90% of the performance of assembly language. And that was before they made incredible strides in compiler technology. PL/1 would have been a disaster, assuming it could run at all on a 16-bit machine.
- alkonaut 5y agoWait is this article saying that there is no good/obvious/standard function to parse a string to a number and has the two obvious outputs of such a function (the number, and a bool or error code)? Even a person in the 60s would realize that that’s the api for conversion from a string to a number (or any conversion that might fail)! What happened? Why do these functions even exist?
- gumby 5y ago> It exists because it became part of the POSIX standard way back when a pdp7 was an advanced computer… The PDP-7 was long obsolete by the time the POSIX effort started. By then the most common Unix host was a VAX (32 bits), though it, or Unix-alikes, ran on a variety of 16 and 32 bit machines, hence a desire for standardization.