7 ms·
The famous C dilemma: we want to be as close to the machine as possible, but don't want to change anything when the machine changes
by usrnm 2mo ago
The famous C dilemma: we want to be as close to the machine as possible, but don't want to change anything when the machine changes
- tonyhart7 2mo agoso what they gonna do ??
- pjmlp 2mo agoBecause contrary to urban myths, C is a normal high level language like everything else. The Assembly like abilities have been growing as language extensions in specific compilers, not as part of ISO C. Going back to K&R C, inline Assembly or intrisics were not even available, all of that required using the Assembler directly.
- uecker 2mo agoNot every high-level language gives you byte-level access to the representation of memory objects. But it is also wrong to reduce a language to what is in the spec.
- pjmlp 2mo agoMany do, contrary to what many C advocates talk about. Apparently reducing the language to what is in the spec is only a thing when talking about C and to some extent C++. When other languages have compiler specific extensions beyond the spec, it is a failure in their design. Yet when C and C++ devs have to reach out to compiler specific extensions, it is not a design failure like it is pointed out to others, rather an advantage. It is also wrong to not apply the same measure when it doesn't suit the message.
- uecker 2mo agoLot of stawman arguments.
- embedding-shape 2mo agoWouldn't be a authentic pjmlp comment unless they shit on C/C++ and/or praise Java/.NET with a bunch of straw-men :)
- pjmlp 2mo agoHow wrong you are, C++ isn't in the same league as C, Microsoft was right not wanting to keep updating their C support. It was already outdated by the time Borland released Turbo C++ 1.0 for MS-DOS, and only got new wind thanks to GNU FOSS and their manifest to prefer C as the main compiled language for GNU projects. Everywhere else outside UNIX, was going with a mix of C++ for OS frameworks, Apple, Microsoft, IBM, Be, Nokia, Epoch,.... Naturally given the option, between C, C++ and something else I might prefer that something else, however I managed a few interesting positions exactly due to my C++ skills, and interests. So don't mix my preferences for C and C++ on the same basket.
- embedding-shape 2mo agoWouldn't trade it for anything <3 Enjoy your Tuesday mate :) > So don't mix my preferences for C and C++ on the same basket. That mistake is mine indeed, I'll remember. Thanks, and I hope "no harm meant" was implicit :)
- gchamonlive 2mo ago> Wouldn't be a authentic pjmlp comment unless... > and I hope "no harm meant" was implicit :) Ad hominen then an apology, mixed signals here or I'm missing something. Maybe sarcasm?
- throwlifeaway 2mo ago> Ad hominen then an apology, mixed signals here or I'm missing something. Maybe sarcasm? You can express annoyance at someone's pattern of behavior without it being personal. embedding-shape isn't the only person annoyed by pjmlp's repeated disdain and snark towards people who use C (or Zig or WebAssembly or Rust or...).
- sylware 2mo ago"When other languages have compiler specific extensions beyond the spec, it is a failure in their design." This one of the failures of Linus T. with the linux kernel: he was not able to keep the assembly source code with plain and simple C code you can compile with a small and alternative C compiler (same failure for the glibc devs I think). I don't blame him, he is already keeping the linux ABI stable, and pulling that off is something.
- jstimpfle 2mo agoMany other languages only have one compiler available to start with. Each additional compiler supported by a project means variance in functionality and thus additional work for the project. That work could make the codebase more robust. Or it could be a ton of useless work. Or anything in between. Depends on the context of the project.
- Dylan16807 2mo ago> assembly source code with plain and simple C code you can compile with a small and alternative C compiler Which part of that big clause is the part that failed? Because I thought you could still compile Linux with TCC.
- sylware 2mo agoAs far as I know, that was for x86(32bits) linux, that decades ago. With those assembly source files (which do not abuse any pre-processor) and plain and simple C, I could build a modern x86_64 linux kernel with cproc/qbe (which gets 70% of gcc speed in my CPU intensive benchmarks... for a few % of gcc code and in plain and simple C, not brain damaged c++). But I kind of don't mind since the future is assembly coding on non-IP-locked standard like RISC-V, and the main issue for that future is the abuse of pre-processors (ffmpeg was bitten by it) or code generators which would not be written in assembly themselves (or with a simple high level language with an assembly written interpreter, asmpython?).
- Dylan16807 2mo ago
- jstimpfle 2mo ago_You_ do that. All the time. And then you fight these strawmans.
- pjmlp 2mo agoAnd you reply to that all the time with the C bias as well, oh well.
- groundzeros2015 2mo agoYou comment this almost everyone it comes up. The hardware also isn’t x86! That’s an abstraction too. The point is in C you have greater control of execution and resources, not that it matches the hardware exactly. It’s a spectrum and C is closer on that spectrum than JavaScript.
- pjmlp 2mo agoBecause just like your comment proves the point, many think only C can do this. So I keep re-educating folks that isn't the case. Every thread has different people reading it, so there is always a first time for many of them.
- jstimpfle 2mo agoYou are again misreading even the most clearly put statement. Compared to e.g. Javascript, C is "closer" to the hardware, gives you "more control" of it. It would be completely ridiculous to deny this fact. And if you move to e.g. C# / Java or similar, if you squint, and you try to be a smart-arse, then you could deny that C is closer to the hardware than C#, because C# probably has everything you need to control it, to the same degree that C allows you to. But if you work in these languages for a while, and look at the code that you ended up producing, then again you will absolutely find that it would be ridiculous to not admit that C gives you better control. And you could even extend this to Rust, because the language encourages you to use high-level prefabricated components. It discourages you from doing low-level things, at least a little bit I think (I'm not a Rust user). I think what you are doing all the time, is you are being a smart-arse, nothing else. What interesting low-level performant things have you actually programmed lately?
- pjmlp 2mo agoSmart-arse is comparing C versus JavaScript, instead of C vs C++, for example. And then coming with such lengthy ad hominem. Let make a fun exercise for the audience, given your performance remark. Paste a random C code that I should replicate in whatever language I feel like. There is one rule. If the sample code is pure ISO C, then I will only use what is in the standard of whatever language I pick up. If the sample code makes use of single language extension not part of ISO C, then I will have the freedom to also pick whatever language extensions I feel like.
- FooBarWidget 2mo ago> Not every high-level language gives you byte-level access to the representation of memory objects. Any code that ventures anywhere near that territory is 99% Undefined Behavior. It's almost impossible to write proper C/C++ code that isn't UB while touching byte-level representations.
- uecker 2mo agoThis is certainly not true. Accessing bytes of objects is well-defined in C.
- FooBarWidget 2mo agoJust look at this: https://blog.habets.se/2026/05/Everything-in-C-is-undefined-behavior.html https://blog.habets.se/2026/05/Everything-in-C-is-undefined-... This is undefined behavior! const int* magic_intp = (const int*)bytes; Heck even something trivial like this is UB: bool bar(char ch) { return isxdigit(ch); } The only safe thing to do is memcpy, but that's super useless. As soon as you try to interpret or manipulate the byte-level data in any way, there are UB traps everywhere you go.
- uecker 2mo agoYes, using an arbitrary type that is different from the one of the object is UB. But any access of a representation byte using a character pointer is well defined, not just memcpy and I would also not call memcpy useless.
- deleted 2mo ago[deleted]
- titzer 2mo agoIt's not normal in any other language that constant-folding in the compiler has different behavior than running an expression on the machine. C exists in a nether world of being neither assembly nor high-level language. People only call it high level because in the 1970s, having blocks, loops, and functions was high level, compared to the SoTa machines available in the day, which were either programmed with assembler or some bespoke thing the manufacturer came up with.
- thin_carapace 2mo agomay I ask specifically what aspects of C lead to examples of quasi low level status such as the one you gave? I'd hazard a guess the C abstract machine is defined in a particular manner differentiable from say the JVM?
- titzer 2mo agoIt has pointers and pointer arithmetic. Pointing into the stack, allocating buffers on the stack (and the resulting decades of stack smashing attacks that came with it). Most languages don't have an underlying model of a flat memory that you can just randomly point at and write things; they have objects and data types and functions that aren't meant to be pointed at (and usually cannot).
- thin_carapace 2mo agoah so other languages emulate harvard to a degree. I wonder if performance could improve by reimplementing a C like language based on a von neumann abstract machine inspired by something other than a PDP. thanks for your response
- Dylan16807 2mo agoIt's a similar distinction, but not really? Harvard is about separating code and data and C still does that with high reliability. Pointers and the stack and the important return addresses on the stack are all data. Strong barriers between different pieces of data are a different concern.
- bigfishrunning 2mo agoI would argue that inline assembly, while not in the standard, is really just a convenience feature -- it is part of the standard to declare an extern reference to a function in the symbol table and jump to it, it just requires a separate ASM object to link alongside your C object. Inline assembly doesn't allow anything you can't do without it.
- convolvatron 2mo agothat's kind of not true. inline asm lets me refer to the register that the compiler placed a value in. lets say I really want to use popcnt in my inner loop. with inline assembly I can just shove it in there. external linkage forces a function call overhead that can't be inlined, which obviates any benefit I might have had from using the specialized instruction.
- treyd 2mo agoYou'd probably also be able to use an intrinsic to avoid the pain of inline asm and make it portable.
- jonhohle 2mo agoIntrinsics usually come after new instructions have existed long enough for the higher level pattern across several architectures to establish a common pattern. If specific hardware is being targeted you may need to use those instructions before intrinsics exist.
- convolvatron 2mo agointrinsics are nicer in every way, assuming they exist. but some instructions are inherently non-portable. performance instruction like my popcnt example are good candidates since they can be implemented at varying costs on other architectures. but for systems programming there are control register and mode switch instructions that really aren't. some of those can be put into separate asm routines, but there are some that are poorly suited. segment long jumps on x86 are maybe an example. rdtsc is another one potentially. its also true that when I unwrap my new spin with fancy new instructions its unlikely to have a robust set of instrinsics around them. inline asm is a real mess, I always regret tussling with it, but its kind of pragmatically necessary if you're actually working at the metal in a high performance or embedded context unless you're doing the whole thing in assembly.
- MisterTea 2mo ago[dead]
- jjav 2mo ago> Because contrary to urban myths, C is a normal high level language like everything else. It's a myth that this is a myth. Have you ever worked with newbies learning C? These students split pretty hard into two camps (of course there are oddball exceptions): those who knew assembly and find C easy, and those who didn't and struggle with pointers until they finally get it (some, never do). Of the mainstream languages, C is the only one where understanding and dealing with direct memory access is a fundamental requirement if you're going to get anything done. The myth argument is that C code doesn't directly translate into the execution flow on the CPU. No, of course it doesn't. That's not the point. What people mean when they say C is lower level than most mainstream languages is because it forces you to deal with details most other languages paper over. Yes, multiple languages have some way of achieving this kind of memory access, but except from C, it is considered an esoteric edge case that mostly nobody needs.
- IsTom 2mo agoI might have learned all this too long ago and lost touch with how it's like to learn this, but: how is learning about using pointers in C different from using object references and array indices in python/js/java?
- mannschott 2mo agoIn C a pointer can point to any arbitrary offset within an object or array, to a local variable somewhere on the stack, to a static variable, to read-only memory, to the first(?) of an indeterminate number of characters (hopefully) terminated by ((char)0). To the first of a statically unknown number of variables located somewhere in memory, to a variable of indeterminate type (thanks to union), to I/O devices, and also to unallocated memory since there's no bounds checking. A reference can only point to an object or an array on the heap. Array indexing is bounds-checked. It seems obvious to me that these restrictions make references easier to reason about than C-style pointers, and therefore easier to learn.
- IsTom 2mo agoI imagine most of these examples (besides strings) isn't something you learn in the beginning and when e.g. doing operations on linked lists you're pointing to either allocated structs or null.
- wren6991 2mo agoExcept for the basic integer types. Change those as much as possible. Hell, CHAR_BIT=12 just to keep them on their toes. Personal pet theory: C is portable as in "you can retarget the compiler to any machine" moreso than "your code will run on any machine".
- trashb 2mo ago> Personal pet theory This is actually how c grew up. This is also one of the reasons why the spec is quite ambiguous in certain locations. C is made to be easily portable not a universal codebase for all platforms (though you can get quite close with some tricks like macros). Remember the spec allows C to run on a Unisys 1100/2200 just as well as on a pdp-11. I may be to embedded for this but if you want your code to handle long long as int64_t use <stdint.h>. I am of the opinion that you should always use fixed width types as portable types are a huge footgun and kind off redundant. Especially when you start doing a little more complex things expecting them to work exactly the same, like 128bit values on a 64bit platform.
- imtringued 2mo agoint was supposed to be the signed integer version of size_t. Meaning that it conforms to the native word size of the machine. But then a lot of software assuming that int means 32 bit got written and even 64 bit ABIs have kept int as 32 bits.
- wren6991 2mo agoI imagine there would also be some cache pressure increase if you inflated all of those ints to 64 bits.
- flohofwoe 2mo agoIt really hasn't much to do with C (e.g. there is no such thing as a "C ABI", and especially no such thing as a "standardized C ABI" - not sure if that's even a hot-take anymore). ABIs are defined by CPU and operating system vendors. Those ABIs usually happen to be quite 'C friendly', but that's not a requirement (for instance the AmigaOS ABI was primarily meant to be used from handwritten assembly code, and Amiga C compilers had to adapt to those ABI rules or they wouldn't be able to call into the operating system DLLs).
- sylware 2mo agoYep, an ABI (function call convention) is computer language agnostic. It is a binary specification. And in real life, only a subset of it is actually used. If they want to find something really obsolete, they better have a look at executable/dynamic lib file formats (PE+, ELF64). In other words, they better look at that first: I am using my own, which is beyond simple (a little RFC would suffice), no loader of any kind, basically userland syscalls. And I do embbed exes in an ELF64 capsule to run them transparently on linux systems (writting a internal linux exe loader would be copying ELF loading code while trashing 90% of its code). (hopefully in some not too far future, I'll try to build a mesa AMD vulkan driver for this very simple format and for that the main issue is.. c++ with its runtime, as always...).
- leoc 2mo agoNo-one cares anymore about ancient computers or weird specialist systems of course, but don't neglect my important requirements by breaking compatibility with any the platforms I am relying on at any point in time. We also need all of the most aggressive optimisations that compiler writers can come up with—this is high-performance code, after all!—but we certainly don't have time to deal with any breaking changes that would force revisions to our big important codebases. Make sure we can realise significant performance gains with just a simple recompilation. But remember to keep everything straightforward and close to the machine: we really hate all that weird UB which it's so easy to trigger by making an obvious, reasonable assumption which turns out to be wrong for some inexplicable reason.