5 ms·
There are multiple factors that contribute, but here's a few that I haven't (to the best of my recollection) seen mentioned so far: tooling (crucial), "do the s
by strlen 13y ago
There are multiple factors that contribute, but here's a few that I haven't (to the best of my recollection) seen mentioned so far: tooling (crucial), "do the simplest thing that could possibly work" attitude brought about by (lack of) a built-in collections library and simple syntax, and lack of rapidly changing requirements during development.
Tools for C are powerful, mature, and available on near every platform. Every single C programmer that I know uses multiple memory debuggers, for example: e.g., valgrind for leak checks and more on smaller code segments, LLVM address sanitizer on unit test runs, jemalloc or tcmalloc/google-perftools, and more.
"Do the simplest thing that could possibly work" is somewhat close to the author's point: simply put, while there are widely used cross-platform hash tables in C, it takes more effort to use one; hence, they aren't used e.g., when N is known to be 100 (or if maximum size is bound, there's likewise no pressure to use the built in hash table/red-black tree and one can use perfect hashing instead...). In another big distinction I see between code I wrote in C vs. higher-level languages is that one can simply use a stack allocated variable length array -- or (if random access doesn't happen) a linked list in place of a heap allocated resizable array.
These exammples, I think demonstrated how C discourages you from interpreting "the simplest thing that could possible work" as "the thing that is fastest to code"; that's not to say development speed is not an issue -- it's a _huge_ issue! -- but it often leads to code where the unneeded complexity (such as using a red-black tree or a skip list instead of a sorted stack allocated array!) is hidden from the programmer but is nonetheless present.
Onto the last point, I agree sentiment expressed by others questions the premise. I think the examples of C programmers where we explicitly know of something as being a C program are heavily biased towards well known and open source software; open source software, language runtimes, operating systems, browsers (usually in a clean and limited subset of C++ like, e.g., Chromium and likely Firefox as well -- I only mention Chromium as I've re-used base libraries from its codebase), and so on. We don't, however, think of parking meters, vending machines, and many other usually embedded (but not life critical) systems that -- unlike open source systems software -- are built from unclear, confused, and ever changing business requirements ("well, this city introduced a new red-zone, so now we have to make sure to charge extra 0.33 cents between hours 8 and 13 except on all Holidays, but not July 4th").
Yet when people bring up Java, it's easy to forget about gmail or Minecraft, but instead to think of the insurance claim management system that you had to use after your rental car got rear-ended near 4th and Harrison, that crashes with a user-visible stack trace when you click the wrong button and (even in 2013) doesn't use Ajax and requires opening a new page (slowly, as they don't use Java's built in thread-pools but spawn new threads, because the system has to run at a customer site that uses Enron's custom implementation of Java 1.3 on an AS/400) to make an changes.
- swah 13y agoThe other day I was quite surprised to learn that the first version of memcached was written in Perl - thinking something like that should be obviously written in C, or else could never make LAMP systems faster... :)
- strlen 13y agoKeep in mind the date it was written -- this was before SSDs and in the days of tiny heaps. Any improvement upon seek times of what were then commodity disks (SCSI/SAS being insanely expensive and out of bradfitz' reach, afaik) would easily dominate the cost of Perl's runtime (which, in retrospect, isn't horribly bad -- the GC and VM are more mature and less trouble-prone than most common scripting languages, although nowhere near LuaJit or any modern Scheme/Common Lisp implementation).
- drdaeman 13y agoWhy it's obvious an user-space daemon should be written C? It's a myth other runtimes don't have enough performance.
- strlen 13y agoIt's not about performance (nothing memcache does or more precisely -- given all new features I am likely unaware about -- when it was first rewritten in C is particularly CPU intensive), it's about memory management and efficiency: a in-memory system serves to reduce latency by using large amounts of memory; tuning the JVM to handle managed heaps larger than ~24gb (with only half to 2/3 of that usable for caching data) is doable but pretty much a black art if you want to avoid huge latency spikes caused by garbage collection pauses (which defeats the purpose of an in-memory store). Even if pauses are rare enough as not to affect 95-pctile latency (or if you don't care about 95-th pctile latency) in a highly-available system they introduce one of the more nasty distributed systems failures -- node A talks to node B and goes into a GC pause; node B thinks node A is dead, node A thinks node B is dead -- because when it didn't receive a reply to its outstanding requests until after it emerged from a GC pause triggering a timeout (all in the mean while certain figures from RDBMS community insist this didn't actually occur because "partitions never happen in a single datacenter...") While it's very much possible to directly allocate and manage memory yourself in Java -- and I have done that ( http://mail-archives.apache.org/mod_mbox/hbase-commits/201306.mbox/%3C20130613181824.1A81E23889ED@eris.apache.org%3E http://mail-archives.apache.org/mod_mbox/hbase-commits/20130... ) -- but without memory safety guarantees, without the great tooling, without the general idea that "if your code compiles, it will probably work correctly". In other words, you're programming a very verbose and ugly (see DirectByteBuffer interface) C. There's many cases where this approach ("off-heap memory") makes sense when building memory intensive apps in Java that clearly benefit everywhere else from being written in Java: JNI is unpleasant to work with, JNA adds additional performance penalties on top of those imposed by JNI calls that it makes under the cover -- which is fine in many cases, but remember that reading files or sockets are also JNI calls -- but less so when you're building a system that distinguishes itself by being in memory. On the other hand, if you use C (or "C+" a.k.a. "C with objects", or -- as I prefer to say to avoid confusion with libraries like glib or apr or the Linux vfs layer, all successful object oriented systems written in regular C -- "C with templates") you have accesses to the existing works (advanced allocators like jemalloc and tcmalloc don't work very well with JVM), excellent memory-debugging tools, and so on... I will say this: I think more software rather than less should be written in higher level languages. I absolutely love OCaml, Erlang, and Lisps (especially those -- like Typed Racket and Clojure -- that have started importing features from Haskell/ML family), like what I see in Go. I use Python (and formerly Perl) on a daily basis for "casual programming", to experiment with new ideas, and automation. I've written a great deal of software in Java, where many people blinked the idea that this category of software could be more than a toy in Java. I think garbage collection could be great improved and I see no theoretical reasons why, e.g., compilers can't be written in OCaml as opposed to C or C++ (I do see practical reasons having to do with runtime, lack of multi-core support but that's a separate and fixable issue). I find projects aimed at making systems programming in high-level languages fascinating (e.g., Mirage OS, Microsoft's experiments, Jikes RVM, etc...) However, even if these theoretical strides are achieved, there's always going to be room for C -- in the end, you need systems that give you great deal of control and act in a very deterministic and predictable fashion . In the mean while, however, there's also need to build practical systems -- so while C and C++ may not be ideal in theory, they are often the only realistic option in practice. [1] Note, however, I didn't say anything about performance -- OCaml, Java, Haskell, LuaJit etc.. do well in the Debian benchmarks; Lisps follow those languages closely. Likewise, many high-level languages can be AOT compiled. While when written by a strong programmer and/or assisted by today's excellent optimizing compilers C will still beat these fast high-level languages in most cases, it should be noted that today it's extremely difficult to write assembly code by hand that beats assembly code written by an optimizing C compiler or even the JVM. I predict that when a language comes about that has a better designed type system than C/C++ and yet still provides ability to control memory layout much as C does, eventually compilers for that language will emit assembly code that beats assembly emitted by C compilers for the same reason compiler-written assembly beats hand-written assembly -- compiler is able to leverage the type information (declared or -- in the case of dynamically typed languages with fast runtimes -- inferred) to aid optimization.