10 ms·
Redis crashes - a small rant about software reliability
- shin_lao 14y agoThis is an interesting post, especially the part about memory testing. We have a simple policy: ECC memory is required to run our software in production. Failure to do so voids the warranty.
- zdw 14y agoThis. For desktop computers, Intel charges a premium on any ECC-capable gear (their Xeon line), so it's really only available in workstation class computers. Most AMD gear (AM2/3/3+ sockets, not A-series) can take ECC RAM, if there is BIOS support. ECC RAM costs about 10-30% more per DIMM, but as memory is so incredibly cheap these days, its probably the cheapest safety net you can buy.
- barrkel 14y agoIt's a pain in the neck. By not supporting ECC RAM, Intel is IMO indirectly responsible for millions of dollars worth of lost work from crashes on consumer hardware in workplaces worldwide. ECC RAM should be standard, given modern memory capacities. The last two machines I built had bad modules that needed weeding out, and I follow anti-static precautions fairly carefully. I used to be a PC technician and I built probably over a hundred PCs in the 90s. Memory was never as fragile and fault-prone as it is these days.
- minimax 14y agoWhat if your customers want to run on EC2 instances?
- antirez 14y agoIt is covered in the blog post. (This is not a critique, just an hint, I understand that reading a very long blog post is time consuming).
- darklajid 14y agoNope? You seem to have misread (well, or it's me of course). You mention EC2 in your blog post, but he asked the person requiring ECC memory or voiding the product warrany what _they'd_ do if the customer wants to run on EC2. In fact, the GP probably used the EC2 part of your blog entry to come up with the question in the first place.
- minimax 14y agoI did read the whole post. It was very informative. I wasn't aware that EC2 did not have ECC RAM. My question was directed at shin_lao's policy about not providing a "warranty" for his customers running on non-ECC hardware.
- antirez 14y agoOh sorry I get it now...
- shin_lao 14y agoOn production servers using virtual machines to run our software is not advised. Nevertheless, we would do our best to please a customer looking to host our software on an EC2 cluster, with the appropriate warnings. ;) A bit of context: we sell a "real time" non-relational database (http://www.quasardb.net/ http://www.quasardb.net/). Our customers come to us for speed and reliability and therefore build dedicated farms to host our database.
- henrikschroder 14y agoWow, that product page is completely lacking any meaningful technical information about your product. :-D How do you stack up against the most common open source NoSQL systems? Redis, Cassandra, Mongo, Couchbase? Is your db eventually consistent, or partitioned, or replicated, or what?
- shin_lao 14y agoThanks for the feedback, this is currently a landing page we give to our customers we meet face to face. We're working on something more consistent to answer questions like yours. quasardb is a key/value store. It is (a lot) faster in a multi-client context that the engines you listed and can handle entries of any size (provided you have enough space on the servers, of course!). It's fully symmetric which means the load is equally distributed and replicated on all the nodes (no master node). If you have more question feel free to mail us (don't want to highjack this thread).
- politician 14y agoDo you have a blog? Maybe you could do a write-up.. This is the kind of geek catnip that HN likes.
- shin_lao 14y agoWe do have a blog - the subject is vast. Do you have anything in particular you would like to read about?
- raverbashing 14y agoDoes Amazon or other provider offers this warranty, I mean, ECC servers (or something similar?)
- lucian1900 14y agoPerhaps using safer languages (and languages with better error reporting) would be a solution to these kinds of problems.
- meaty 14y agoNot really. We get all the same sorts of errors in our very high level C#/Asp.Net/VMware deployments and it's a shit load harder to debug with all the extra baggage that a VM and hypervisor throw on top as well... A better solution to all the reliability problems is better quality hardware i.e. not X86. X86 has very few reliability features built in past ECC. If you look at UltraSparc based machines, they can predict failures and offline chunks of the hardware (CPUs, RAM regions, IO devices) so they can be replaced without disrupting the system. Prevention is better than debugging :)
- lucian1900 14y agoI was thinking more of something like Rust, where only a small subset of your code would be poking at memory manually. Then when you do get a mysterious crash, you only need to look at the unsafe portions of your application (or perhaps at the compiler).
- shin_lao 14y agoRust isn't production ready.
- lucian1900 14y agoI never claimed it was. I'm just thinking that perhaps the approach it (and Haskell, for that matter) takes is better.
- shin_lao 14y agoI think you're falling in to the "silver bullet" trap. http://en.wikipedia.org/wiki/No_Silver_Bullet http://en.wikipedia.org/wiki/No_Silver_Bullet Basically, making reliable software is hard. Changing the language doesn't bring anything. There are a lot of tools to make sure your C/C++ programs doesn't have obvious errors. The problem are non-obvious errors, and these errors exist in all the languages, with different forms. Another way to put it: "You cannot reduce risk, you can only replace it with another".
- ComputerGuru 14y agoPage is down. Here is a formatted copy: https://gist.github.com/4154289 https://gist.github.com/4154289
- nateberkopec 14y agoThe irony is too thick to read through.
- jrajav 14y agoFormatted more nicely: http://gist.io/4154289 http://gist.io/4154289 By the way, check the page first, give bloggers the traffic their content deserves!
- JimWestergren 14y agoIs he not caching his blog with Redis?
- antirez 14y agoThe blog uses Redis as primary data store, unfortunately there is Ruby between the user and the DB ;-) Btw here the problem was mine, I was running the Sinatra app wit "ruby app.rb", and Apache was mod_proxing to this running on port 4567. By default mod proxy will suspend the connection 60 seconds with an error if the proxed thing returns something wrong. Idiotic default that can be avoided just with: ProxyPass / http://127.0.0.1:4567/ retry=0 See "retry=0".
- antirez 14y agoSorry, the Sinatra based site is deployed with "ruby app.rb". Probably not enough...
- akx 14y agoMay I recommend uWSGI for hosting? http://uwsgi-docs.readthedocs.org/en/latest/Ruby.html#running-rack-applications-on-uwsgi http://uwsgi-docs.readthedocs.org/en/latest/Ruby.html#runnin...
- 14y ago
- codeflo 14y agoIn theory, there's nothing stopping the OS from remapping the pages of your address space to different physical RAM locations at any point during your test. So even if you have a reproducible bit error that caused the crash, there's a chance that the defect memory region is not actually touched during the memory test. Now, this may not be such a huge problem in practice because the OS is unlikely to move pages around unless it's forced to swap. But that depends on details of the OS paging algorithm and your server load.
- jgrahamc 14y agoHis point about logging registers and stack is interesting. Many years ago I worked on some software that ran on Windows NT 4.0 and we had a weird crash from a customer who sent in a screen shot of a GPF like this: http://pisoft.ru/verstak/insider/cwfgpf1.gif http://pisoft.ru/verstak/insider/cwfgpf1.gif From it I was able to figure out what was wrong with the C++ program. Notice that the GPF lists the instructions at CS:EIP (the instruction pointer of the running program) and so it was possible by generating assembler output from the C++ program to identify the function/method being executed. From the registers it was possible to identify that one of the parameters was a null pointer (something like ECX being 00000000) and from that information work back up the code to figure out under what conditions that pointer could be null. Just from that screenshot the bug was identified and fixed.
- danielweber 14y agoI remember generating .map files as part of the build process that were invaluable in figuring out where Windows desktop programs were crashing. It was about 30 minutes of work that made 3-hour debugging sessions into 10 minute debugging sessions from then on.
- malkia 14y agoUnfortunately address space randomization techniques make this much harder.
- geal 14y agoNot necessarily. The screenshot indicates the bytes pointed by the IP, so it would still be possible to find them a binary you just built, and debug it from there.
- malkia 14y agoonly if this is nota relocated code, and still there can be much code duplication, especially with C++ templates/inlines
- apaprocki 14y agoCan't agree with this more.. And he is just talking about logging crashes. One of the best debugging tools you have at your disposal in a large system (a lot of programmers contributing code -- bugs can be anywhere) is logging the same stack information in a quick fashion under normal operation in strange circumstances so as not to slow down the production software. The slowest part of printing that information out is the symbol resolution in the binary of the stack addresses to symbol names. This part of the debugging output can be done "offline" in a helper viewer binary and does not need to be done in the critical path. We frequently output stack traces as strings of hex addresses detectable by a regex appended to a log message. The log viewer transforms this back into an actual symbolic stack trace at viewing time to avoid the hit of resolving all the symbols in the hot path.
- deleted 14y ago[deleted]
- grundprinzip 14y agoI totally like this post, because main-memory based software systems will become the future for all kinds of applications. Thus, handling errors on this side will become more important as well. Here are my additional two cents: At least on X86 systems, to check small memory regions without effects on the CPU cache can be implemented using non-temporal writes that will directly force the CPU to write the memory back to memory. The instruction required for this is called movntdq and is generated by the SSE2 intrinsic _mm_stream_si128().
- jimwhitson 14y agoAt IBM, we were very keen on what we called 'FFDC' - 'first- failure data capture'. This meant having enough layers of error-detection, ideally all the way down to the metal, so that failures could be detected cleanly and logged before (possibly) going down, allowing our devs to reproduce and fix customer bugs. Naturally it wasn't perfect, and it depending on lots of very tedious planning meetings, but on the stuff I worked with (storage devices mainly) it was remarkably effective. In my experience in more 'agile' firms - startups, web dev shops and so on - it would be very hard to make a scheme like this work well, because of all the grinding bureaucracy, fiddly spec-matching and endless manual testing required, as well as the importance of controlling - and deeply understanding - the whole stack. Nonetheless, for infrastructure projects like Redis, I can see value in having engineering effort put explicitly into making 'prettier crashes'.
- ricardobeat 14y agospec-matching is a specialty of good agile companies, but web-dev shops don't usually write their own db/server software.
- erichocean 14y agoAlthough we use ECC in our servers already, I've recently been experimenting with hashing object contents in memory using a CityHash variant. The hash is checked when the object moves on chip (into cache), and re-computed before the object is stored back into RAM when it's been updated. Although our production code is written in C, I'm not particularly worried about detecting wild writes, because we use pointer checking algorithms to detect/prevent them in the compiler. (Of course, that could be buggy too...) What I'm trying to catch are wild writes from other devices that have access to RAM. Anyway, this is far from production code so far, but hashing has already been very successful at keeping data structures on disk consistent (a la ZFS, git), so applying the same approach to memory seems like the next step. The speed hit is surprisingly low, 10-20%, and when you put it that way, it's like running your software on a 6 month old computer. So much of the safety stuff we refuse to do "for performance" would be like running on top-of-the-line hardware three years ago, but safely. That seems like a worthwhile trade to me... P.s. Are people really not burning in their server hardware with memtest86? We run it for 7 days on all new hardware, and I figured that was pretty standard...
- aidenn0 14y ago1) Yes, lots of people don't run memtest86 at all. 2) Even those that do run it typically run it for no more than 24 hours 3) Many people don't build their own hardware these days, its a VPS or EC2 4) If you've selected ECC RAM then you know way more about memory failures than >99% of Redis users
- dap 14y agoGreat post, showing admirable dedication to software reliability and a solid understanding of memory issues. One of the suggestions was that the kernel could do more. Solaris-based systems (illumos, SmartOS, OmniOS, etc.) do detect both correctable and uncorrectable memory issues. Errors may still cause a process to crash, but they also raise faults to notify system administrators what's happened. You don't have to guess whether you experienced a DIMM failure. After such errors, the OS then removes faulty pages from service. Of course, none of this has any performance impact until an error occurs, and then the impact is pretty minimal. There's a fuller explanation here: https://blogs.oracle.com/relling/entry/analysis_of_memory_page_retirement https://blogs.oracle.com/relling/entry/analysis_of_memory_pa...
- antirez 14y agoThank you for the interesting link dap.
- jacquesm 14y agoI take it you know about /var/log/mcelog ?
- ComputerGuru 14y agoI don't think enough people appreciate just how awesome of an OS Solaris was. I never had opportunity to deploy it full-scale for any projects, but I lamented the loss of great potential when it "died."
- dap 14y agoIt didn't die. It was forked by the community when Oracle close-sourced it. The community fork (called illumos) is being actively developed by multiple companies, which have done significant new feature work (e.g., http://dtrace.org/blogs/wdp/2011/03/our-zfs-io-throttle/ http://dtrace.org/blogs/wdp/2011/03/our-zfs-io-throttle/).
- ArbitraryLimits 14y agoThe first and only time I used Solaris, I tried to run our application and got the error "System out of colors" or some such. Swore then and there never to use it again if I could help it.
- nicpottier 14y agoThis kind of attention to detail is all too rare these days. I love Redis, because I have never, not once, ever had to wonder whether it was doing its job. It is like a constant, always running, always doing a good job and getting out of the way. It only does a few things, but it does them exceedingly well. Just like nginx, I know it will be fast and reliable, and it is this kind of crazed attention to detail that gets it there.
- pnathan 14y agothere is an approach to hard real time software where antirez's idea for a memory checker is done.
- js2 14y agoIt's crazy that an application should have to test memory. It should simply be handled by the HW and OS. e.g. Some details about how Sun/Solaris deal with memory errors: http://learningsolaris.com/docs/DRAM_errors.pdf http://learningsolaris.com/docs/DRAM_errors.pdf Note the section on DRAM scrubbing, which I was reminded of from the original article's suggestion on having the kernel scan for memory errors. (I remember when Sun implemented scrubbing, I believe in response to a manufacturing issue that compromised the reliability of some DIMMs.)
- CrLf 14y agoI find this idea of a lack of ECC memory on servers disturbing... This is the default on almost all rack mountable servers from the likes of HP or IBM. Of course, people use all kinds of sub-standard hardware for "servers" on the cheap, and they get what they pay for. I haven't seen a server without ECC memory for years. I don't even consider running anything in production without ECC memory, let alone VM hypervisors. I find it pretty hard to believe that EC2 instances run on non-ECC memory hosts, risking serious data loss for their clients. Memory errors can be catastrophic. Just imagine a single bit flip in some in-memory filesystem data structure: the OS just happily goes on corrupting your files, assuming everything's OK, until you notice it and half your data is already lost. Been there (on a development box, but nevertheless).
- antirez 14y agoI hope there is a way to get some official statement from Amazon, Linode, and other very used VM providers about the kind of memory used in their servers. This would help users understanding the real risks.
- CrLf 14y agoI think they don't mention it because they think it to be obvious (I hope). However, with all the special built servers that big providers use to reduce costs, there is some margin to doubt. I think Google may be able to get away with it. With enough checksums along the way, memory (and other hardware) errors can be detected in software pretty easily if you have independent machines checking the data and can afford the processing penalty. Now, for virtualization I seriously doubt it. Not unless their instances run simultaneously on more than one machine to check for inconsistencies between them (something that the mainframes do since the dawn of time, but that I don't see as feasible in a distributed environment).
- jcrites 14y agoA single machine is never going to be completely reliable. At any time it can halt for a variety of reasons: power loss, hardware failure, disaster in the data center like flooding, etc. Thus, a configuration that relies on the availability of a single machine is already risking serious outage or data loss by not being machine-redundant. Reliable systems require the coordination of many machines (at least two), and the replication of data across them if data's involved. It is useful to have component-level redundancy (e.g., RAID or ECC memory), but in some environments it may be cheaper overall to have machine-level redundancy using inexpensive machines. It also only takes the failure of a single critical subsystem for a machine to suffer an outage. You might have ECC memory and RAID, but do you have only a single Ethernet card and power supply? Single machine availability is a "weakest link" phenomenon from its components. I acknowledge that building software to run across a fleet of machines is more difficult than software that runs on only a single machine, but (1) the software development cost is largely a fixed cost, not a variable cost in the number of machines (2) building a distributed system is sometimes needed for scaling reasons anyway. If you scale a single machine vertically (i.e., get a bigger box), its cost rises faster than its capabilities; so an efficient high-scale system typically also means running a fleet of cheap machines (scale horizontally). I think these effects contribute to the rise of commodity-server computing, and cost is a reason not to consider it disturbing. In other words, crunch the numbers and see when it makes sense :-)
- chewxy 14y agoAnd people wonder why I recommend redis. Having run redis for over 1.5 years on production systems as a heavy cache, a named queue and memoization tool (on the same machine), redis has never once failed me. It's clear with antirez's blog post, his attention to detail. This post is fantastic.
- BoredAstronaut 14y agoThis post reminded me of my time as a consulting systems support specialist. Lots of weird problem turned out to be bad hardware. Usually memory or disk, sometimes bad logic boards. For end users, this would often lead to complete freezing of the computer, so it was less likely to be blamed on broken software, but there were still many times it was hard to be sure. Desktop OS software can flake out in strange ways due to memory problems. I used to run a lot of memory tests as a matter of course. I think the title of the article could be more accurate, considering how much is devoted not to issues about software reliability per se, but to distinguishing between unreliable software and unreliable hardware. I think an implicit assumption in most discussions about software reliability is that the hardware has been verified. I personally do not think that it is the responsibility of a database to perform diagnostics on its host system, although I can sympathize with the pragmatic requirement. When I am determining the cause of a software failure or crash, the very first thing I always want to know is: is the problem reproducible? If not, the bug report is automatically classified as suspect. It's usually not feasible to investigate a failure that only happened once and cannot be reproduced. Ideally, the problem can be reproduced on two different machines. What we're always looking for when investigating a bug are ways to increase our confidence that we know the situation (or class of situation) in which the bug arises. And one way to do this is to eliminate as many variables as possible. As a support specialists trying to solve a faulty computer or program, I followed the same course: isolate the cause by a process of elimination. When everything else has been eliminated, whatever you are left with is the cause. I'm still all jonesed up for a good discussion about software reliability. antirez raised interesting questions about how to define software that is working properly or not. While I'm all for testing, there are ways to design and architect software that makes it more or less amenable to testing. Or more specifically, to make it easier or harder to provide full coverage. I've always been intrigued by the idea that the most reliable software programs are usually compilers. I believe that is because computer languages are amongst the most carefully specified kind of program input. Whereas so many computer programs accept very poorly specified kinds of input, like user interface actions mixed with text and network traffic, which is at higher risk of having ambiguous elements. (For all their complexity, compilers have it easier in some regards: they have a very specific job to do, and they only run briefly in batch operations, producing a single output from a single input. Any data mutations originate from within the compiler itself, not from the inputs they are processing.) In any case, I believe that the key to reliable programs depends upon the a complete and unambiguous definition of any and all data types used by those programs, as well as complete and unambiguous definitions of the legitimate mutations that can be made to those data types. If we can guarantee that only valid data is provided to an operation, and guarantee that each such operation produces only legitimate data, then we reduce the chances of corrupting our data. (Transactional memory is such an awesome thing. I only wish it was available in C family languages.) One of my crazy ideas is that all programs should have a "pure" kernel with a single interface, either a text or binary language interface, and this kernel is the only part that can access user data. Any other tool has to be built on top of this. So this would include any application built with a database back-end. I suppose that a lot of Hacker News readers, being web developers, already work on products featuring such partitioning. But for desktop software developers who work with their own in-memory data structures and their own disk file formats, it's not so common or self-evident. Then again, even programs that do rely on a dedicated external data store also keep a lot of other kinds of data around, which may not be true user data, but can still be corrupted and cause either crashes or program misbehaviour. In any case, I suspect that this is going to be an inevitable side-effect of various security initiatives for desktop software, like Apple's XPC. The same techniques used to partition different parts of a program to restrict their access to different resources often lead to also partitioning operations on different kinds of data, including transient representations in the user interface. Can a program like Redis be further decomposed into layers to handle tasks focussed on different kinds of data to achieve even better operational isolation, and thereby make it easier to find and fix bugs?
- tylerneylon 14y agoThe memory check algorithm is a nice solution of the challenges he presents - easy to understand and effective. Here is a variation which, unless I'm missing something, would be a little simpler still and require less full-memory loops: 1. Count #1's in memory (possibly mod N to avoid overflow). 2. Invert memory. 3. Count #0's in memory. 4 Invert memory. I think this would catch the same errors (stuck-as-0 or stuck-as-1 bits). One difficulty is that multiple errors could cancel each other out, at which point you can do things like add checkpoints in the aggregation, or track more signals such as number of 01's vs number of 10's. In the end, this is like an inversion-friendly CRC.