12 ms·
Random bit-flip invalidates certificate transparency log – again?
- nickf 3y agoThis happened before a couple of years ago, and the problem seems to have repeated itself. https://news.ycombinator.com/item?id=27728287 https://news.ycombinator.com/item?id=27728287
- topspin 3y agoIs coincidence plausible here? Anyhow, Ian Fleming's thoughts on such things seem applicable: "Once is happenstance. Twice is coincidence. Three times is enemy action"
- jeroenhd 3y agoBitflips just happen. There's a good chance the device and connection you're using right now has had several bit flips in (unused) memory and you'll never even know. Everything is fine until these flips happen in critical code or data paths. ZFS corruption is a famous example, but there's also DNS corruption that happens quite regularly (and has been demonstrated to be usable for malicious purposes). I'm a little surprised the certificate transparency protocol doesn't validate the incoming data well enough to detect these bit flips, but on the other hand most software I've seen just assumes the bytes in memory and the bytes received through the network are all what they're supposed to be.
- theamk 3y agoeh, I bet you are wrong. I've run memtest many times, often for days, and it never detected any bitflips in RAM. And memtest is specifically designed to exercise every bit of memory and detect bitflips. While I've seen my share of PCs with memory errors, when you see one you replace memory / tweak settings, and then run memtest for a day or so to ensure it does not reappear. Also, "unused memory" should not be a thing in modern PC: the read cache should expand to fill it - it's super fast to evict and provides tangible benefits in case of hit. The errors usually occur on the boundaries - like in the network or in the SATA connection to disk or even in USB bus. For example the original "ZFS corruption" story (which seems t be gone from regular web, but I think its [0]) pretty clearly mentions damage "on the way to disk". [0] https://web.archive.org/web/20091212132248/http://blogs.sun.com/elowe/entry/zfs_saves_the_day_ta https://web.archive.org/web/20091212132248/http://blogs.sun....
- sliken 3y agoNot sure what it is about memtest86, but in my experience it doesn't find most memory errors. I've had serious memory issues, that trigger memlogd, dmesg, or similar about EDAC/ECC errors. I try memtest86 over night, no errors. Restart the node and get more errors. This seems to happen most of the time, only in the rare case can I reproduce errors I see in a production system with memtest86 reports.
- ThePowerOfFuet 3y agoSame guy too (Andrew Ayer). Someone is keeping an eye on it!
- KirillPanov 3y agoI cannot understand why they don't use two-of-three voting to produce the log entries. It doesn't have to be three machines owned by separate organizations, or even in separate buildings. Just three servers in the same rack, doing the same computations, and nobody signs anything unless one of the other two produces the exact same result.
- politelemon 3y agoIt's on the 4th line where it says 00000030: 9126 9384 .... Instead of 9284 ....
- dlahoda 3y ago[flagged]
- j16sdiz 3y agoIn CT, we want every inconsistency manually checked and reconfirmed. They are published to public maillist like this to deterrent attacks.
- gorgoiler 3y agoThe numbers from this Google SIGMETRICS09 paper are my usual benchmark for thinking about ECC DIMMs: https://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf https://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf Their metric of “25,000 to 70,000 errors per billion device hours per megabit” is a bit hard to grapple with. If you assume each error is a single bit then that’s 20 to 50 bytes per GB DIMM per month, or one bit per GB every two hours.
- p-e-w 3y ago> or one bit per GB every two hours How is that possible? Wouldn't such an error frequency lead to user-observable problems all the time? For example, in the average statically compiled codebase, flipping a single bit in source code being edited (with the code then being saved to disk) will make it fail to compile with high probability, which would be noticed immediately. Yet I've never encountered this situation in practice, nor heard of anyone else encountering it, and like most people I don't even use ECC RAM. That seems incongruent with the figure quoted above.
- viraptor 3y agoYeah, that's not realistic. This is as much code as goes through a CI I'm working with many times a day. I'd constantly see errors on an unknown variable if this was a real rate.
- teaearlgraycold 3y agoInterestingly, Google has a team of engineers dedicated to detecting hardware prone to these kinds of errors by perpetually QAing devices in the field. You can then replace the node before it breaks something important.
- nsteel 3y agoI would hope, at the very least, everyone doing important work is using ECC and actively monitoring their correction counters. This should be automated. We do a lot of extra testing on top of our vendor's to catch weak bitcells before a device is shipped to customers. Over many generations of tech, RAM faults have always been a small but constant source of faults.
- teaearlgraycold 3y agoRAM isn’t the only source of these errors. They can originate inside of a CPU core as well.
- nsteel 3y agoYes. And in long wiring paths (in fact, these are far more frequent than core logic faults). However, we've found atpg stuck-at and transition fault coverage to be vastly better with each generation. Due to improvements in the methodologies, the tools themselves, and our vendor's attitude. Of course, even with coverage in the high 90s that still leaves many paths unchecked but it's a small percentage of the faults (that we find...). But for our devices, RAM faults don't seem to be getting much better, they're always a pain point.
- Thorrez 3y ago>Unfortunately, it is not possible for the log to recover from this. That sounds bad. What does this mean? Does the entire log need to be thrown out, and a new log needs to be create to start from scratch?
- ryanwhitney 3y agoComment by the author (from the last time this happened) seemed helpful: "OP here. Unless you work for a certificate authority or a web browser, this event will have zero impact on you. While this particular CT log has failed, there are many other CT logs, and certificates are required to be logged to 2-3 different logs (depending on certificate lifetime) so that if a log fails web browsers can rely on one of the other logs to ensure the certificate is publicly logged. This is the 8th log to fail (although the first caused by a bit flip), and log failure has never caused a user-facing certificate error. The overall CT ecosystem has proven very resilient, even if a bit flip can take out an individual log. (P.S. No one knows if it was really a cosmic ray or not. But it's almost certainly a random hardware error rather than a software bug, and cosmic ray is just the informal term people like to use for unexplained hardware bit flips.)" -https://news.ycombinator.com/item?id=27731210 https://news.ycombinator.com/item?id=27731210
- vivegi 3y agoI have no knowledge of how the CAs maintain the CT logs. What is the process for a CA to rebuild the CT log, if one exists? Is it something like what is illustrated below? Let CA1, CA2, CA3 and CA4 be different certificate authorities. The set enumerated alongside each CA is the set of certificates logged into its CT log. CA1 : {c1, c2, c3} CA2 : {c2, c3'} CA3 : {c1, c2} CA4 : {c1, c3} Suppose CA2 is where an issue was detected with certificate c3 (the anomalous cert is denoted by c3') and CA2 trusts CA3 and CA4, then the set {c1, c2, c3} can be constructed after verifying the CT logs of CA3 and CA4 and merging their logs. Is that kind of how it would work or would CA2 just truncate its log and restart from this point forward?
- detaro 3y ago
- sushidev 3y agowhat happened here?
- swixmix 3y agoan append-only log became non-writable earlier than expected
- aeaa3 3y agoDo we know whether the machine on which this occurred had ECC memory?
- agwa 3y agoIt appears to be hosted on AWS, which claims to use ECC memory.
- Animats 3y agoThis is perhaps one of the few legitimate use cases for a distributed blockchain. Then several nodes have to agree for the chain to advance.
- p-e-w 3y ago> This is perhaps one of the few legitimate use cases for a distributed blockchain. It's incredible how a global propaganda machine has turned one of the most transformative technologies of our time into something that's considered shady and even crime-adjacent by default, and for which "legitimate" use cases are somehow considered special, when in reality there are countless potential applications for distributed ledgers. But the powers that be want to maintain centralized control at all costs, and their pushback has clearly reached even the minds of HN users.
- automatic6131 3y agoOr, hear me out, the technology sucks. And a great deal of software engineers here can see that the Emperor has no clothes.
- p-e-w 3y agoIf you can name an alternative technology that solves the same problem (integrity consensus without centralized authority), I'm all ears.
- automatic6131 3y agoTry again when you can solve the problem without introducing a dozen others that make everything, overall, worse. Until then, trusting a central authority will do just fine. It works well enough now.
- bob1029 3y agoWhy do we care about the centralized authority part so much? There are a lot of problem domains where zero trust is not desirable or understood to be a fool's errand. In a B2B setting where you have 3 parties working together on an integration project wherein 2 parties are B2B vendors and the 3rd is a B2C org, where does the centralization reside? What about vendors of the B2B vendors? I honestly don't know what we are getting at with this word anymore. If we were to apply something like SQL Server w/ Ledger tables (i.e. centralized blockchain) to this kind of problem in my shop, we would almost certainly find a solution that all parties would find agreeable. This forces you to trust 1-2 parties (i.e. Microsoft themselves), but in the above we agree that for many (most?) areas this may likely be explicitly desirable. There are also technologies (again, centralized) that provide non-repudiation through the hosting layer itself. Example of this being something like Azure Confidential Ledger. The part where this seems to get frustrating for people is the desired crystalline & immediate nature of the system. If you operate with the tiniest amount of extra flexibility you can get so much more done - E.g. perhaps the business can review a data tamper event tomorrow with their partners on the phone.
- molticrystal 3y agoReminds me of BitSquatting where cosmic rays, hardware, or other errors flip a bit in a domain name and an advantage you can gain by purchasing bitflipped domains. >Over the course of about seven months, 52,317 requests were made to the bitsquat domains [0] [0] https://media.blackhat.com/bh-us-11/Dinaburg/BH_US_11_Dinaburg_Bitsquatting_WP.pdf https://media.blackhat.com/bh-us-11/Dinaburg/BH_US_11_Dinabu...
- throwaway290 3y agoSection 5.3 of the PDF is veeery interesting. They check where the most requests to bitsquat domains come from. They chose microsoft.com as a more neutral reference than e.g. FB. And looking at the graph, the lion's share of requests to bitsquat domains for microsoft comes from... China?? Followed by Brazil? (and only then by US)
- zinekeller 3y agoWait, it's DigiCert's again? (Previous: https://groups.google.com/a/chromium.org/g/ct-policy/c/PCkKU357M2Q/ https://groups.google.com/a/chromium.org/g/ct-policy/c/PCkKU...) Do we have a list of all failed CTs?
- agwa 3y agoNot directly, but you can search for "Failed" on this page: https://sslmate.com/app/ctlogs https://sslmate.com/app/ctlogs
- amluto 3y agoIt seems to me that CT (or its operators?) should take a lesson from adversarial blockchains (cryptocurrency) here: a new state should not be propagated without verification. I think that, for CT, this should be fairly straightforward. Some machine with access to the signing keys should generate new nodes and signatures and push those internally to some front-end machines. The latter (on separate physical machines) should fully validate the result before propagating it any farther. No outside user sees the result until at least, say, 3 machines fully validate it. Then, if validation fails, the state could be rolled back internally. When a rollback occurs, the signing machine would think it’s signing a new, conflicting state, but that’s fine: no one outside the log operator has seen the old conflicting state.
- j16sdiz 3y agoIn Blockchain, there are concurrent, multiple, anonymous appends to that log. That's why you need concenses. In CT, all append are controlled. Nothing is anonymous.
- lima 3y agoA distributed consensus mechanism provides byzantine fault tolerance - which is helpful even with trusted actors, as this event demonstrates.
- amluto 3y agoIt shouldn’t even need much distribution or any protocol change. If every CT operator required three separate nodes to validate a proposed new block before distributing that block, then the only way to corrupt the log would be for an invalid block to pass verification three times (due to a bug or to a vanishingly unlikely coincidence of hardware errors) or for something to accidentally publish the block without verification. Unlike cryptocurrency, this would require no fancy public protocols, negligible computational resources, and no additional verification overhead outside the log operator at all.
- deleted 3y ago
- ransackdev 3y agoRelevant video from Veritasium on cosmic rays flipping bits and causing chaos https://www.youtube.com/watch?v=AaZ_RSt0KP8 https://www.youtube.com/watch?v=AaZ_RSt0KP8
- chunk_waffle 3y agoWe don't take bit flips seriously enough, practically every consumer device uses non ECC memory, very few folks use filesystems (e.g. ZFS) that can detect corrupt blocks. Even when using those things together it's still not perfect. Everything is terrible.
- chunk_waffle 3y agoWhy the downvote? Is it really acceptable for the machine reading your passport at the airport to have bit flip? Or the person processing you at the DMV, the doctor reading your medical history, or a million other things that power the modern world?
- sigio 3y agoOn reads it's easy enough to do a re-read.... bit-flips when writing something are somewhat more critical, as this is non-repeatable.
- dboreham 3y agoMy lifetime experience says to suspect software did this (I've had careers in both hardware design, designing large memory subsystems, and in software development). Yes it's one bit changed which makes the mind go to the ever present alpha particle, but code also flips single bits. If some library code inside a process generating this data wanted to update a bitmap structure but got the address wrong, you'd have the same outcome.
- H8crilA 3y agoFunny enough my experience says that it is hardware. Software bugs rarely manifest at such low frequencies (say 10^-15 or so) compared to hardware faults.
- jjoonathan 3y agoI just don't have the same level of trust in HW. My priors include a bunch of dead hard drives and a few bad sticks of RAM, all of which caused significant subtle damage before the big catastrophic failure that drew attention. None of these were ECC or RAID, but I've also seen enough foot-guns with ECC and RAID to place the probability of "solution degraded to consumer reliability" considerably north of zero.
- kevingadd 3y agoFlipping a single bit is a lot harder to do by accident in code than corrupting a byte or multiple bytes. You really are not doing bitwise operations on a regular basis in most types of software.
- benlivengood 3y agoDo some folks still not validate new entries in their certificate transparency logs on at least one other machine before publishing them? This is getting to the point (2 log failures in just under 2 years) that I wouldn't be surprised to see some certificates invalidated because they only used 2 transparency logs and both failed within the lifetime of the cert.
- bombcar 3y agoWith things like RAID and Reed-Solomon codes, we have the ability to have verifiably correct data even with some percentage lost; how come something like this isn't used?
- TanjB 3y agoDRAM is not greatly affected by radiation, because the capacitors are large structures relative to radiation events. SRAM is affected, which is why SRAM arrays should always use SECDED ECC. The dominant cause of DRAM failures is bit flips from variable retention time (VRT), where the cell fails to hold charge long enough to meet refresh timing. These are believed to be caused by stray charge trapped in gate dielectric, a bit like an accidental NAND, and they can persist for days to months. This is why the latest generation (LPDDR4x, LP/DDR5) have single bit correction built into the DRAM chip. Along with permanent single cell failures due to aging this probably fixes more than 95% of DRAM faults. The DRAM vendors sure could do a lot better on publishing error statistics. They are probably the least transparent critical technology used in everything, but no regulation requires them to explain and they generally refuse statistics on faults even to major customers (which is why folks like AMD run large experiments at supercomputer sites to investigate, and most clouds gather their own data). That said, DRAM chips are pretty good. The DDR4 generation probably had better than a 1000 FIT rate per 2GB chip, so in a laptop with 16GB that would have been less than 10 error per million hours, or under 1 per 50 laptops used for a year. For many of us the vast majority of data is in media files. I personally notice broken photos and videos every now and then. I would love to have a laptop with a competent ECC level, but they do not exist. Even desktop servers often come without. It is unclear how much better the LP/DDR5 generation will be since the on-die ECC still does not fix higher order faults in word lines and other shared structures, which may sum to as much as 10% of aging faults. All simply educated guesses, since the industry will not publish.
- jjoonathan 3y agoSo this might be a dumb question but it's been bothering me and you sound like you might know. What's the big advantage of DRAM over SRAM? In school we learned that DRAM was cheaper -- but surely the difference between 1T1C and 6T isn't more than 6 and my intuition says C is big so it's probably 2 or 3 or something for a given process generation. The problem is that the latency of DRAM is absolutely dreadful. On one hand I see a staggering amount of engineering that goes into hiding DRAM latency, and on the other hand I see that DRAM has become so cheap that many systems are over-provisioned by a factor larger than its theoretical cost advantage purely by accident. The "obvious solution" would seem to be DIMMs of SRAM with (comparatively) wicked fast timings -- but this doesn't happen, despite the fact that the memory industry is extremely competitive and filled to the gills with smart people, so presumably there's another factor that stops "DIMMs of SRAM" from being viable. Do you happen to know what it is?
- peter_d_sherman 3y agoIf there might be a software remedy (or at least amelioration) of the hardware problem of random bit flips affecting a software data structure -- it might involve creating multiple (more than 1, 2 or greater) redundant data structures in memory -- and checking each one for consistency against the others at specific intervals... If the random bit flips affect code -- then if the code is deterministic, then one solution (or at least amelioration) might be running multiple copies of the same code -- but loaded at different memory locations and where the results of one set of code's calculations are checked against the results of the same set of code's calculations -- but loaded and executed from a different memory location... Kludgy? Yes -- but if the underlying hardware is buggy (random bit errors which cannot be removed for whatever reason) -- then it may be the only effective way to make the system work, despite the kludgyness... Which brings up a strictly academic question -- what would an OS where each OS data structure and OS code path was replicated/redundant (and the results of running each redundant code path / data structure and the results checked at various intervals) -- look like? (I know NASA did something like that a long time ago with using something like five redundant computers where each computer checks the results of the computation of the group, and if there's an inconsistency, the computer producing the inconsistent result would be shut down...) Related: https://history.nasa.gov/computers/Ch5-5.html https://history.nasa.gov/computers/Ch5-5.html