35 ms·
It is always the DNS
- geocrasher 5y agoObligatory DNS Song reference: https://soundcloud.com/ryan-flowers-916961339/dns-to-the-tune-of-let-it-be https://soundcloud.com/ryan-flowers-916961339/dns-to-the-tun...
- tptacek 5y agoI propose we stop attributing to DNS reliability problems caused by DNSSEC. It's not fair to DNS, which gets a bad rap already. The bug here as I understand it: this hosting provider has a client whose DNS authority server DNSSEC-signs their records. That server is misconfigured. When you look up a nonexistent record in DNS, you get a simple response back that says there aren't any matching records. Easy enough! But you can't do that in DNSSEC, because denial "needs" to be signed. What DNSSEC does instead is essentially return a descriptor of a lexicographically ordered span of records that does exist, and doesn't include the record you were looking for. You notice that, and interpret as denial. See, simple! Oh, also, the names in those records are password-hashed, because otherwise DNSSEC would be publishing raw zone dumps to serve errors. Except those password hashes are very straightforwardly cracked. But I digress. The records carrying these spans in modern servers are called NSEC3s. Think of a tuple (start-name, stop-name, record-types). If you asked this client's server for a nonexistent A record, you'd get an NSEC3 back that properly established the covering span of names in which A was nonexistent --- but it erroneously signaled that other record types for that span (most notably MX) were also nonexistent.† The fun part of this bug is that it's effectively a vulnerability: an attacker can query random recursive caches for nonexistent record types to generate a cached denial records for records that do exist. Which is not something that can't go wrong in ordinary unsigned DNS! It's a vulnerability that DNSSEC adds to the DNS. This is your periodic reminder that virtually none of the industry's largest tech companies (and concomitantly well-resourced security teams) use DNSSEC, which you can confirm quickly for yourself by running `host -t ds <domain>` and noting that they don't have DS records. Previous (recent) thread, in which DNSSEC blew up (checks notes) all of Slack: https://news.ycombinator.com/item?id=29378633 https://news.ycombinator.com/item?id=29378633 † (If I've got the particulars of this wrong, that makes it even funnier, since it's suggesting that it's difficult even to describe what can go wrong with DNSSEC).
- wahern 5y agoWe don't need DNSSEC because we have TLS certificates. And as we all know, no large tech company has ever had trouble with misconfigured or expired TLS certificates.
- tptacek 5y agoWhy not both sets of problems?
- belorn 5y agoThe core issue behind both sets is single point of failure. A misconfigured security system can causes down time. The problem is that security systems tend to be single point of failure by design that fail "safe" by not allowing access. This would be an excellent security research topic. How to design non-single point of failure security that both fail safe and do not cause down time when misconfigured.
- cookiengineer 5y agoIn addition to what you wrote, I'd add from my own experience: - never trust a single DNS server as a source of truth - always rely on transport encryption, which DNS on port 53 (via UDP) does not offer - verify your results statistically, rotate servers (DNS ronin), and endpoint protocols in order to verify there's no MITM - deactivate message compression if possible. Naive implementations are vulnerable to both stack and buffer overflows (see NAMEWRECK). - don't rely on DNS cookies, and always assume there can be multiple records with different types in a reply to a single question. - DNS TTL 0 is a fix for stupidity, and nothing else. If any of the above isn't respected it's your fault as a DNS client, because the protocol was designed when the internet was a nice place and everybody literally knew everybody via the /etc/hosts file dumps. That's why I'm a strong believer in DNS via HTTPS, because having a better and more reliable transport protocol would fix a lot of the issues people/clients have to face. DNS via TLS turned out to be useless because the dedicated port 853 for it literally gets blocked by ISPs and nobody can do shit about it.
- ate53 5y agoWhat was the implementation that produced the bad NSEC3 proofs?
- wruza 5y agoDNS issues rumors/news always puzzled me (not that I ever made my honest research). Domains name system is to me basically a half-replicated/half-redirected database of {domain-name : categorized-ip-list}. Can someone please ELI5 why can’t we just GET https://<dns-ip-and-its-cert-from-dhcp>/get-ips?d=<domain> https://<dns-ip-and-its-cert-from-dhcp>/get-ips?d=<domain>, follow 3xx, get json and that’d be it? This question lies in the same category of questions which I’ve never managed to get answers to. Why fetching email over https is instant, but over imap it may take minutes to “sync” (or what it does there)? Why browsing is instant, but ftp was dead slow? Why all the formats (mailbox, uue, etc), when you could have a proper database with column encodings? I’m not asking anyone to answer these, cause it would be too offtopic, but these arcane issues have never left me alone. I understand that most of it is historical, but what the heck, internet.
- xxpor 5y agoThats essentially what dns-over-https is :)
- cookiengineer 5y agoNote that this is only partially true, because the RFC includes only the DNS wireformat (for TCP, with a leading bytesize). [1] The DNS JSON API is a google only thing, because even cloudflare's DNS use different MIME types and don't support the complete API and have varying support for base64 with/without padding characters and the full 65k frame size that the spec and drafts define. EDNS request options are also mostly unsupported in practice. Even multiple questions (A/AAAA) are partially supported and oftentimes fail dependent on which server you get as an endpoint. The URL schema is also different from server to server because everybody implements their own thing. [1] https://www.ietf.org/rfc/rfc8484.txt https://www.ietf.org/rfc/rfc8484.txt
- msravi 5y agoSomewhat related. If I use DNS-over-HTTPS with a nextdns edge server the response time is of the order of 10ms. With google 8.8.8.8 or nextdns 45.90.* it's about 130-140ms. Ever since I set up a local pihole with nextdns DoH as the upstream server, things have been very spiffy. The pihole cache also helps with sub-ms response times, especially if you set a minimum ttl of 40 mins or so.
- throwaway984393 5y agoIt seems like the world today operates mostly on technology that is replaced every 3 years, so nothing new ever becomes a standard. It's so wonderful when I need to interface with a standard internet protocol and it works the same way it used to. Last year my company decided to move one of its public services to a new hosted vendor and change their domains entirely. This would involve (among other things) changing the domain that customers were using and changing live traffic to point from one vendor to another. Of course nobody budgeted for any downtime, and the system was so little-known that it would have taken forever to figure out what all would be affected by a maintenance window. A less experienced engineer might consider this trivial; just change the DNS records and flip from one live system to another. Easy, right? Luckily, something like 15 years before, I had managed (among other things) DNS records for a large internet site. Because we were essentially hand-crafting Bind zone files and making all sorts of weird changes to tons of records, I learned all about the DNS ecosystem. I learned about all the different layers of the stack, all the different forms of caching, the common problems (UDP vs TCP, EDNS0, firewalls, Apex record limitations, negative cache, ignored TTL), spam features, service identifiers, XFERs, and the fact that there's a dozen different RFCs all laid on top of each other to form "the DNS protocol". Keeping in mind all the potential problems, I was able to plan our migration steps thoroughly enough for the 144-hour-long change window. I knew to prepare my SOA, NS, A & CNAME changes, verify the records on the nameservers (and from client devices) pre-change, wait 2 days after changing the TTL limit, verify with resolvers around the world to confirm everything was picked up, and validate on each device that the change worked correctly. The process worked beautifully and the customers never noticed a thing. And if there had been a problem, my changes would have resulted in no more than about 60 seconds downtime. What struck me at the end wasn't how much bizarrely complex and deep knowledge was required just to make the switch-over successful. It was that I had learned everything I needed to know 15 years before, and all that technical knowledge remained the same. I think that's only happened one or two other times in my career (one related to TCP weirdness, another related to C programming)
- abagheri43 5y agooh yes . a good nese