4 ms·
You’re putting words in my mouth again. I didn’t say those were the only two options. I asked what your silver-bullet alternative to mTLS is that has no issues
by mmalone 6y ago
You’re putting words in my mouth again. I didn’t say those were the only two options. I asked what your silver-bullet alternative to mTLS is that has no issues of its own. The position seems to be that mTLS has problems, so don’t use it. But everything has problems.
Re: revocation... the point is that you can revoke attacker access without revoking a certificate.
Or, if you want, you can revoke certificates. I’m not totally opposed to that. CRL is finicky, but architecturally it works well with short-lived certs and it’s not very different than secret rotation. If you can push new secrets, you can push a CRL. In fact, if you can push new secrets, you can also push new roots.
There are many people who use mTLS. Along with consistent service logging, it’s one of the main reasons people use service meshes. It’s also very common in IoT.
- tptacek 6y agoIf you can revoke attacker access without revoking the certificate, the certificate isn't doing all the authentication stuff you say it is. If you have to inform all the services in your network that a certificate (or an identity bound exclusively to a certificate) is invalid, you've implemented revocation. You can see how this looks like jazz-hands, right?
- mmalone 6y agoThe certificate is still containing blast radius when the attacker is in your network and limiting the insider threat from people who have access to some subset of production. Without a cert, they’d be able to access everything. With a cert, they can only access stuff that the identity bound to the certificate can access. It’s doing precisely what I want it to do. If you find someone malicious in your network, you lock them out of your network. You don’t need to revoke a certificate to do this. OR, you can do CRL. This isn’t jazz hands, it’s a choice. As you push mTLS harder, and need things like active revocation to satisfy your threat model, it’s worth reconsidering whether mTLS is the right choice. Sometimes it is, sometimes it’s not. Sometimes you don’t have a feasible alternative.
- tptacek 6y agoWait a minute. You're not being coherent. I'm the one saying mTLS is blast-radius containment. What that means is: it's not your primary security control. It's certainly not your primary authentication modality. It's a thing that constrains the impact of another security flaw. You can't "lock someone out of your network". That's handwaving. If you can do that reliably, you don't need any other security controls. Just lock all the bad people out! The problem with mTLS in configurations where you rely on it for authentication is, if an attacker manages to obtain a client certificate key, it no longer matters if you nuke the machine from orbit, because they have the keypair! The keypair still works! They can use that keypair from any other place on the network they have access to (if there's no such place, you just defined away the need for mTLS). Which is why you need active revocation if you're going to rely on mTLS: when you lose confidence in your sole custody of a client keypair, you need to actively revoke that keypair, immediately. You can't just wait 7 hours for the step-ca default key lifetime to expire! "Or, you can do CRL" is literally active revocation. Your project has a prominent call-out about how most people don't need to do this. But that's literally the opposite of the truth. It's true in the WebPKI, but I think you've become confused about the difference between WebPKI server certs and internal client certificates. Telling people they should deploy lots of mTLS but not worry about revocation is malpractice. What's going to happen to people in practice is that they're going to have to re-key their entire fleet any time there's any question about the integrity of a service. Or, more likely: they won't, because the mTLS deployment mode you're encouraging creates so much goddamn friction to protecting a key that people will roll the dice with the safety of their users rather than confront the fact that the correct engineering solution is an outage-inducing all-hands-on-deck rekeying. What blows my mind about this is, at 24-hour expiry, in a medium-sized application, you're going to have machines needing to refresh keys practically every hour of the day; your CA will need to be available 24/7. At that point, you've basically reinvented Kerberos. Ops teams fucking hate Kerberos! And for all that effort, you still face fleetwide updates any time something sketchy happens anywhere. I like mTLS for things like "we set up a Consul cluster; let's make sure just the machines that use Consul can reach it". It works fine for that. I don't think it's a good idea to take it much further.
- mmalone 6y ago