5 ms·
Appreciate the thoughtful response! Reading it over, I think we mostly agree on the facts. It's easy to do mTLS and x509 wrong. The question, then, is what's e
by mmalone 6y ago
Appreciate the thoughtful response!
Reading it over, I think we mostly agree on the facts. It's easy to do mTLS and x509 wrong. The question, then, is what's easier / more secure: doing mTLS/x509 right or doing something else? I think that's somewhat subjective: it depends on your requirements, your environment, and your skillset.
One point that I'd like to reiterate is this: if you want a consistent cryptographic solution that works everywhere, TLS is pretty much your only choice. You could use something else for client authentication, but you probably still need TLS.
As a strawman, here's a sketch of how I'd recommend doing TLS in a microservice system. I consider this "right" for most garden-variety microservices-in-cloud scenarios and don't think it's particularly hard to do. Most of this is already implemented in https://github.com/smallstep/certificates https://github.com/smallstep/certificates:
* Deploy the root cert via automation (so it's quickly rotatable) and/or keep it in a managed HSM/KMS. You might harden root rotation a bit by signing your new root with your old root. But, generally, trust config management or container orchestration to push root(s) (you already trust it to push code and secrets). Root rotation (and, thus, bulk revocation) is now as fast as secret rotation (secrets are generally pushed the same way).
* Issue short-lived certificates per logical entity. If it gets a box and a name in your architecture diagram, it should get an identity and each instance should get a certificate. Use domain names and email addresses that you control for names. Keep certs simple: one SAN. Certificates bind a name to a public key. That's it.
* Automate certificate issuance. ACME can work for this, but there are other options (single-use tokens issued by config management, cloud-managed instance identity documents or service accounts, an existing device certificate issued by a manufacturer, etc.)
* Automate certificate renewal. A simple mTLS HTTPS request works for this. This is easy to implement and easy to scale out with multiple intermediates. "Revoking" a certificate just marks it as "not renewable". To reduce risk of outage, in this architecture, it's safe to renew an expired certificate as long as it's not revoked (ACME-STAR basically does this, but it's push instead of pull).
* If you really need active revocation, fine. One good solution is to push CRL to a cloud storage bucket. Short-lived certs will keep your CRLs small. If you need to do a mass rotation, rotate roots (push new root, wait for rotation, pull old root).
* Use secure NTP for time.
* Index issued certificates. CT (trillian) is cool if you want to be fancy. Your existing database or SIEM also works. zcertificate can parse x509 and output a JSON representation of a certificate that you can map to something like an Elastic Search schema: https://github.com/zmap/zcertificate
I want to respond specifically to your first and final points.
On your first point: I understand that in theory an attacker could slip a request across a secure channel, and binding authentication to a request could in theory prevent that. I don't understand how that's likely to happen in the context I'm thinking of here. Which may be different than the context you're thinking of. So let me clarify.
Suppose I have `<end-user> -> <service-a> -> <service-b> -> <database>`. Let's focus on `<service-a> -> <service-b>`. I don't see how using end-to-end mTLS, terminating in `<service-a>` and `<service-b>` application code, would be any more vulnerable to this variety of attack than an HTTP Basic header like `Authorization: Basic base64(service-a:password)`. Surely, the logic in `<service-a>` is simply "insert HTTP Basic header into requests on their way out to `<service-b>`". It doesn't matter if we're authenticating the request or the channel. If you're able to smuggle something malicious into that request, it's gonna get sent over to `<service-b>` with proper authentication attached.
Are we talking past one another? Are you trying to make `<end-user>`'s authenticated identity carry through `<service-a>` to `<service-b>`? If that's the case, then yes: I see what you're saying and you shouldn't use mTLS for that. I'm not sure if there's a term-of-art here, but I call this "end user identity propagation". You need something like a top-of-stack ticket service (a bearer token) for that. Or, better yet, macaroons. I consider those two separate things, though. mTLS is for authenticating your immediate peer. For end-user identity propagation mTLS is a poor choice.
On your final point: you could, in theory, express claims in x509. I'm sure you're aware, but it's been tried before (e.g., SPKI/SDSI). However, I agree that, unless you really know what you're doing, x509 is too complicated for that. Don't do it. You'll likely screw it up. If you're parsing x509 and ASN.1, you're doing it wrong. If you're processing strings that you've extracted from a certificate, and you're not in the habit of writing your own formal languages, you're definitely doing it wrong. Just put a flat name in a SAN. The only thing you should ever need to do with that string is an exact string comparison. If you need to know roles or groups or some other metadata look them up in a database.
(Or use macaroons)
- tptacek 6y agoAt some point, in most designs, you're going to end up with a client and server service where the server does something on behalf of a user. The identities in this design now include [client-identity, server-identity, user-identity]. I'll stipulate to mTLS resolving client-identity and server-identity. But the client is making requests of the server that pertain to a user. The certificate doesn't (and can't) attest to a user. In most designs I've actually evaluated, what ends up happening is that the server just trusts the user provided by the client, and the client tries hard not to ask for things for the wrong user. But beyond whether the client has authorized every code path that generates a request, you have an additional problem here, because even after the client authorizes a path that generates a request, a bug in the client can give an attacker influence over the request itself. The server has no way of distinguishing between a corrupted request and a real request, even though you're relying on a secure channel. Authenticated requests mitigate this problem: an attacker might be able to corrupt a request, but it's not enough to corrupt it; you need to know whatever secret authenticates the request itself. To do this right, you now need two authentication schemes in play: one for the secure channel, and one for the requests themselves. But if you can reliably authenticate requests, why are we rigorously authenticating the secure channel? We're spending complexity chits to buy only marginal extra security. I think it can make some sense to mesh up services with mTLS as a "you must be this high to get on the ride" mechanism. But since that's all it's doing, we don't need a lot of complicated mechanism to give precisely the right certificates to services, because even if something gets screwed up, request authentication, not secure channel authentication, should be what's protecting your application. Especially when we think about short-lived certificates and ACME and rollover: what is this really buying us in a strong system with request authentication? A WebPKI ACME certificate will live for 90 days. If certificates matter to our design, we can't let a compromised cert live for 90 days. Our certificate lifespans have to be much shorter --- the duration in which we'd be OK having a compromised credential still viable for an attacker. Remember the differing consequences between a WebPKI certificate compromise and an internal credential compromise: one is a second-order flaw that lets attackers with some other vulnerability MITM a subset of users; the other is a game-over flaw. How long would you let a compromised developer prod SSH key live? So now we're talking about very short-lived certificates, and you can see, this is a design that spends all the complexity chits asymmetric encryption costs, but is asymptotically approaching the operational resiliency of Kerberos. Yikes. The WebPKI has to assume certificates might be compromised because there are 18 zillion different ways people might mishandle a certificate. But that's not at all true for a data center deployment of a microservice ensemble! Certificate keys are as secure in our system as the microservice itself is. If you lose one to an attacker, you lost the microservice too --- even if you rotate the cred, whatever flaw gave the attacker that certificate is just going to give the attacker the next one, too. If you stop trusting a certificate (and, by extension, the instance it was resident on), you zap the whole instance, revoke, and root-cause the flaw that caused the problem. Rotation just isn't winning you much. At bottom the issue here is that the WebPKI has a very particular threat model, and it's a shitty threat model, and we have built lots of tools for that threat model, and many of them are by necessity shitty because we live in a fallen world. But your microservice ensemble wasn't born with original sin! You don't need to inherit the complexities (e.g. X.509) and limitations (e.g. X.509 revocation) that the WebPKI has to deal with. Which is a reason some people get itchy when we talk about importing WebPKI technology to things like K8s. I tend to be pretty chill about little islands of mTLS, like "all the Consul clients need a cert". That, to me, solves a practical problem. I am way less chill about attempts to create coherent PKI namespaces for all the components of an app, tied together with mTLS. People should use Macaroons!