4 ms·
Someone just sent me a tweet from @halvarflake where he asked for the threat model under which service-to-service TLS is the best solution. I wanted to cross-li
by mmalone 6y ago
Someone just sent me a tweet from @halvarflake where he asked for the threat model under which service-to-service TLS is the best solution. I wanted to cross-link these threads since I think they're related: https://twitter.com/halvarflake/status/1337308903512678401 https://twitter.com/halvarflake/status/1337308903512678401
First, agreed. Certificates can't and shouldn't attest to the user's identity. You need another mechanism for that.
If you're securing requests, then securing the channel might be redundant. However, I do want authenticated service identity (not just user identity). The shape of policy I want to enforce is: "<service x> should be allowed to make <rpc> to <service y> in the context of responding to <request> from <user z>". I need the authenticated identity of both the service and the user to do that. Either that, or the authorization could take place higher up in the stack and I need a capability token passed through. Even with a capability token, it's nice to know the authenticated identity of the calling service for other purposes, like audit. Not to beat a dead horse, but I realize that macaroons satisfy all of these requirements :).
I also want confidentiality. Macaroons don't do this. This is an important point that I want to make more strongly this time, because I feel like it's getting lost in the noise. Authentication isn't the only thing people are using TLS for.
I appreciate your perespective. I think the best way to respond is to actually answer @halvarflake's question: what's the threat model that justifies service-to-service mutual TLS?
One answer is: that's the wrong question. Many people use mTLS for flexibility, not security.
Most microservice systems are born as a bunch of web services thrown in a VPC (or, more recently, a kubernetes). They can all talk to each other. Some of these systems are successful, and eventually outgrow their perimeter. Sub-components need to be able to communicate across untrusted channels (e.g., the web). TLS can give you a "cryptographic VPN" to do this.
There are a lot of VPN-y solutions with a lot of pros and cons. TLS isn't always the right way. The big pro for TLS, however, is that it works everywhere. Even if you don't control the host stack (e.g., functions as a service) or if you're on a constrained device, etc. TLS is also baked into most infrastructure: proxies, queues, databases, etc. You can easily terminate TLS easily at your perimeter if that's what you want. Alternatively, you can connect directly from a service in one perimeter to another service or to a piece of infrastructure like a database in another, without a proxy. Again, TLS is ubiquitous and it's flexible.
At this point I think we're at your "you must be this high to get on the ride" model. Why go further? This is where it makes sense to start talking about threat model.
First... I already mentioned this, but, once again: confidentiality. I think this is less controversial, and it doesn't require client certs. So I'll not elaborate further except to say that people need it.
The interesting topic is: service authentication. Why use client certificates for service authentication? I contend that service-to-service mTLS plus a bit of simple, coarse authorization actually does address a number of common security threats. Specifically:
* Malicious insiders: if you're using microservices and practicing DevOps and agile and have engineers on-call you probably have a lot of people and a lot of code that can get inside your perimeter. Coarse segmentation of service-to-service interactions can help here. An engineer who has the ability to `ssh` or `kubectl exec` or tunnel into your perimeter won't immediately have access to every sensitive service you run.
* Lateral movement: very similar to the above, but this time a service is popped or an engineer's laptop is compromised. Service authentication helps control the blast radius if someone gets inside your perimeter.
* Credential leaks: secrets are hard to manage. They get committed to repos, accidentally logged, dumped in error messages and served to end users, pasted in to slack channels, walk out the door with dismissed employees, etc. Using certificates with a relatively short lifetime gives you assurance that compromised credentials will be eventually rotate once an attack vector is closed. I contend that it's actually a lot easier to set up client certificates that automatically rotate than it is to rotate secrets for all of your proxies, databases, services, etc. An asymmetric crypto key is also inherently less likely to leak to logging infrastructure or end users than a password or bearer token since it's not actually in any requests.
Request authentication could mitigate these threats, too. And, again, carrying an authenticated end-user identity through the stack to downstream services is valuable. But, in a microservice system, I'm less worried about code lying about the end-user it's acting on behalf of than I am about malicious humans. If I already have TLS for "my VPN", authenticating clients and coarsely authorizing service-to-service interactions starts looking pretty attractive.
Regarding compromised keys, certificate lifetimes, and revocation... that's all important, but less critical to the threat model than it may appear. If a key is compromised I'm still much better off than I was without my coarse segmentation: the only stuff you can access is the stuff that the service is authorized to access. If an engineer exfiltrates a key to "test something", that key will eventually get rotated, even if I never know it happened. Same with an attacker exfiling a key: they have to maintain a presence on compromised systems to exfil new keys as they rotate. (FYI: by default, the smallstep toolchain issues certs for 24 hours. Personally, I'd push that down to a few minutes.)
It's also worth remembering that we still have our perimeter (or perimeters). If we find an attacker, we can lock them out at the perimeter instead of revoking a certificate. Even if you can TLS through the perimeter, we only need to enforce revocation at stuff directly exposed to the internet (e.g., in a proxy, not necessarily in every service). Active revocation is nice. But it's not necessarily critical. As another example, since you mentioned SSH: smallstep's SSH product uses certificates without active revocation. We revoke access in policy, by pushing ACLs to hosts and locking users out via NSS and PAM. The certificate may still prove I'm "mike", but mike doesn't have access anymore.
Why not do all of this host-to-host instead of service-to-service? Well, depending on your requirements, that may make sense. Certainly in kubernetes-land the line between "a host" and "a service" are blurring. TLS still has a few advantages:
* It works in places where you can't control the host (e.g., functions): it's a consistent mechanism that works everywhere
* It makes applications more self-contained and, therefore, more portable
* It's logically closer to what you want: you're trying to authenticate services not hosts
* There's more deployment flexibility: you can terminate in a sidecar or in the application
* It's easier to pass authenticated identity information into the application
* It works for authenticating to infrastructure: proxies, databases, queues, etc (again, it's a consistent mechanism that works everywhere)
Fin.
(Are we using macaroons yet?)
- tptacek 6y agoIn a lot of this, I see you trying to frame the argument as "should we do mTLS or should we do no TLS at all". And while there are threat models where TLS doesn't make any sense (Halvar's same-host or same-rack K8s model), I don't think people are generally advocating against encrypting network traffic. Encrypt your network traffic. Obviously, you don't need mTLS to do this. Another sleight of hand I see here is the ambiguity between network authentication, service authentication, and user authentication. When it's convenient for the argument, it seems, mTLS is just a way to ensure both ends of a network connection are part of the super best friends club. But later, when it's helpful to argue that mTLS is "flexible", certs carry credentials. Maybe even JWTs! There's a subtext to the "be wary of mTLS" argument, and it's that unified site-wide PKIs don't work well in practice. Finally, you've lost me completely with this "revocation is less critical to the threat model than you think it is" stuff. The step-ca website is even explicit about this distinction you're trying to draw between "active revocation" and "passive revocation" (short-lifetime certificates). No. That doesn't work in internal application environments. step-ca defaults to 24-hour certificates. Have you considered how weird that is? 24 hours is short enough that your CA needs to be highly available --- ask around about production outages that have been caused by certificates expiring at much longer lifetimes --- but long enough that you can't possibly "wait out" a compromised cert. I feel like something that is happening here is that WebPKI-style mTLS advocates think that a server cert and a client cert are two different instances of the same thing. They are not. They are wildly different entities. A server certificate prevents a MITM active attacker from intercepting a TLS connection. That's a rare attack, and, more importantly, a second-order attack. A client certificate is an access token for the application. Lose a prod-viable SSH key and then tell me I should wait 7 hours for it to expire. The reality is that every application that deploys "passive revocation" is ultimately going to have to implement "active revocation", and, if they follow the advice on the step-ca website, they're going to have to do it on no notice, and in the most disruptive way possible, because when an internal access credential is lost, it has to be shut off immediately.
- mmalone 6y agoIt's not intended to be sleight of hand. Confidentiality is the gateway drug. The argument is that, once you're using TLS for that, the incremental cost of adding client certificates and doing mTLS to reduce blast radius is lower. It's not zero. It is lower. I'm also not arguing that some active mechanism to immediately remove attacker access is unimportant. Of course it is. I'm arguing that active certificate revocation is not the only way to achieve that. I gave two examples: remove access at the perimeter or push new policy to revoke access for the compromised entity. Most microservice systems already have tools here: they're typically adding mTLS on top of a perimeter model with decent privileged access management. This is a bit of an asymmetric argument, in any case. I have all of my cards on the table. The opposition isn't really proposing any concrete alternative, so it's hard to make a meaningful comparison. If you want active revocation, implement active revocation. It's not easy, but OCSP and CRL exist and they do work. I even suggested a way to do it well: push a CRL to a cloud bucket, which is scalable and highly available. You can call that "cargo culting all of the idiosyncracies of the WebPKI", but I still haven't heard a concrete alternative. So where does that leave us? The only concrete thing that's been proposed... that I had to propose myself, since no one else would even answer this very simple question, is macaroons. Which I've said repeatedly I think are a good idea. However, it's abundantly clear to me why people, in practice, are choosing client certs over macaroons. Macaroons can't layer into an existing system the way that mTLS can, and if all you want is "blast containment" you can get that with a fairly simple client certificate setup and a bit of coarse authorization. Regarding CA availability: if you can't remediate an availability issue with a shared-nothing web service in 8 hours (the downtime you'd need for a cert to expire in a default step-ca setup) then you've got bigger problems than certificate management. The 24 hour default is long enough for people to remediate an outage and short enough to get people operationally comfortable with the idea of short-lived certificates. If I were to use shared secrets instead of asymmetric keys, and I created a system that rotated all of my service-to-service shared secrets automatically every 24 hours would you say that's a bad idea?