3 ms·
I'm curious why the connection from Rails is "mTLS-y" rather than actual mTLS. The macaroon repo mentions a Noise transport. Maybe to avoid dealing with X.509 c
by ramchip 3y ago
I'm curious why the connection from Rails is "mTLS-y" rather than actual mTLS. The macaroon repo mentions a Noise transport. Maybe to avoid dealing with X.509 certificates by distributing trusted public keys via LiteFS?
> We didn’t use the pre-existing public implementation because we were warned not to. [...] Macaroons decided to use untyped opaque blobs to represent caveats. We need things to be as rigidly unambiguous as they can be.
Another place where the implementations are limiting is in the nonce for discharge macaroons: it has to match the "challenge" exactly. I think the fly.io implementation makes the nonce a structured object that includes the challenge / key ID as a field, and extracts that field during third party caveat verification, which is nice. It makes it possible to include a random nonce (in the cryptographic sense) or revocation ID for instance.
I experimented recently with discharges containing service-specific info, e.g. a discharge from an auth service could contain the user's name (for logging), etc. It felt dangerous though, because there's potential for confusion on which service is allowed to provide what info - if we have a 3P caveat for an auth service and a revocation checking service, we want to be sure that the user profile info comes from the auth service, not the other one, and it gets complicated to encode that in the token.
Maybe this is what "proof" macaroons solve? They're not mentioned in the blog post or macaroon-thought.md, but seemed to be about making positive statements on something.
- paulmd 3y agoThis has been a perpetual problem we’re trying to solve at our org and while I was on that team I was pushing for the idea of macaroons (we just use JWT and we suffer many headaches because of it). My suggestion at the time was macaroons with some hierarchal PKI - servers themselves have a macaroon that delegates the permission to sign tokens that meet certain criteria, delegated by us the root organization. So in most situations a typical key hierarchy would look like “root key -> signing servers -> normal services -> bearer token”. If a token is stateless then this just becomes another caveat and signature - show me that you have the authority to sign this key under this attenuation. Signing server? You should only be signing keys for key services. Data service? You should only be issuing tokens with X caveat. Etc. If you want revocation though it’s not really stateless - at minimum you have to store a list of revoked tokens until the parent tokens that issued them have expired. In fact you would want to build this as a routine capability on every single level - signing services should rotate their own keys (this is that “clearing caveats” thing) and issue their own revocations, and then the assertion is that any signing from that key is invalid from that moment forward. Similarly, if a browser session wants to refresh its token, it needs to commit to never using the old one again, for the lifespan of the IAM token, and we need to track that for the lifespan of the IAM token. But there is a defined TTL after which you can clear the revocation - and it’s the lifespan of the token used to issue it. Which is a very Redis-with-cross-region-replication sort of task, or perhaps Postgres for persistent backing. The good news is that there is no logical split-brain, the mere fact that another region issued a revocation is ipso facto all you need to know, when and why it happened is kinda irrelevant to the fact that the token is now revoked. In other words - signing servers probably have an IAM representing them as an instance, with their own internal secret/private key, which is used to issue their own "bearer token" (public key) which rotates periodically. And when they do a clearing etc they are the authority that signs that the new macaroon is valid. And after some time their own bearer token is revoked (eg after 2pm any signings are invalid) and they rotate to signing with a new bearer token. If you are going to rely on propagating revocation efficiently it needs to be a routine fact that’s occurring regularly, right? And if you don’t see a revocation from a signing server on some expected timeframe, it is probably in fact a sign of split-brain occurring… you’ve lost a server or a region and it’s in fact dubious to continue accepting that token anyway. So this sort of leads to an inherent “valid -> historically valid but not for new tokens -> expired” lifecycle that mirrors the IAM/bearer model that clients also use. This actually closes the "fail open" nature of the revocation - it is impossible to have a "split brain" in the usual sense of revocation not propagating to a region etc, because you know that a signing-server key isn't valid for more than 15 minutes anyway etc. And perhaps you could even add a "parentKey" field and use the bearer token itself as a carrier of these revocations (if the user has a new key and it's validly signed, then SOMEONE must have revoked the old key...) although I haven't rigorously thought that one through. I am drawn to this revocation idea like a moth to a flame despite knowing how much complexity is going to live in making sure that service doesn’t ever go down or lose data or go split-brain. But if you are doing this “clearing” idea at least you don’t end up with a giant tree of historical caveats and sub-signatures. On the other hand the reality is that really nobody wants true stateless especially for browser tokens… “log me out everywhere” is practically table stakes and that implies some kind of revocation. And revocation is implicitly state - it’s either valid or revoked, that’s state even if you aren’t mutating the token itself, you’re still mutating the validity of the token. Even if it’s a global “user revoked all tokens at 2pm” it’s still mutating state etc. The idea of revocations lets the “happy path” be stateless, and handle this smaller amount of replication of revocations etc, but otoh it’s still very load-bearing and revocations fail-open. But since the token is limited in time and scope anyway, that’s potentially a valid tradeoff… But I haven't actually implemented any of this, so, take it for what it's worth.
- tptacek 3y agoI cut this back drastically from earlier drafts which talked more about how the verification service was deployed, but wanted to save at least one diagram with a Muppet in it, and spaced on blanking out the mTLS/TLS thing. Long story short: we originally deployed the verifier on NATS, not with direct HTTP, and so TLS wasn't an option; we wrote a Go implementation of Noise. It's just HTTP now (but still Noise for now). The signing interface is mTLS-like (client-authenticated) because you can't just let any component sign. But really any component can verify, so all you need is server authentication.
- ramchip 3y agoMakes sense. The Muppet was definitely needed. Thank you!