6 ms·
Reading the whitepaper, the inference provider still has the ability to access the prompt and response plaintext. This scheme does seem to guarantee that plaint
by ryanMVP 11mo ago
Reading the whitepaper, the inference provider still has the ability to access the prompt and response plaintext. This scheme does seem to guarantee that plaintext cannot be read for all other parties (e.g. the API router), and that the client's identity is hidden and cannot be associated with their request. Perhaps the precise privacy guarantees and allowances should be summarized in the readme.
With that in mind, does this scheme offer any advantage over the much simpler setup of a user sending an inference request:
- directly to an inference provider (no API router middleman)
- that accepts anonymous crypto payments (I believe such things exist)
- using a VPN to mask their IP?
- Terretta 11mo ago> the inference provider still has the ability to access the prompt and response plaintext Folks may underestimate the difficulty of providing compute that the provider “cannot”* access to reveal even at gunpoint. BYOK does cover most of it, but oh look, you brought me and my code your key, thanks… Apple's approach, and certain other systems such as AWS's Nitro Enclaves, aim at this last step of the problem: - https://security.apple.com/documentation/private-cloud-compute https://security.apple.com/documentation/private-cloud-compu... - https://aws.amazon.com/confidential-computing/ https://aws.amazon.com/confidential-computing/ NCC Group verified AWS's approach and found: 1. There is no mechanism for a cloud service provider employee to log in to the underlying host. 2. No administrative API can access customer content on the underlying host. 3. There is no mechanism for a cloud service provider employee to access customer content stored on instance storage and encrypted EBS volumes. 4. There is no mechanism for a cloud service provider employee to access encrypted data transmitted over the network. 5. Access to administrative APIs always requires authentication and authorization. 6. Access to administrative APIs is always logged. 7. Hosts can only run tested and signed software that is deployed by an authenticated and authorized deployment service. No cloud service provider employee can deploy code directly onto hosts. - https://aws.amazon.com/blogs/compute/aws-nitro-system-gets-independent-affirmation-of-its-confidential-compute-capabilities/ https://aws.amazon.com/blogs/compute/aws-nitro-system-gets-i... Points 1 and 2 are more unusual than 3 - 7. Folks who enjoy taking things apart to understand them can hack at Apple's here: https://security.apple.com/blog/pcc-security-research/ https://security.apple.com/blog/pcc-security-research/ * Except by, say, withdrawing the system (see Apple in UK) so users have to use something less secure, observably changing the system, or other transparency trippers.
- amelius 11mo ago> Folks may underestimate the difficulty of providing compute that the provider “cannot”* access to reveal even at gunpoint. It's even harder to do this plus the hard requirement of giving the NSA access. Or alternatively, give the user a verifiable guarantee that nobody has access.
- astrange 11mo agoThat's what non-targetability is for.
- 7e 11mo agoAt the end if the day, Nitro Enclaves are still “trust Amazon”, which is a poor guarantee. NVIDIA+AMD offers hardware backed enclave features for their GPUs which is the superior solution here.
- iancarroll 11mo agoAren’t they both hardware backed, just changing the X in “trust X”?
- almostgotcaught 11mo agobe sure to let us know when you can run eg nginx on a GPU in said enclave.
- Terretta 11mo agoYou think Nitro Enclaves aren't hardware backed?
- sublimefire 11mo agoYes but at the end of the day you need to trust the cloud provider tools which expands the trust boundary from just hardware root of trust. Who is to guarantee they will not create a malicious tool update and push it then retract it? It is nowhere captured and you cannot prove it.
- anon721656321 11mo agoat that point, it seems easier to run a slightly worse model locally. (or on a rented server)
- rimeice 11mo agoWhich is apples own approach until the compute requirements need them to run some compute on cloud.
- bigyabai 11mo agoJust a shame they spent so long skimping on iPhone memory. The tail-end of support for 4gb and 6gb handsets is going to push that compute barrier pretty low.
- brookst 11mo agoEh, maybe a bit, but those era devices also have much lower memory bandwidth. I suspect that the utility of client models will rule out those devices for other reasons than memory.
- bigyabai 11mo ago> much lower memory bandwidth Not really? The A11 Bionic chip that shipped with the iPhone X has 3gb of 30gb/s memory. That's plenty fast for small LLMs if they'll fit in memory, it's only ~1/3rd of the M1's memory speed and it only gets faster on the LPDDR5 handsets. A big part of Apple's chip design philosophy was investing in memory controller hardware to take advantage of the iOS runtime better. They just didn't foresee any technologies beside GC that could potentially inflate memory consumption.
- rasengan 11mo agoWe are introducing Verifiably Private AI [1] which actually solves all of the issues you mention. Everything across the entire chain is verifiably private (or in other words, transparent to the user in such a way they can verify what is running across the entire architecture). [1] https://ai.vp.net/ https://ai.vp.net/
- immibis 11mo agoIt's probably illegal for a business to take anonymous cryptocurrency payments in the EU. Businesses are allowed to take traceable payments only, or else it's money laundering. With the caveat that it's not clear what precisely is illegal about these payments and to what level it's illegal. It might be that a business isn't allowed to have any at all, or isn't allowed to use them for business, or can use them for business but can't exchange them for normal currency, or can do all that but has to check their customer's passport and fill out reams of paperwork. https://bitcoinblog.de/2025/05/05/eu-to-ban-trading-of-privacy-coins-from-2027/ https://bitcoinblog.de/2025/05/05/eu-to-ban-trading-of-priva...
- deleted 11mo ago[deleted]
- macrael 11mo agoHowdy, head of Eng at confident.security here, so excited to see this out there. I'm not sure I understand what you mean by inference provider here? The inference workload is not shipped off the compute node once it's been decrypted to e.g. OpenAI, it's running directly on the compute machine on open source models loaded there. Those machines are cryptographically attesting to the software they are running. Proving, ultimately, that there is no software that is logging sensitive info off the machine, and the machine is locked down, no SSH access. This is how Apple's PCC does it as well, clients of the system will not even send requests to compute nodes that aren't making these promises, and you can audit the code running on those compute machines to check that they aren't doing anything nefarious. The privacy guarantee we are making here is that no one, not even people operating the inference hardware, can see your prompts.
- bjackman 11mo ago> no one, not even people operating the inference hardware You need to be careful with these claims IMO. I am not involved directly in CoCo so my understanding lacks nuance but after https://tee.fail https://tee.fail I came to understand that basically there's no HW that actually considers physical attacks in scope for their threat model? The Ars Technica coverage of that publication has some pretty yikes contrasts between quotes from people making claims like yours, and the actual reality of the hardware features. https://arstechnica.com/security/2025/10/new-physical-attacks-are-quickly-diluting-secure-enclave-defenses-from-nvidia-amd-and-intel/ https://arstechnica.com/security/2025/10/new-physical-attack... My current understanding of the guarantees here is: - even if you completely pwn the inference operator, steal all root keys etc, you can't steal their customers' data as a remote attacker - as a small cabal of arbitrarily privileged employees of the operator, you can't steal the customers' data without a very high risk of getting caught - BUT, if the operator systematically conspires to steal the customers' data, they can. If the state wants the data and is willing to spend money on getting it, it's theirs.
- macrael 11mo agoI'm happy to be careful, you are right we are relying on TEEs and vTPMs as roots of trust here and TEEs have been compromised by attackers with physical access. This is actually part of why we think it's so important to have the non-targetability part of the security stack as well, so that even if someone where to physically compromise some machines at a cloud provider, there would be no way for them to reliably route a target's requests to that machine.