3 ms·
Interestingly enough, it is possible to do private inference in theory, e.g. via oblivious inference protocols but prohibitively slow in practice. You can also
by sbszllr 9mo ago
Interestingly enough, it is possible to do private inference in theory, e.g. via oblivious inference protocols but prohibitively slow in practice.
You can also throw a model into a trusted execution environment. But again, too slow.
- ramoz 9mo agoModern TEE is actually performant for industry needs these days. Over 400,000x gains of zero knowledge proofs and with nominal differences from most raw inference workloads.
- sbszllr 9mo agoI agree that is performant enough for many applications, I work in the field. But it isn't performant enough to run large scale LLM inference with reasonable latency. Especially not when we compare the throughput numbers for a single-tenant inference inside a TEE vs batched non-private inference.
- ramoz 9mo agoWe just served Deepseek R1 on this bad boy in CC+TEE (and an integrated signing layer we developed for vLLM). https://pasteboard.co/k1hjwT7pWI6x.png https://pasteboard.co/k1hjwT7pWI6x.png reach out if interested in collab.