4 ms·
To throw out some real and up-to-date numbers from [1] for FHE at "128-bit security level", to sort 8x 8-bit unsigned integers on the most ordinary of desktop P
by dhx 2mo ago
To throw out some real and up-to-date numbers from [1] for FHE at "128-bit security level", to sort 8x 8-bit unsigned integers on the most ordinary of desktop PCs, wait 3 seconds for the result. Want to sort 32x 8-bit unsigned integers instead? Come back 34 seconds later for the result.
update: also see [2] for some primitive unsigned 64-bit integer operation benchmarks with the TFHE-rs library (winner in the sorting performance comparison of [1]). Equality at 80ms, addition and subtraction at 100ms, division at 8 seconds, etc.
[1] https://eprint.iacr.org/2026/1495.pdf https://eprint.iacr.org/2026/1495.pdf Oblivious Sorting under Fully Homomorphic Encryption: A Comprehensive Survey and Performance Analysis, Omar Ahmed and Rostin Shokri and Nektarios Georgios Tsoutsos, 2026
[2] https://docs.zama.org/tfhe-rs/tfhe-rs/1.0/get-started/benchmarks https://docs.zama.org/tfhe-rs/tfhe-rs/1.0/get-started/benchm...
- pamcake 2mo agoBenchmarking code in repo: https://github.com/google/heir/tree/main/benchmark https://github.com/google/heir/tree/main/benchmark Project intro talk from 2023: https://www.youtube.com/watch?v=kqDFdKUTNA4 https://www.youtube.com/watch?v=kqDFdKUTNA4
- j2kun 2mo agoThe first link has a misleading name. Instead use these two links for a better picture: https://github.com/google/fully-homomorphic-encryption/tree/main/demos https://github.com/google/fully-homomorphic-encryption/tree/... https://fhe-benchmarking.github.io/ https://fhe-benchmarking.github.io/
- tbenst 2mo agoThat is sobering for sure, I wonder what the theoretical bounds are on what is possible if known. Would be such a dream to use a Frontier LLM one day with homomorphic encryption, but this sounds wildly implausible based on where things are today.
- joshspankit 2mo agoWe’re not even close to the limits of AI optimization so finding the theoretical bounds is going to have to wait
- redeeman 2mo agothey would never allow it, atleast not for regular plebs such as you or I, consider if you made it say something politically incorrect? cant have that
- odo1242 2mo agoIt’s slightly better for LLMs because FHE is really bad at branches (it ends up essentially having to try both branches), making sorts nearly the worst possible thing to try since it’s all branches. In the case of AI most things are just addition and multiplication which can make some things faster since there aren’t as many branches. But we’re still nowhere near viability.
- matthewdgreen 2mo agoI’m genuinely not an expert, but isn’t the beauty of MoE models the fact that we explicitly don’t evaluate every parameter on inference? We evaluate exactly the subset that are needed to evaluate a prompt. Seems like this will bring back data-dependent branches again.
- charcircuit 2mo agoIt would also kill speculative decoding. You would have to run a full inference pass for every token instead of being able to generate multiple tokens with a single pass.
- odo1242 2mo agoPretty much, and this does a good job of illustrating the fundamental issue with branching. You could use an encryption scheme that allows the server to determine what MoE expert to load (the simplest would be to have the client decode the value and send it back to the server, though this can sometimes be possible to do without the round trip), but then it’s not fully homeomorphic because the server has some info about the computation that could be used to recover stuff about the original text. Taking the above point to the extreme, a very simple yet mildly effective “homeomorphic encryption” scheme would be to run the first layer(s) of the ML model on-device, run the majority of the model in cloud, then run the remainder of the model on the device. But then you leak a lot of information that can essentially be used to get back the original text. (Usually in this type of scheme, to defend against this, the provider of cloud services doesn’t have access to the full model, it’s been used before on vision applications involving medical data)
- adastra22 2mo ago
- randomImmigrant 2mo agoThis seems a fine tradeoff to me, depending on the context. There are datasets and operations on them where speed being sacrificed for privacy/security seems appropriate. Ideally, give me a dial, to ask for encrypted intelligence when I need it. Kind of like a private chat, but with deeper privacy protections.
- XorNot 2mo agoExcept that's for pathetically small datasets. What real datasets exist where this would be a worthwhile trade off versus simply owning the hardware? The numbers are so bad that underpowered local hardware would still beat it.
- inigyou 2mo agoyep I can beat 0.0008 tokens/s on GLM-5.2 on my CPU