6 ms·
My master's thesis is on a topic in this field (Privacy Preserving ML) and from my understanding HE and other techniques have very high overheads(~10^3) on infe
by sabretooth1405 2mo ago
My master's thesis is on a topic in this field (Privacy Preserving ML) and from my understanding HE and other techniques have very high overheads(~10^3) on inference tasks and thus aren't very commercially viable.
- clayhacks 2mo agoDo you think that’s like a fundamental limit or something that will improve with time and new algorithms?
- bhu8 2mo agoI wouldn't be very bullish. Homomorphic encryption got significantly efficient with the first few iterations, but I don't really see the necessary orders of magnitude savings coming soon. You could reduce this by some partial encryption schemes (e.g., for LLMs you need a handful of basic operations) but a better alternative already exists: multi-party computation. Source: I did research in this area in the past.
- dietr1ch 2mo agoExactly my concern, and worse overhead that what I recalled. Cost-wise the only viable private compute is local compute. It's more expensive than cloud, but true private compute in the cloud is definitely pricier.
- abetusk 2mo ago1000x slowdown is bad but not a complete deal breaker. Do you have a sense of what a reasonable achievable factor is? Do you have sense for how long before we get to that achievable factor?
- jacquesm 2mo agoIt's a ridiculous waste of energy, just use local compute.
- joquarky 2mo agoGood point, but if we can advance research on this via the AI bubble, it could improve privacy in other areas.
- furkanturan 2mo agoLocal compute is preferable where possible. There are cases where computation needs to be performed remotely. For example, when collecting data from remote entities while preserving privacy by allowing each entity to retain ownership of the encryption keys used to protect its data. At Belfort, we are exploring such applications, such as - https://belfortlabs.com/blog/belfort-partners-with-lg-on-encrypted-advertising-recommendations https://belfortlabs.com/blog/belfort-partners-with-lg-on-enc... - https://belfortlabs.com/blog/encrypted-fraud-detection-with-swift https://belfortlabs.com/blog/encrypted-fraud-detection-with-...
- jacquesm 2mo agoFor me it is the simplest reasons of all: privacy / confidentiality. There is absolutely no way any of this data leaves my systems.
- jerf 2mo agoThe article conspicuously fails to go into much detail about that. I poked around with an AI a bit (to rapidly cover all the linked pages) and it seems the best numbers we can get are from this arxiv paper: https://arxiv.org/html/2506.18150v4 https://arxiv.org/html/2506.18150v4 Which says: "We evaluate HE-LRM on UCI (health prediction) and Criteo (click prediction), achieving inference latencies of 24 seconds on UCI and 228 to 489 seconds, respectively, on a single-threaded CPU." There don't seem to be any direct comparisons available, probably because nobody else has any reason to limit themselves to one single-threaded CPU with normal techniques, but for reference the AI seems to expect that normal times for conventional setups are in the milliseconds range, fairly comfortably, even on CPU. I didn't find a clean primary source to link to for this claim, but clicking through various things that don't cleanly state the situation it did seem plausible. So we seem to still be in the range of single-digit orders of magnitude slower, possibly as much as 5 or 6, which is to say, we're still talking the range where we need to take the log of the difference to get sensible numbers, we're not using percentages. (To run it yourself, I basically just fed the URL from the HN link, mentioned that FHE is known to be slow, and asked if anything linked in the blog post gave concrete times.)
- j2kun 2mo agoThe linked repository has demos you can run (though you have to install bazel), and some of the smaller models run inference in about a second, while the larger ones take minutes. That said, there is a lot of ongoing work on GPU acceleration. Cf. the recent FHE-based CIFAR demo that runs in 200ms: https://sofar.belfortlabs.cloud/ https://sofar.belfortlabs.cloud/ Still maybe 1000x slower than cleartext, but progress!
- laksjd 2mo agoWith GPU acceleration, that 228 to 489s becomes a fraction of a second :) https://belfortlabs.com/blog/belfort-partners-with-lg-on-encrypted-advertising-recommendations#:~:text=We%20achieve%20a%20400x%20speed-up%2C%20bringing%20the%20latency%20of%20a%20recommendation%20down%20to%200.56s https://belfortlabs.com/blog/belfort-partners-with-lg-on-enc...
- u1hcw9nx 2mo agoThat's the reason for HEIR like optimization and parameter selection. It narrows the 10^3 - 10^6 penalty to 10x - 100x.
- michaelmrose 2mo agoWhich seems massively worse than a real local device in fact 2x is probably untenable to the point of uselessness because actually privacy sensitive matters need actual privacy that can't be defeated by your government telling Google to serve you compromised js and spy on you anyway and most people don't give 2 shits about privacy so they won't pay 10% more let alone 2x. I'm glad people fund things that are only of interest to nerds but this will never be useful.
- u1hcw9nx 2mo agoI think you have completely wrong use cases in mind. You will not use this for normal compute workloads. Typical use cases are for doing biometric authentication without giving your biometric information, or sensitive queries using medical information. Apple has homomorphic encryption in image search. You can use your own photos encrypted into the cloud to search for landmarks in the image without revealing photos. People can also coordinate and compare information without sharing sensitive data.
- jgerrish 2mo agoBut for some fields, almost EVERYTHING they touch is sensitive. I've been reviewing PPML for wildlife management purposes, which my brother is involved in and my other brother, a ML expert, may want to be involved in. With wildlife management, you're dealing with health issues, like rabies outbreaks. That requires privacy. You want to preserve customer confidentially because it's often embarassing. And private property cameras and sensors can leak information about private citizens or kids in a neighborhood without appropriate social and technical protections. There are hundreds of fields like this. Not just healthcare and policing. We are likely to see a lot of the hardware required for this to move out into space data centers for batch jobs at least. And along with that calls for reduced RF and light pollution like StarLink. And that is going to be helped by a large number of angry liberal citizens who are being riled up about data centers. And that anti-tech rhetoric is already leading to violent responses and debate. Which hurts the liberal cause for universal healthcare. Conservatives see angry liberal anti-tech actions and tarnish calls for healthcare reform and other liberal causes. The people who are going to benefit the most from cheaper PPML in orbiting data centers are in many ways making it harder for the rest of us. It's not a small issue, and I wish I had had the reputation to reduce the anger. It seems unrelated to PPML. But PPML is a clever political gas pedal to get space control.
- bevekspldnw 2mo agoCommercially viable for Google boils down to can they attribute ads behaviors to it or not. Then there’s a second tier of things that just make those wheels turn and if they do or don’t make ads revenue is nominally immaterial. The teams doing this stuff at Google are purely for show, none of this makes it into any real products. There’s the narrow exception of stuff like gboard, that does use privacy preserving ML/fed learning, but this stuff isn’t in the same zone. I find it a bit embarrassing when Google publishes this stuff to be honest.
- luckydata 2mo agoYou are very wrong about all of this btw.
- deleted 2mo ago[deleted]
- bevekspldnw 2mo agoYou literally don’t know who I am or the roles I had. So unless you can tell me how many steps you were from Kent Walker and what you worked on I’m gonna bet a hell of a lot I know more than you. Edit to clarify my prior point: some of the technology makes it into the product, but the putative data protections do not. Why? Because there is always a work around, and ads legal will approve it every time.
- bitpush 2mo ago> You literally don’t know who I am or the roles I had. I'm now curious. Who are you?
- bevekspldnw 2mo agoSaying that would make it unable for me to use HN as certain companies monitor my social media comments.
- luckydata 2mo ago
- dhx 2mo agoTo throw out some real and up-to-date numbers from [1] for FHE at "128-bit security level", to sort 8x 8-bit unsigned integers on the most ordinary of desktop PCs, wait 3 seconds for the result. Want to sort 32x 8-bit unsigned integers instead? Come back 34 seconds later for the result. update: also see [2] for some primitive unsigned 64-bit integer operation benchmarks with the TFHE-rs library (winner in the sorting performance comparison of [1]). Equality at 80ms, addition and subtraction at 100ms, division at 8 seconds, etc. [1] https://eprint.iacr.org/2026/1495.pdf https://eprint.iacr.org/2026/1495.pdf Oblivious Sorting under Fully Homomorphic Encryption: A Comprehensive Survey and Performance Analysis, Omar Ahmed and Rostin Shokri and Nektarios Georgios Tsoutsos, 2026 [2] https://docs.zama.org/tfhe-rs/tfhe-rs/1.0/get-started/benchmarks https://docs.zama.org/tfhe-rs/tfhe-rs/1.0/get-started/benchm...
- pamcake 2mo agoBenchmarking code in repo: https://github.com/google/heir/tree/main/benchmark https://github.com/google/heir/tree/main/benchmark Project intro talk from 2023: https://www.youtube.com/watch?v=kqDFdKUTNA4 https://www.youtube.com/watch?v=kqDFdKUTNA4
- j2kun 2mo agoThe first link has a misleading name. Instead use these two links for a better picture: https://github.com/google/fully-homomorphic-encryption/tree/main/demos https://github.com/google/fully-homomorphic-encryption/tree/... https://fhe-benchmarking.github.io/ https://fhe-benchmarking.github.io/
- tbenst 2mo agoThat is sobering for sure, I wonder what the theoretical bounds are on what is possible if known. Would be such a dream to use a Frontier LLM one day with homomorphic encryption, but this sounds wildly implausible based on where things are today.
- joshspankit 2mo agoWe’re not even close to the limits of AI optimization so finding the theoretical bounds is going to have to wait
- Fordec 2mo agoThe primary path to speed ups appear to be in custom ASICs by startups like Niobium. Combined with the recent Taalas acquisition by AMD, I think I see where this is going. But yeah, for hot path traffic it's probably going to be swamped by the input data rate. But I expected identity tables and cached lookup data will need to be a core component so duplicate checks is avoided in every way available.
- furkanturan 2mo agoCount Belfort too. In addition to our GPU acceleration efforts, we have ongoing ASIC initiatives to further accelerate encrypted compute.
- fragmede 2mo agoThe one that I'm waiting for is a women's period tracking app that uses FHE on the backend to be fully private.
- glaslong 2mo agoSaving your comment for the Weekend Project idea backlog, if you don't mind :)
- lupire 2mo agoWhy on earth do you need a backend for this? The backend only exists because it leaks the data.
- fragmede 2mo agobecause you lose your phone and don't have access to the account anymore.
- xg15 2mo agoThen the app can make a bog-standard encrypted-at-rest backup to somewhere and make all the computations on the device on the cleartext data. I don't see the need to do computations on the encrypted data here, which is what FHE would provide in addition to traditional encryption. > and don't have access to the account anymore. This would be trouble with or without FHE. Even if the backend wouldn't need to decrypt the data, the user will - so as soon as you actually want to show something in the app, you have the same key management problems as without FHE.
- fragmede 2mo agoOkay, so the platform becomes valuable to its users when it's able to suggest things like "based on millions of users, people with cycles like yours typically ovulate around day 16." In order to do the data mining in order to make those kinds of claims, traditionally you'd need to have access to the data. As you point out, encrypted-at-rest is solved. But what about when it's not at rest? In-use and in-transit is when FHE kicks in. Sure, you could just do it locally, but then you miss out on the aggregate data mining. Not for advertisers, but because it helps women with their bodies. The compelling product claim is "we literally cannot read your period data." Not "we pinky swear not to" but "we actually really really actually can't!"
- mrcwinn 2mo agoSounds like the start of every journey. With all respect to your thesis, I probably put my chips on Google's research and security teams.
- elgertam 2mo agoI saw a paper about this in early 2020 (pre-COVID shutdowns) at the ScaledML conference. I looked into it and had the same conclusions. At some point, running your own models in the clear is just more practical.
- Jabrov 2mo agoIt might still be useful for classification usecases
- petters 2mo agoOnly 1000x overhead would make some image classification tasks go from 1ms to 1s. That’s viable for some applications!
- furkanturan 2mo agoExactly. That is what we have today at Belfort; not enough for making all AI work privacy preserving, but fast enough for many applications, where otherwise unencrypted compute is not acceptable.
- monster_truck 2mo agoOne of the biggest problems IMHO is that they aren't trying to usefully accelerate it on anything other than specialty hardware or 64+ core EYPCs so nobody gets to play with it at home. ex: A 7900XTX barely gets 0.5 TOPS of u/i64 naively w/ hip-direct, 5-10s just to bootstrap! I needed more throughput for non-crypto i64 diff eqs so I slopped up a lib that uses RNS & CRT w/ Int8 GEMM... it's good for ~3.9 TOPS (~90% theoretical peak of the RDNA3) at prod relevant FHE sizes (2048/4096). This lowers bootstrap time to 200-500ms. It was basically free real estate lol It isn't done yet (not worth the heat in the summer), going to finish it in the fall. Have been accumulating cloud credits to do CDNA3/4 validation in the meantime (If anyone has some to offer do let me know!) It's neat but very dry, uses semantic contracts so you tell it what kind of mult you need and it chooses the validated best backend. If you're doing lots of smaller ops (512, 1024) it will use custom WMMA/MFMA kernels, dual issue, and grouped dispatch to land >70x over hip-direct. https://github.com/doublemover/RNS8/ https://github.com/doublemover/RNS8/
- therealmarv 2mo agoDo you mind sharing your master's thesis? oO Would be interesting to read it (and no judgement!)