6 ms·
I am still amazed by the amount of people that still mix security and encryption in-use with privacy. And truth is the whole privacy/security industry is doing
by v4dok 4y ago
I am still amazed by the amount of people that still mix security and encryption in-use with privacy. And truth is the whole privacy/security industry is doing nothing to change that.
Take this for example
"Valuable insights through AI (artificial intelligence), big data, and analytics can be extracted from data—even from multiple and different sources—all without exposing the data, secret decryption keys, or, if need be, the underlying evaluation code."
FHE gives you NO guarantee about the code that is running on the encrypted data. I can run a leaky AI model or a SELECT * on encrypted data and still get the output. What I can do (and that's assuming there is open-sourced, auditable code) is to make sure that anyone with hypervisor access on that machine cannot dump my data out during processing.
A very powerful concept for remote processing, supply chain security, and overall reducing trust; but completely unrelated to privacy.
- bawolff 4y ago> FHE gives you NO guarantee about the code that is running on the encrypted data. I can run a leaky AI model or a SELECT * on encrypted data and still get the output. What I can do (and that's assuming there is open-sourced, auditable code) is to make sure that anyone with hypervisor access on that machine cannot dump my data out during processing. I might be misunderstanding but i think this is misleading. Any code can be run, but the person running the code cannot see the results (or any side effects), so they cannot leak data.
- v4dok 4y agoDepends on what do you mean by "person running the code" If by person you mean the admin of the machine, then yes. If by person you mean the developer of the FHE-based application, then "maybe" If by person you mean the analyst who would in the end order an AI python model to be executed through the FHE-based software on a machine. Then no, that person will in the end get back human-readable results. Be that a model, or a table from an SQL DB running in FHE
- vesinisa 4y agoThe way I have understood FHE is that any algorithm that would operate on the data would, by definition, be unable to produce any result that was intelligible to anyone except the person holding the original key. At no point during the execution of an FHE algorithm is the data decrypted. The amazing thing is exactly that the code running on the data does not understand the data it is consuming nor the data that it is producing. Maybe someone who has actually studied homomorphic encryption can chime in.
- v4dok 4y agoThats true. But it will still produce some data, and that data will be viewed by someone eventually who owns the key to decrypt it. FHE tells you nothing about what this product should be. It could as well be a full copy of the original data. For example: I run an ML model using FHE on some data I shouldn't have access to in plaintext. The expected outcome of this workflow is a trained ML model on that data. FHE tells me nothing about the quality of this model. It could as well be an overfit model that spits out all the sensitive data.
- jxf 4y agoWouldn't doing the same thing without FHE also result in the same problem?
- v4dok 4y agoYeah, and many more. But I've seen multiple people argue that using FHE will magically solve all their privacy problems and its far from true. FHE (and similar technologies) solve a piece of that "puzzle" and most providers somehow gloss it over.
- deleted 4y ago[deleted]
- vesinisa 4y agoSorry but I still fail to see how that would be a problem, since the output of the program (e.g. the ML model parameters) would themselves not be intelligible to you. To make _any_ (non-cryptanalytical) inference on the plaintext of the homomorphically encrypted data _necessarily_ requires that the attacker at some points can access or execute some classical code on the plaintext. This would obviously violate the "fully" part of FHE. Edit: Okay so I might now understand you refer to a scenario where the user submits their data in homomorphic form to the cloud, where an AI model is trained on it. The AI model parameters are later returned to the user's device, which then decrypts them with the user's key and executes a classical model with those parameters, and then resubmits the user's data after processing with the said ML model (unencrypted) back to the cloud. It's true the user usually has no way of auditing the code / model that runs on their device, but isn't that rather easily alleviated by opening up the APIs for communicating with the cloud part of the service?
- m1ghtym0 4y agoI agree! That's why remote attestation or simply verifiability is such an important feature of these schemes. Semantic attestation ofc means having access to the source code and that IMO makes open source the natural choice. Audits might be an option where OSS is not desired for other reasons (likely business related). Not an expert on FHE, but confidential computing provides that attestation feature and it's just a matter of the software to make use of it.
- ketzu 4y ago> A very powerful concept for remote processing, supply chain security, and overall reducing trust; but completely unrelated to privacy. It is still related to privacy, but the privacy "attacker" is the execution place, allowing for outsourcing of computations and storage without running into data leaks or violations of data protection laws. Maybe you use a different definition of privacy?
- v4dok 4y agoYou are right. It is related to privacy, but its not the whole story. Thats what I am trying to say. Running FHE will not magically solve the "data leakage" problem of your AI models, and I believe that the people/companies who don't make that distinction are misleading.
- swores 4y agoThis isn't an area I know much about so I'll stay out of the main conversation, but would like to point out that considering you wrote "completely unrelated to privacy" in your first comment, following it with "It is related to privacy, but its not the whole story. Thats what I am trying to say." makes me unsure what your point is or if you actually mean or understand it. Sorry for being blunt.
- v4dok 4y agoYou are right to be blunt. I didn't want to define the separation of input and output privacy in a comment. In retrospect maybe I should have but I can't edit anymore. It is however common HN practice to pick out the words that suit someone and construct an very specific argument based on these words, often missing the spirit of the comment. Yes, when I wrote the comment I had in mind "output" privacy while FHE is dealing with "input" privacy. It is related to privacy, but not in the way most people think about it. If you go to a random person and ask them about privacy they will not think about the threat model of a cloud provider leaking their data, but they will think of the thread model of a pharma company knowing exactly what drug they bought and when. That notion of privacy is not covered by FHE(alone). And even the first notion of privacy is covered only if the FHE program has a way to attest itself so you know that what you expect to run is indeed what is running.
- carrotcypher 4y ago> FHE gives you NO guarantee about the code that is running on the encrypted data. You mean in order to validate the data is authentic? Otherwise the code running on the data is irrelevant, as it can't access the data itself (and thus preserves the before-mentioned privacy).
- coldtea 4y agoThe code decrypts the data at some point (for use, presentation, etc). If the code is crappy, insecure, etc. then the data will be exposed, and the data being encrypted wont help at all...
- carrotcypher 4y agoDecrypted on the local computer, not the untrusted remote computer as it were.
- deleted 4y ago[deleted]
- segfaultbuserr 4y ago> FHE gives you NO guarantee about the code that is running on the encrypted data. I'm open to correction, but it's my understanding that the strongest form of FHE allow users to submit an encrypted executable with embedded data as input, which is then processed by an untrusted server. I'd definitely call it an ultimate form of privacy. The computational cost is prohibitively expensive, and conditional branch is impossible in the standard implementation, so it's largely an academic exercise. But last time someone on HN told me currently the achievable performance on a modern computer is roughly equivalent to a 1970s mainframe, so I guess some niche applications are still possible. Weaker forms of FHE don't have this level of privacy, and they do not claim so. Nevertheless, relevant development still represents progress on cryptography and privacy researches as a whole.
- matthewdgreen 4y agoIn the context of cryptographic protocols we sometimes use "privacy" to refer to the notion of "confidentiality". The latter is, I think, a cleaner word that avoids the collision with human notions of privacy. In this case the real danger is that the availability of "privacy-preserving technologies" like FHE, MPC and Differential Privacy will actually do more to undermine human privacy than all the non-confidential tech that come before. This will mostly occur by allowing corporations to build sophisticated statistical/ML models using data that would previously never have been allowed out of its confidential silo.
- hansvm 4y ago> run a leaky AI model or SELECT * That's the point though isn't it? Only the person who wants the results can get them or even see the inputs. That restricts the data available for shitty AI and precludes any Joe Schmo from scanning the whole database. If your threat model is instead that you don't trust the FHE endpoint, then much how you want HTTPS termination to happen in a place you control you also in this case just encrypt the stuff you care about on your own devices.
- y7 4y agoSecure multiparty computation (MPC) does help here. Suppose you have multiple organizations that want to run some computation on their joint data, without revealing their data to each other. Each organization has their own machine that runs the MPC protocol. They have full control over their machine, and can inspect that the code correctly executes the protocol. Only once all organizations agree, will the computation take place, and within the security model of the protocol, it is guaranteed that only the correct computation output is revealed to the designated parties.
- LawTalkingGuy 4y agoFully homomorphic encryption is a toolset, it's not a specific configuration. Your scenario has these parties, 1) a patient whose data we're discussing, 2) the hospital they shared it with, and 3) a pharma company looking to use the data. The hospital wants to promote this use without leaking any PII. You're right that the hospital has no idea about the queries ("the code") but they control the server and which messages it will send in response. As you point out, the hospital wouldn't run a FHE database capable of full-text extraction specifically because that would amount to simply sending all the data to the pharma company. Instead they'd run a specialized FHE-DB server which would, for instance, return only row counts. The pharma company would run secret queries and if the hospital had one or more patients who matched the query the pharma company would know to the contact the hospital and then once paperwork is signed they could rerun the query with a signed token from the hospital and finally the query would return the actual PII.