10 ms·
The Apple Psi System [pdf]
- deleted 5y ago[deleted]
- gjsman-1000 5y agoTopic should be renamed the Apple PSI System or Apple Private Set Intersection System (not Psi or something related to tire pressure.)
- floatingatoll 5y agoThe HN style guide lowercases acronyms like PSI to Psi, unfortunately.
- joshschreuder 5y agoSource? I can think of a number that aren’t, and a bunch currently on the frontpage (as specific as “CSAM”, and as widely used as “IP”).
- easton 5y agoHN automatically converts to titlecase on submit, if you click edit and change the title to how you want it won’t mess with it again.
- floatingatoll 5y ago(Mods still might, if they’re in the mood.)
- moonchild 5y agoTire pressure is measured in PSI, or Pounds per Square Inch.
- gjsman-1000 5y agoAlso of potential interest besides the OP link, there is also the paper "A Concrete-Security Analysis of the Apple PSI Protocol", also known as the "Alternative Security Proof of the Apple PSI System." https://www.apple.com/child-safety/pdf/Alternative_Security_Proof_of_Apple_PSI_System_Mihir_Bellare.pdf https://www.apple.com/child-safety/pdf/Alternative_Security_... Basically it's a second opinion on the mathematics from a different perspective. The original post link is the formal proof by Apple employees and Stanford, the Alternative Proof is by the University of California.
- tptacek 5y agoFor clarity: the "official" proof has Dan Boneh's name on it, and the "alternative" proof has Mihir Bellare's name on it.
- asymmetric 5y agoWhat does this clarify?
- tptacek 5y agoThat Dan Boneh and Mihir Bellare worked on proofs of this.
- josephcsible 5y ago> Privacy for the server: A malicious client should learn nothing about the server’s dataset X ⊆ U other than its size. In particular, it is important that the client learn nothing about the intersection size |id(Y̅ ∩ X)|. Otherwise, the client can use that to extract information about X by adding test items to its list Y̅, and checking if the intersection size changes. If one of your pictures is a false positive hash collision, you'll have no idea until your front door gets broken down. > Privacy for the client: Let X be the server’s input from which pdata is derived. A malicious server must learn nothing about the client’s Y̅ beyond the output of ftPSIAD with respect to this set X. Apple can't check whether a hash match is a false positive or not, because they only get the matching hashes and not the pictures that triggered them. So if you have a bunch of false positives, your front door is getting broken down, with no opportunity for a human to realize the problem and intervene. > The protocol need not provide correct output against a malicious client. That is, the protocol need not prevent a malicious client from causing the server to obtain an incorrect ftPSI-AD output when the protocol terminates. The reason for this is that a malicious client can always choose to hide some of its data from the PSI system in order to cause an undercount of the intersection. Their protocol isn't (and can't be) secure against the one attack that the people that this system is supposed to catch would actually commit. > Moreover, a malicious client that attempts to cause an overcount of the intersection will be detected by mechanisms outside of the cryptographic protocol. This seems eerie to me but I can't put my finger on why.
- gjsman-1000 5y ago"So if you have a bunch of false positives, your front door is getting broken down, with no opportunity for a human to realize the problem and intervene." Well, technically speaking, once enough security vouchers have been submitted and reached the threshold as stated, a report will be sent to an Apple employee (somewhere). The vouchers will then, combined, be decryptable and contain grayscale low-res versions of the original image for confirmation, in which case the NCEMC (National Center for Exploited and Missing Children) will be alerted and law enforcement. I'm just pointing out that your door being broke down is by multiple false flags and a human will get a chance to "realize the problem and intervene" before it goes to the FBI or whatever. Not saying I like this system, just making a nitpick on your criticism. I don't really know how else you could have a human intervene without "breaking down your door."
- sydthrowaway 5y agoHow to learn enough crypto to understand this paper?
- shuckles 5y agoDan Boneh’s CS255 course reader and CS355 lecture notes should be a good beginning.
- nullc 5y agoThe actual scheme is relatively simple, the exact details and terminology make it hard for a layman to understand. I'll give you an EL18 description of a basic private set intersection: I have a database of image fingerprints which I want you to test your images against and tell me if there are matches. I can assume that you're going to faithfully run my protocol because I use DRM to control the software that runs on your computing device. The obvious way to accomplish the matching would be for me to just send you the database of fingerprints-- they are hashes after all and don't tell you anything about the images other than letting you match them-- and for you to tell me about the matches. But I don't want to tell you the database or tell you when an image matches because if I do you'll realize that I'm targeting images connected with a particular ethnicity, which I intend to mass murder. So, instead I tell you I want to search for child porn and I get you to agree to the following protocol, and you're foolish enough let me keep the hashes secret for no obvious reason. The first building block we need is an encryption scheme which is additively homomorphic, such as elgamal encryption. With this special encryption scheme the following properties hold: Enc(Data1, key) + Enc(Data2, key) = Enc(Data1+Data2,key) and x*Enc(Data1,key) = Enc(Data1*x,key). Or, in English, the sum of two ciphertexts gives you a ciphertext of the sum of the plaintexts, and a ciphertext multiplied by a value gives you a ciphertext for the plaintext multiplied by that value. With that in hand we can build a private set intersection. (1) I pick a private key, send you the public key, and I encrypt each of the hashes in my database. I send you the encryptions-- which, thanks to the encryption, teaches you nothing about the database except an upper bound on its size. (2) For each database entry you take the hash of image you want to test, encrypt it with the same key and subtract it from the encrypted database entry. If they matched you have an encryption of zero (which you can't tell is zero, due to encryption), if they didn't match you have an encryption of some non-zero value-- the difference between the image hash and the database entry. You then pick a new random number and multiply the result with it. You now either have an encryption of a totally random number (if there was no match) or an encryption of zero (since random*0=0). (3) You send that to me, I decrypt it... and if it decrypts to zero I add you to the list of people to be executed at some time in the future. If it doesn't decrypt to zero I learn absolutely nothing about your hash, other than it didn't match, because the result is literally a random number. The Apple scheme makes a number of elaborations on this basic idea to improve efficiency (the database is a cuckoo hash table, so instead of sending you one encryption per database entry per image of mine I only need to send you a few encryption per image-- however much fanout the hash table has), to make it so that matches result in leaking a decryption key so they can decrypt the image, and additional complexity to make it so that the matching isn't fully triggered unless you have more than some threshold number of matching images and to partially obscure the exact number of sub-threshold matches.
- ljhsiung 5y agoThanks for the read. Pretty cool. The following is my layman's understanding, mostly because it helps me to understand things by typing. Please criticize if I miss stuff. Noteworthy is Dan Boneh's contribution on this. Given his reputation in the crypto community, it seems he really does believe in this, despite all the controversy as of recent. They discuss and build upon several security properties, but the most contentious point-- privacy/leakage/scanning of your photos-- is addressed as follows: Firstly some background-- the private set intersection[1] (PSI) technique in general permits 2 sets to be compared with both parties learning *only* the intersections. In a nut shell, Apple uses this concept such that, if the number of intersecting elements is greater than a threshhold, they're notified. There are several modifications (Shamir's secret sharing, Cuckoo tables, PKI-ish schemes) to create what Apple calls a ftPSI-AD protocol to optimize desired properties-- performance, integrity, and most notably to me, zero false passes. That is, innocent people will be minimized at the cost of real child-pornographic images slipping by. Couple noteworthy things that still raise red-flags-- 1) they prove that, for honest servers <--> malicious clients, and vice-versa, privacy is not violated for either party, but to me this considers the client as the phone and Apple the server. I'd argue that the phone is actually "Mallory", and you are the client. You might be honest, but how do you trust the Apple + phone short of reverse engineering it? This is the biggest hole to me, and so I don't fully understand this proof (or, perhaps I have the parties mixed up). 2) Several things are handwaved and/or left "variable to implementation". E.g. Section 5, on "near-duplicate images" that may count twice to this threshhold-- >> Several solutions to this were considered, but ultimately, this issue is addressed by a mechanism outside of the cryptographic protocol What the?? Hello?? Perhaps this is addressed in another whitepaper, given this is a theory/protocol heavy paper, but this does not instill confidence. Or, take this bit from remark 3-- >> If needed, these false negatives can be eliminated with a tweak to the data structure used Uh, I thought not sending innocent people to jail was a pretty critical property. You're telling me the server/Apple, who controls the Cuckoo table, can just change this on a whim? How would I hold them responsible/be notified of this? These "variations" are remarked on several times in the paper. Again, not exactly confidence building. Overall, while I really applaud this effort, and I'm not as outraged as I initially was, I'm only slightly less so and have a handful of more questions than before. Again, please correct me if my annoyance might be misguided, given these technical details. [1]: https://en.wikipedia.org/wiki/Private_set_intersection https://en.wikipedia.org/wiki/Private_set_intersection
- stefan_ 5y agoWhere is the cryptographic proof that the only thing you are scanning for is CSAM? This is just nice window dressing.
- 015a 5y agoIts ok, because they're comparing against "a database of known CSAM image hashes provided by NCMEC" You can trust them. You have to trust them. Oh, "provided by NCMEC and other child-safety organizations." [1] Other unnamed organizations. Don't worry about who they are, its not relevant. Stop thinking so hard. Stop Screeching. Trust Apple. [1] https://www.apple.com/child-safety/pdf/CSAM_Detection_Technical_Summary.pdf https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
- YokoSix 5y agoThat's the big point they don't address. Apple can't possibly check every hash of every file some obscure organization is sending to their database, especially in foreign countries. If some men in black walk into the offices of some nice little child protecting service in Ohio and demand that they put some additional hashes into the database because otherwise something bad could happen to them or their families, does Apple really think they would decline? It doesn't matter how secure the system is, the vulnerability is that virtually anyone can input virtually anything into the database without Apple even knowing and expose selected users that way. This failure point is now at the heart of iOS and macOS and it's baffling that Apple doesn't see that or doesn't want to see it. My guess is that they're somehow forced to implement this and try to talk their way out of it with some strange PR pieces which are only convincing their most naive users.
- sennight 5y agoSomething weird is definitely going on. I'd normally chalk it up to a powermad hubris that seems to be common among technocrats, but this is the latest example of a corporation voluntarily providing politically motivated non-profits the ability to control their platforms. This has been going on for a while on financial networks - in the pursuit of the elusive, but totally real, neonazi resurgence. Paypal recently announced such a "partnership". Microsoft has allowed some UK based group to censor their search results for a long time now. That was originally sold under the banner of "bbbut the children!", but the only reason I know about it is because they started blacklisting sites that archived video related to the war in Syria.
- xucheng 5y agoWhile it is good that Apple seeks formal security analysis over its proposed system, there are some fundamental security assumptions baked in the system: - It assumes that the server would not tamper its dataset (i,e., the list of CSAM). So it is OK to disclose the information if the client has enough matchings. But in reality, nothing prevents a malicious server adding arbitrary content to the list. - It fails to consider the vulnerabilities of the perceptual hash. This includes false positives and adversarial collision attacks (https://arxiv.org/abs/2011.09473 https://arxiv.org/abs/2011.09473). Another potential long term issue is that it is unclear how long Apple will store the safety vouchers. As a storage service, Apple may store them forever. The system is based on Elliptic-curve cryptography. Despite it is the current state-of-the-art encryption technique, it will be broken when the quantum computer becomes a reality in the future. So it is possible that every encrypted safety vouchers can be decrypted in the next 50 years.
- joshschreuder 5y agoSome apps like WhatsApp save images from messages automatically to the camera roll. So a theoretical attacker could send you a message on WA with a set of adversarial collision images (not even child pornography potentially), and effectively SWAT you
- ezconnect 5y agoAnd everyone will be on the list as pedophiles.
- selsta 5y agoThis is already possible on both iOS (iCloud Photos) and Android (Google Photos) as both scan for CSAM server side.
- dmitrygr 5y agoGoogle photos does not by default upload photos in WhatsApp folder till you enable it, default is no, so no...
- markrobin 5y agoneed enough crypto understand
- doorknobs 5y agoThis document is a deflection from the main concern -- while it's commendable that they took the effort to prove the cryptographic properties of their system, what good is that when your hash database is a government-controlled black box? Can't exactly publish a white paper on that, huh?
- drenvuk 5y agoThis is incredibly annoying. They're providing all of this information which says that this program that runs on our devices is incredibly safe if we're not bad people. That the people at Apple and then law enforcement don't see anything about our photos until there is a human set threshold hit by a human programmed algorithm. Technically, this should be sound since the proof has checked out with multiple cryptography experts. But none of that matters. The Apple PSI System is spyware. They are providing all of this info to justify putting spyware on our devices. They are attempting to put spyware on our devices to see if we can be sent to jail. That's all that matters. That is the end effect. Apple is justifying putting SPYWARE ON ALL OF THEIR PHONES. Any discussion of the technical merits of a SPYWARE system implemented against you is missing the point. It should not exist.
- liamcardenas 5y agoWe need to draw the line somewhere. Maybe this is ok. But where does it end? What if we have a Neuralink-like device? Is there ever a line too far?
- akersten 5y agoThis is the line and it's not ok. I am astounded that anyone here is entertaining the idea that it is somehow ok. Take out the "CSAM" noun from the equation - it really poisons the argument with emotion. "Apple devices check your files against government blacklist" is the headline, and not enough of us are saying "no."
- gruturo 5y agoThe line is here. And this is absolutely not OK. It is ripe for abuse by governments, which Apple absolutely has a history of bowing to. Saudi Arabia WILL use this to track dissidents and homosexuals. China WILL use this for whatever they need, on a daily basis. Hell even the US will use this, I'm sure. The only way this would be OK was if the input data (the CSAM database) was cryptographically signed and no government, no single entity and not even Apple, could change the content, with the only signing key split in 10 parts to be held personally by Bruce Schneier, the Pope, Linus Torvalds, the Orthodox Patriarch, Keanu Reeves, a couple head Rabbis and a few of whatever their equivalent in Islam is, and they had to personally review the images one by one and certify that perceptual hash 0FB89C8A7DF6AA1945B is indeed CSAM content and agree to collectively sign it for addition. This would only work if the full IOS source was fully published, and compiled as a reproducible build, so everyone could confirm the scanning code does what it's supposed to, and is not altered with any subsequent update. P.S. Don't nitpick on the names, it was a deliberately absurd list of people either with a good reputation or with a lot to lose in the respectively chosen afterlife.
- madmod 5y agoYeah what happens when someone abuses this for easy “untraceable” swatting. What if the next iOS malware uploaded some CSAM hashes to iCloud unless you pay 5 btc in the next 24 hours?
- r00fus 5y agoThere's also the other side - assuming NCMEC and law enforcement are good actors is a flawed premise (as has been proven time and again). What if the CSAM database has non-CSAM images there, or non-CSAM hashes? One could argue which is more likely, but the fact remains that the entire premise of this system is flawed and it seems the only way to play this game is to not use iCloud photos at all (though it's possible that malware could bypass that and turn on iCloud sync as well as upload the photos too...)
- ezconnect 5y agoThis is to facilitate the easy approval of court orders for the spy agencies to get the content of someones phone. It doesn't matter if the photo is illegal or not.
- nullc 5y agoOne theory I've seen advanced why the claimed reports to the NCMEC are many orders of magnitude larger than the prosecutions is that the conjecture that the whole thing is a CIA operation to collect kompromat on massive numbers of people. The alternative that the database is just stuffed full of non-illegal images seems more likely but I dunno if its that much more comforting.
- selsta 5y agoThis is already possible today on both iCloud Photos and Google Photos. Both scan for CSAM server side.
- darkhorn 5y agoCan they scan for face recognition? For specific faces?
- eknkc 5y agoThis is a way worse of a spyware than I initially thought. - They have a database of file hashes. - They can’t validate its contents by design. - They can’t explain who supplies it in detail. - The suppliers of these data work with / close to the US gov. - The match your own files on your own device against this database that contains who knows what. - When it matches a couple of times, they alert the authorities. - There is zero fucking visibility for both Apple and you by design and this is a very good thing for Apple. - I assume the source of this data can update / add new hashes in time and your device will happily comply. And their only concern is to say that the algorithm is so perfect, they can not see what’s happening and there won’t be false positives (hopefully). You know what, I trust your ability to do it properly. No need to explain more. Problem is that the very thing you are building is fucked up by design. And obviously Apple does not address it in any way. But think about the children…
- somedude895 5y agoSo if I understand correctly, they'll have a db of hashes which are probably tagged with some info on what's in the picture? Who is supplying the hashes? It seems like a LOT of power will be concentrated there. They can basically decide who gets in trouble with seemingly no accountability whatsoever?
- rahkiin 5y agoThis already happens server-side now so what you mention here does not actually change.
- eknkc 5y agoNot on iCloud. They have all the content already encrypted on your phone. Can’t fingerprint after that. Them adding this as an on device process is due to that. And the fact that Apple has been talking about privacy all this time (and rightly so) means their past mitigations makes this more painful. Any other vendor could just do it on the server, add a line to their Privacy Policy and nobody would know.
- elisharobinson 5y agoshow me the code or STFU
- cgio 5y agoWhat are the storage and processing requirements for this functionality and how will they evolve over time? Can it be so that it is only loaded on my device if I use the relevant services? Other than impact on privacy it also has impact on property even though I know with Apple that ship has sailed a long time ago. My last Apple purchase was a couple days before I heard about this, and it will be my Last. Not like they’ll care, but I do, as I only accepted their walled garden in exchange for privacy, however naive that proves me to be…
- atoav 5y agoYou don't own your apple devices anyways. It is a happy walled garden for those who don't care too much, those who lost all hope and those who can lie to themselves very well. Not that I am a open source extremist, but the moment where we couldn't control the way our own machines run is the moment where they stopped belonging to us. Supporting a true open source phone OS might be a good idea even if you don't use it. Because one day you might have to.
- dmitrygr 5y agoAssuming perfectly trustworthy governments and perfectly flawless programmers, it is all good. Now where do I get me some of those in this imperfect world?
- supernes 5y agoThe crypto seems sound, but a lot of the tricky questions are waived aside with the blanket statement of "mechanisms outside of the cryptographic protocol." The part that worries me the most is that no one outside of Apple can verify that the hash set they're pushing hasn't been tampered with (Section 4, Remark 5). This allows them for example to add leaked product image hashes to hunt down and prosecute people who share info about their products before release. In fact, the system seems designed to be impossible to audit, with only a subset of the whole hash set being pushed to clients, so that researchers can't even tell when more hashes have been added. As a consequence of that design, they acknowledge that a "small number" of false negatives will be missed, and justify that with an argument that it improves performance (Section 2, Remark 3) False positives on the other hand will be common (as detailed in Section 5, "Duplicate images") - simply copying a file on two client devices that don't share a cloud owner ID will count towards the threshold, and again fall back on "mechanisms outside of the cryptographic protocol". And last but not least, let's spare a thought for the Apple employees that will be required to sift through potentially traumatizing imagery (assuming the company doesn't outsource that to a third party.)
- nojito 5y agoThis is absolutely earth shattering work in the world of differential privacy. Kudos to this research team for spending the time to work it out and publish. Just wow.
- Renaud 5y agoI'm sure the crypto part is well implemented. However, I make the prognostic that within the next 2 years, the Chinese government will force Apple to use its own database of "objectionable" content and will require that the on-device photo roll be scanned, not just iCloud (they already have access to that). And probably not just China, because any LEA and secret services would love to have the ability to use such a system for ad-hoc searches: dear Apple, for those users, please extend the database source to this new one we maintain and enable scanning of all on-device pictures.
- nullc 5y ago> I'm sure the crypto part is well implemented. If it is, that's bad for human rights. The crypto part protects Apple and their under-specified list sources from accountability. Without the crypto the system would just send you the database of bad hashes and snitch on you based on it. Your privacy would be no worse, but at least researchers would have some chance of detecting when the system was being used off-label to enable genocide.
- doe88 5y agoAll other flaws asides, that the hashes are not auditables by users on their own computers is deeply wrong and undermine the trust on any algorithm as sound as it could be. Moreover, this blackbox provides Apple plausible deniability if something turns ugly, they will say, we're just the middle man, we only relay the officialy approved database, we can't vet it.
- robertoandred 5y agoThey do vet it. They manually review every flagged account before the process goes further.
- nullc 5y agoYou misunderstand the purpose of the review. Per current case law if apple automatically handed the data over rather than reviewing it themselves the reading of the report would constitute an unlawful search. A human at Apple needs to inspect it in order to quash your fourth amendment protection against a warrantless search. Apple inspecting your images destroys your privacy, the extra step of review doesn't protect your privacy -- it's there to destroy the constitutional protection of your privacy.
- unstatusthequo 5y agoTime to notify the attorneys general. I can tell you personally that plaintiff firms are already gearing up to launch class actions against Apple on this. If you oppose it, notify your AG as well. HN Link: https://www.naag.org/find-my-ag/ https://www.naag.org/find-my-ag/
- naasking 5y agoOthers have done a good job addressing the political and social issues with this sort of endeavour, but I'm puzzled by something else. Isn't this just trivially bypassed? Just changing a few bits in the image won't be noticed by the human eye but the hash will be wildly different.
- mkl 5y agoYou're missing the fact that it uses perceptual hashing, not cryptographic hashing. Minor changes to an image result in a similar or identical perceptual hash. Here's an article and a big discussion about it from the other day: https://news.ycombinator.com/item?id=28091750 https://news.ycombinator.com/item?id=28091750
- naasking 5y agoGood to know, thanks for the link!
- kfprt 5y agoPolice departments are a business. They are incentivized to secure the most convictions for the least amount of work. Finding people that posses CSAM files is easy and gets convictions with relatively little work. It does not however do much to deter child abuse. Police departments and CPS routinely ignore calls for investigations into alleged acts of CSA. Take the Sophie Long case for instance. A question for the reader, is it more important to spend resources stopping CSAM or CSA? Is it really worth giving up our fundamental privacy rights when the police already routinely ignore CSA?
- ahel 5y agoHashes can be of any file. How long will it take before your hard drives are scanned for matching hashes of copyrighted material?
- renonce 5y agoEven assuming the protocol is sound, Apple could choose a set of innocent photos that are frequently found on people's devices, in addition to whatever photos they deem unlawful. In case they didn't show the thresholds to the user, they can set it arbitrarily low. The client will never know which photos were used to determine they were guilty. It's still an interesting read from a cryptography point of view though.