42 ms·
I'm somewhat confused about how they would be doing that. Do you have any references to blogs/papers on this technique?
by vitno 10y ago
I'm somewhat confused about how they would be doing that. Do you have any references to blogs/papers on this technique?
- bjterry 10y agoOne typical strategy for spam detection is to convert text to a "bag of words" representation[0]. If you take this bag of words representation and hash all the values, then rather than training words like ED, you are training on word number 19213123. The number of these hashed values is smaller than the number of words, just like a hash table, and this generally doesn't harm the accuracy of the machine learning. When you receive feedback on the classification (from people reporting spam or people marking things as not spam), you just turn the reported email into a bag of hashed words and feed that change into your model. Because the order of the entries in the bag of words is arbitrary, and the words have been hashed, it is impossible to go back from a bag of words representation to the original email. I don't know if this is what google does, but it is pretty normal to do so. 0: https://en.wikipedia.org/wiki/Bag-of-words_model https://en.wikipedia.org/wiki/Bag-of-words_model
- cvwright 10y agoThis is actually a really tricky topic. Things that sound like they should give very good security, often don't in practice. The "hashed bag of words" technique that you describe here sounds an awful lot like some recent attempts at letting legacy systems search on encrypted data [1,2]. We took at look at this recently [3], and it turns out that mapping the word numbers back to the original words is actually a lot more doable than you'd think. [1] ShadowCrypt: Encrypted Web Applications for Everyone http://dl.acm.org/citation.cfm?doid=2660267.2660326 http://dl.acm.org/citation.cfm?doid=2660267.2660326 [2] Mimesis Aegis: A Mimicry Privacy Shield–A System’s Approach to Data Privacy on Public Cloud https://www.usenix.org/conference/usenixsecurity14/technical-sessions/presentation/lau https://www.usenix.org/conference/usenixsecurity14/technical... [3] The Shadow Nemesis: Inference Attacks on Efficiently Deployable, Efficiently Searchable Encryption https://www.sigsac.org/ccs/CCS2016/agenda/ https://www.sigsac.org/ccs/CCS2016/agenda/
- bjterry 10y agoYeah, I definitely don't think that this would give you mathematically provable security, especially if you are including n-grams which would allow you to chain together sentences combined with a language model. How much that matters in practice depends on the dimensionality reduction. In the context of google I doubt if this is even much of a specific goal, since no matter what they train on, they actually have access to the source text if they want it.