9 ms·
Federated Learning
- ahelwer 7y agoAll right, I'm cynical as all heck about ad companies and privacy, but this has me optimistic. Somebody disillusion me, why shouldn't I be optimistic?
- deleted 7y ago[deleted]
- carlosdp 7y agoIt could always end up being just fluff, but given it's an active research area and they've open sourced a framework based on the concept already, plus the fact the model is cost-advantageous given you get to offload training to a fleet of customer devices instead of paying for your own servers, it could be the real deal. I started reading it looking at it from a cynical view, but ended it with a "hmm... this could actually work," especially after reading about Secure Aggregation. That's badass.
- patrickthebold 7y agoLooks like everyone gets the same model. So it won't be used for things like targeted ads. Cynical view is that it's only used when Google doesn't want your individual data. This will help muddy the waters in discussions about privacy.
- carlosdp 7y agoSure, it doesn't address that problem head on, however if they moved most stuff to a federated model, they're only a hop, skip, and a jump away from doing ad-selection on-device too, right? I think that if the model proves successful, it could end up proving out the concept enough to incentivize at least trying it out.
- singron 7y agoThere are other reasons not to do things on-device. For instance, being able to pick your own ads might hinder the ability to defend against ad-viewing bots. Also, traditional ad-tech also has to deal with meeting global target spend rates (e.g. if client A pays for 1000 impressions, and client B pays for 9000 impressions, then 90% of the ads must be for client B even if every device would prefer ads for client A). Usually there are realtime databases keeping track of impression counts for all the campaigns and it would be infeasible to synchronize that with every device so that they could make local targeting decisions. I.e. you can't necessarily make the decision locally if it's not local problem, and ad-targeting usually isn't a local problem.
- jameslevy 7y agoJust because everyone gets the same model doesn't mean you can't have targeted ads. It just means the prediction model isn't modeled specific to your behavior. It can still make personalized predictions based on your search input, email content, etc.
- cavisne 7y agoWell the cynical view would be. 1) This still lets you have personalized models, just trained on more than 1 user, thats fine at google's scale anyway 2) Their competitors (FB, AMZN) dont have the edge compute (Android) to do this, and to a lesser degree don't have the ML stack (however Android implements this at the API level will be very Tensorflow focused) 3) Now google can push for privacy regulations that prevent FB and AMZN from storing your raw data 4) Profit That said theres nothing stopping FB doing federated learning within their app on mobile, I just don't think they have the privacy background to bother.
- gbrown 7y agoIf #3 happens I'd be shocked (and pleased).
- pm90 7y agoThey would be absolutely compelled to by market forces. The entire ad-tech industry would be wiped out if something like this becomes more common. You would need an army of highly trained machine learning engineers to build complex systems that would work great but still be efficient. Guess who is the only company which has that army? The cynic in me sees this as a great play by the Googz to cut the Amazon-adtech venture in the bud, and to establish and maintain dominance over the adtech business, and advertising in general.
- defen 7y ago> Their competitors (FB, AMZN) dont have the edge compute (Android) to do this It's not Android scale but Amazon has sold 100 million Alexas.
- hatsunearu 7y agoDoubt Alexas have enough juice to do on device training...
- nixpulvis 7y agoWell the skeptic in me doesn't want my model touched by other peoples models. Of course it depends on the application... For example, I might not want my keyboard autocomplete learning from others, but I might want my self driving car to do it. For applications where we want a common model, I see no other way to do it but this. The idea companies collect massive stockpiles of data forever is infeasible.
- krick 7y agoI think this is pretty cool from the technical perspective, but indeed I don't see the reason to be optimistic from the "privacy concerns" perspective. In fact, if this is seriously gonna be used as a "better privacy" argument (as some people seem to be already doing in this very thread), I'm calling it a PR victory for the "bad guys". First off, if you were worried about what Android was sending to Google, there's no reason to believe it's going to stop. In fact, I believe that the biggest problem always was the (carefully cultivated) confusion about what data is being sent: you have a hundred of menus to "opt-out" of something and it's not even exactly clear if it changes anything. In fact, we know for a fact, that when "opting out" in some cases more data is being sent. And even if we are not entirely happy about it, most of us still allow this to happen, because there's nothing we can directly do to prevent it and everybody says "well, I do need a smartphone after all, right?" (yeah-yeah, somebody doesn't, but we are not talking about the weird minority here) And all of it happens when it's relatively straightforward to see what data is being sent, because of minimal aggregation on the device. And still even somewhat technically-minded people don't really know what Google (Facebook, Amazon, whatever) really knows about them. Second, what really is "federated learning"? Well, let's imagine no humans speak Chinese, but there is this program (owned by Google), that does. And speaking Chinese is how it actually operates internally, when deciding to show this or that ad to you, or sending a ballistic missile to your location. So, in order for it to learn, we normally were sending English sentences, which were translated to Chinese server-side. Federated learning is when they are translated to Chinese client-side (which might be considered "lossy" conversion, but to what degree is not really specified), and then sent to Google to be aggregated. So, yeah, no raw data has been sent, but the central Chinese-speaking machine still somehow knows it all. What exactly it knows, depends on what we really meant by "translating into Chinese" in our metaphor. But effectively, we just offloaded some processor work to the client side, which, as I said, seems really cool to me from the technical perspective, but there's no way it automatically protects us from anything. Third is basically 1 + 2: we didn't know what is being sent when it was all raw-data, we will know even less, when it's client-side aggregated in some unintelligible-for-the-humans way. And it scares me even more, because if it allows some PR victories for the Google&Friends, then sky is the limit for what more surveillance can be done this way. I mean, if it would be publicly known that Android sends all the sound and all the image from your mic & camera to Google, I think (I hope!) that people would seriously oppose to that. But if it's not the real images, but just some matrix of weights, learnt from them — it might be less clear if anybody has to object to that. And I think they absolutely have to! Because if we don't make any very restrictive assumptions about what we mean by "learning" in this very specific case, then the only thing that matters is that the "central brain" still saw all these images, it just isn't known what exactly it "remembered". After all, we, humans, also don't store all the pictures we've seen in our brains: it doesn't make you much happier if I saw your transaction history (or whatever else you don't want me to know), because it was never the picture I was after, but only the "aggregated info".
- frenchtoastto 7y agoWell if large companies have access to metadata like this aggregate or not they know what people are doing, looking at, and talking about. Which allows them to play the financial markets like a drum. Why do they need the aggregate data? To make money plane and simple, tis capitalism 101.
- arnioxux 7y agoI wonder how they are going to handle security/abuse? Once the training is on end user devices, can't the user give fake results? For example back when recaptcha was introduced, trolls tried to transcribe everything as the n-word (since you only had to get one of the two words correct). These cases can obviously be noticed and fixed but it's harder now that the training data is opaque right?
- unreal37 7y agoSo instead of sending the data to Google encrypted for them to analyze, it analyzes the data on your device and sends that data to Google encrypted for them to combine the results. But your data still gets sent to Google. I don't see the difference. It's just another layer on top.
- cycrutchfield 7y agoDid you bother to read how this works?
- carlosdp 7y agoWell no, if you read the part about Secure Aggregation, Google has no way of knowing which piece of training results comes from which device, they can only see the aggregated results of a batch. So sure, technically the training results based on your data are still sent to Google, but that's not really the concern they're addressing. They're addressing Google having a record in a database of every shop you visited in the last week and such and that data getting in the wrong hands (or being used wrong by them). What if they could benefit from training on that sort of data, without ever actually storing it themselves?
- arkades 7y agocall me jaded but: If you’re paying for PR firms to produce cartoons about how good you are for privacy, you’re probably terrible for privacy. This feels like Google’s Joe Camel moment.
- cycrutchfield 7y agoThink about the researchers and engineers who worked hard on this reading your unwarranted and uninformed cynicism and how that must feel.
- carlosdp 7y agoIt wasn't a PR firm, it was this cartoonist: https://lucybellwood.com/ https://lucybellwood.com/ I mean, the comic addresses the fact the current model is bad for privacy right out of the gate and then shows how this team is trying to solve it, what more do we want from them?
- Thorrez 7y agoGoogle Chrome launched with a comic by Scott McCloud in 2008. It looks like Scott McCloud helped on this Federated Learning comic as well. https://www.google.com/googlebooks/chrome/big_00.html https://www.google.com/googlebooks/chrome/big_00.html
- lern_too_spel 7y agoI've been saying that about Apple's privacy lip service and disingenuous privacy marketing for a long time.
- grantlmiller 7y agoFirst, I've loved that Google open sourced Tensor Flow Federated as a way to encourage the rest of the world to adopt this method of decentralized machine learning. Second, I was a bit disheartened that this concept had to be explained with a comic strip to make it accessible because I hoped the benefits were clear to everyone. Third, I read the comic strip, learned new things (secure aggregation protocol, wtf, amazing!), kicked myself for being smug and appreciated the huge amount of effort that someone invested to communicate this.
- pm90 7y agoPopular culture is a tool for the education of the masses. Even for people who may be technically inclined, its not always evident what certain technologies really do. I am a software engineer but mostly work on DevOps-y stuff. This was a very accessible, low-investment way for me to understand exactly what "Federated Learning" really meant. Some of the best teachers at Univ had a way of explaining things in simple terms. This comic strip has captured that experience in a more permanent form a lot better than a textbook would.
- dmix 7y agoUniversities have a captive audience, they can take the time to walk you through incrementally. Most websites and online communication don't have that luxury. It's interesting how the comic works well in these situations, while still pushing a long-read format. Google did the long-form comic thing with Chrome too and I remember reading it page-to-page back then. But at the same time, is it a good idea as your primary website homepage as it is here? Which would be unusual if there was anything more to it like documentation, code, etc. Right now this website is clearly in an educate-the-public mode only which is how they can get away with this being the primary content.
- jhanschoo 7y agoThe comic strip format helps in that the audience is not just potential developers, but also the general privacy-conscious consumer.
- deleted 7y ago[deleted]
- satokema 7y agoRenting my phone out to process data gives me a bad feeling. The airplane mode guy is now just straight up turning off the phone and battery.
- anonytrary 7y agoThis is reminiscent of bitcoin mining, except the thing being mined here is an AI's "intelligence", and consumer's data is the key to it. The benefit is that the consumer doesn't have to give up their data, just their compute power. Obviously, this should be an opt-in service and people should be getting paid for the compute time they loan out.
- wybiral 7y agoHow do they assure you that the training algorithm isn't just exfiltrating your data? Edit: By that I mean... What's stopping the model from being as simple as "learn my personal information"?
- pas 7y agoNothing, of course. But they are probably going to use an open source implementation with probably some kind of deterministic build/compile systems, so they can show you what the training algorithm is (that it does not contain a hidden user id or such to magically "overfit" for that user), and similarly the training questions should also be knowable and provably non-user-specific.
- AlexCoventry 7y agoIt looks from the comic like they aggregate data with linear structure (probably derivatives on the model parameters.) Each device adds a mask to their part of the data, and somehow the masks are coordinated across devices so that when the data are summed on the central training server, the masks cancel out. It's unclear from the comic how the masks are coordinated, or how they compensate for the risk that a participating device drops out (which will make all the other data from that iteration useless, if you set the masks up in a naive way.)
- wybiral 7y ago> It looks from the comic like they aggregate data with linear structure (probably derivatives on the model parameters.) Thanks, that's the part I must have glossed over. It looks like they're using secret sharing to distribute as shares that all need to be together to reassemble [1]. [1] https://storage.googleapis.com/pub-tools-public-publication-data/pdf/ae87385258d90b9e48377ed49d83d467b45d5776.pdf https://storage.googleapis.com/pub-tools-public-publication-...
- AlexCoventry 7y agoThanks, just came back to share that link. :) From the introduction, it looks like they group participants into smaller clusters, coordinate between those via the centralized server in a star topology, use Diffie-Helman to share secrets between the participants in a cluster, and construct the canceling noise within that cluster. The Shamir secret sharing squicks me a bit. It looks like if the adversary can control cluster membership (and Google is the adversary, here), they can recover the gradients. > To prevent the server from simulating an arbitrary number of clients (in the active-adversary model), we require the support of a public key infrastructure that allows clients to register identities, and sign messages using their identity, such that other clients can verify this signature, but cannot impersonate them. In this model, each party u will register to a public bulletin board during the setup phase. The bulletin board will only allow parties to register keys for themselves, so it will not be possible for the attacking parties to impersonate honest parties.
- pas 7y agoWhat happens with the zero-sum cancelling out phase if one device disappears during the process?
- deleted 7y ago[deleted]
- tylerhou 7y agoFrom the paper: https://storage.googleapis.com/pub-tools-public-publication-data/pdf/ae87385258d90b9e48377ed49d83d467b45d5776.pdf https://storage.googleapis.com/pub-tools-public-publication-... > We rely on Shamir’s t-out-of-n Secret Sharing [50], which allows a user to split a secret s into n shares, such that any t shares can be used to reconstruct s, but any set of at most t − 1 shares gives no information about s.
- archgoon 7y agoSo, correct me if I'm wrong, but this basically only works when you've already done your data exploration phase, you've committed to a particular topology, and now you just want to optimize your weights? It seems that this won't work so great if you don't have any initial data to bootstrap yourself with. So, perhaps the idea is you bootstrap with a few people, do your explorations, and then scale it out with federation?
- walterbell 7y agoGoogle mentioned at I/O that speech recognition will soon (this summer?) be performed locally on Android devices, with no voice data being sent to Google, because they have been able to reduce the size of the model dramatically. Is that related to federated learning? Paper: https://arxiv.org/abs/1811.06621 https://arxiv.org/abs/1811.06621
- ma2rten 7y agoNo, federated learning is about training on the device not running prediction on the device. Training a speech model on the device would be hard, because there is no labeled data. We don't know what the user said.
- strin 7y agoI can imagine the world relying more and more on unsupervised pre-training approaches, such as BERT and GPT-2. Then we’ll just need a few labeled data to generalize.
- arthurcolle 7y agoHow can the data be sent in an encrypted manner that can then be useful without the server having a copy of the private keys used to encrypt the data itself?
- nemo1618 7y agoSurprisingly, there are a few ways in which you can perform operations on encrypted data: https://en.wikipedia.org/wiki/Homomorphic_encryption https://en.wikipedia.org/wiki/Homomorphic_encryption However, the best techniques we have for doing so are many orders of magnitude slower than their un-encrypted counterparts, so it's not feasible today.
- ddtaylor 7y agoHomomorphic Encryption
- arthurcolle 7y agoDoes this actually work today? I was tangentially involved in some 'zero protocol'/zcash-related projects a few years back and the lack of ability to communicate and transfer information while being able to perform computation on it was a major drawback to most of the interesting ideas in the space. Are they actually using this in this intended federated learning plan? If so that's a truly major innovation.
- CyanTas 7y agoFully homomorphic encryption, in which you can do arbitrary computation on encrypted data, is still quite slow. But partially homomorphic decryption, in which you can add encrypted values together but not multiply (or vice versa), is quite efficient. And since the secure aggregation protocol only needs to add together encrypted values to get an average, it only needs partially homomorphic encryption properties.
- ddtaylor 7y agoI believe there is also a proof that says any partially homomorphic system can be reworked into a FHE.
- ivan_ah 7y agoThis is very interesting for many reasons. First we have the privacy stance, which is a tremendous step for big G. Whoever managed to push this through in the "machine" of internal office politics deserves applause. The very fact of acknowledging that users might want to control their data locally rather than rsync everything all the time is a big step—it takes us off the "give me all your data" train that we have been on for some time. Talking about specific applications of your users' data makes a lot more sense: "If you share X with us, you're helping to build a better model Y that helps you with Z." Then the prompt "Do you want to share X?" makes a lot more sense than the current generic prompts "App V wants to access all your data W?" which doesn't tell you anything. The anonymisation-by-aggregation aspect is interesting on it's own since it provides a practical approach we can use today and not have to wait for homomorphic encryption. There will probably still be "data leakage" but I can see how aggregation can be fundamentally better than trying to shared anonymized data by fuzzing identifiers, randomization, and binning, which are notoriously hard to pull off and suffer from de-anonymisation attacks by cross linking with other datasets. Research-wise this could be a whole new field. Let's revisit all the ML algorithms and look at the ones that lend themselves to federated updates. Perhaps certain ML algorithms have been overlooked historically because they are not "cutting edge" but lend themselves better to distributed model updates? (I bet this is already a thing...) The communication complexity aspects are also very interesting since it forces us to think about bandwidth needed to communicate model updates and training batching. For high-bandwidth settings we could consider training a model from scratch, for medium bandwidth you can send model updates regularly, but what would be particularly interesting to see async and VERY low bandwidth updates—like just a few MB every, exchanged once in a while when connectivity is available.
- pm90 7y ago> "If you share X with us, you're helping to build a better model Y that helps you with Z." Then the prompt "Do you want to share X?" makes a lot more sense than the current generic prompts "App V wants to access all your data W?" which doesn't tell you anything. That would lead to wayy too much notifications. Just like ToS, people would say yes or no blindly.
- 7y ago
- im3w1l 7y agoMy gut feeling tells me not to believe their promises that it's impossible to deduce the data from the model updates. That there should be attacks. My stylistic criticism is that they portray white men in a demeaning way that they would never dare do to any other group. edited to make a weaker claim
- CyanTas 7y agoDo you have a specific technical criticism of the secure aggregation protocol? That’s what’s supposed to make it impossible to deduce the data from model updates. Or is your concern something else?
- im3w1l 7y agoI hadn't really read it at that point. It more seemed like a too big achievement for me to believe that anyone could solve. One issue with their approach I found while causally browsing is "for the proof against active adversaries, we assume that there exists a public-key infras-tructure (PKI), which guarantees to users that messages they receive came from other users (and not the server). Without this assumption, the server can perform a Sybil attack on the users in RoundShareKeys" Basically assume there is some trustworthy entity that solves Sybil attacks. I don't think such an entity exists. So question is how they solve that in practice.
- ximeng 7y agoLinked paper on using this for Google Keyboard (https://arxiv.org/pdf/1903.10635.pdf https://arxiv.org/pdf/1903.10635.pdf) highlights that there are nevertheless still privacy issues with this approach: While Federated Learning removes the need to upload raw user material — here OOV words — to the server, the privacy risk of unintended memorization still exists (as demonstrated in (Carlini et al., 2018)). Such risk can be mitigated, usually with some accuracy cost, using techniques including differential privacy (McMahan et al., 2018). Exploring these trade-offs is beyond the scope of this paper.
- CyanTas 7y agoYou’re not wrong, but it does say those privacy risks can be mitigated with differential privacy. That McMahan et al. paper (which is also Google) makes the accuracy cost seem low. https://arxiv.org/abs/1710.06963 https://arxiv.org/abs/1710.06963
- gok 7y agoFederated learning is a potentially really great idea, but it's important to be upfront about its limitations. Just because I can't prove that a piece of data came from your device doesn't mean that a machine learned model trained on that data isn't violating your privacy. For example, say we deployed federated learning to train a predictive language model, and allowed it to learn from emails, say, inside Google. Looking at what the model predicts when you type "Here at Google our next secret project is..." could very likely reveal something they wouldn't want widely revealed.
- jonathanhd 7y agoI'm genuinely still unsure if this is a parody or not. The first half of the comic just describes Google's business model and the second seems to be trying to outsource the cost of G/TPUs to the end user. Then at the end they go bankrupt and (presumably) sell their control over the data to a vulture fund. None of this addresses the fundamental problem of advertising companies, once people learn what they're doing they just want them to feck off and leave them alone, without any regard for future promises.
- AlexCoventry 7y agoIt's no joke. https://federated.withgoogle.com/#learn https://federated.withgoogle.com/#learn