4 ms·
There's a lot of problems. 1. How can I confirm that you've done the computation? 2. Privacy and security issues. Can I trust you too process my sensitive info
by rysertio 4y ago
There's a lot of problems.
1. How can I confirm that you've done the computation?
2. Privacy and security issues. Can I trust you too process my sensitive information?
3. Availability: is there a guarranty you won't just do half of it and then be on a hiatus for 2 months.
But for everywhere these problems are solved we have decentralized cloud computing. For others you need to solve these problems.
- bitxbitxbitcoin 4y agoSome blockchain projects attempted to tackle this. Check out Akash.
- latchkey 4y agoAkash is more about k8s, paid for by tokens.
- latchkey 4y agoThere are a number of entities working on this problem. Plenty of papers in Arvix on it as well.
- kordlessagain 4y agoThe Lightning Network may be used for invoices, identity and payment of use of model. It can also likely be used to handle prompts. Regarding privacy, not all things need to be private and if they do they should be run on a single tenant infrastructure. OpenAI is a multi-tenant cloud model, so guarantee of privacy there may be protected only by license and use agreement.
- alasdair_ 4y ago>1. How can I confirm that you've done the computation? The same problem applies to things like Mechanical Turk and other croudsourcing. The way I've dealt with the issue in the past is to start with zero trust and to have them do computations that I already know the answer to. After that, they do computations that are matched with a random other participant (the two results should match, if they don't, compare against a third random participant). Later, when trust has been developed, you can start assuming their work work is trustworthy, but still check it randomly with computations you know the answer to, or a second person doing the same computation. Yes, this adds overhead (roughly 10% in aggregate) but it works fairly well. It works even better if there are penalties that you can impose for ffraudulent results (like cancelling ALL owed payouts).
- bobkazamakis 4y ago> have them do computations that I already know the answer to. By definition, this is not having them do any computation. The proper solution at this time would be some trapdoor function that is easily verifiable (proof of work), at least while P != NP
- alasdair_ 4y ago>By definition, this is not having them do any computation. If I ask you to compute the first million digits of pi to the power of 1.23456 and I already know the answer to validate it, how is this "by definition" not computation?
- mitchellpkt 4y agoGood questions. Hmm do you think maybe it would be viable if the payouts were based on [model] performance rather than ostensible training time? There's a useful asymmetry we can exploit: finding weights that perform well is computationally intensive and takes time, but scoring a set of weights is fast and easy. A number cruncher could spend 2 weeks training a model, and then when they submit the results it takes me 10 seconds to score the model - to verify the quality of the results, and calculate the performance-based payout. In the #1 or #3 scenario where they didn't do or didn't complete the computations, they wouldn't have a well-trained model to submit for payout. (The lost time in #3 is inconvenient in time-sensitive situations, but mechanisms exist to address that - SLAs, up-front collateral, etc) Regarding privacy, that's an EXTREMELY good and important question. There's some really neat prior art for privacy-preserving machine learning that could be useful here, e.g. https://arxiv.org/abs/2106.07229 https://arxiv.org/abs/2106.07229 "Privacy-Preserving Machine Learning with Fully Homomorphic Encryption for Deep Neural Network" (note I'm approaching this as an interesting DistML thought experiment, not proposing it as an immediately viable or sensible initiative)
- kmeisthax 4y agoSETI@home and Folding@home had to deal with these problems decades ago - even with a closed-source client people would mod it in questionable ways to cheat the leaderboards. Any computation can be verified as having been done by, at a minimum, checking for reproducibility. This requires having each work unit be done twice and only issuing credit if both units match. For deep-learning applications "match" is relative: different compute accelerators are going to give different results. So, instead we can insist that all the floating-point outputs on the model have to match up to the first n bits of mantissa. Neural networks are actually really insensitive to small perturbations in their weights, and it's common to train on 16-bit floats to save time. We can also exploit the loss function itself as a verification mechanism. Generally speaking training is more compute-heavy than inference[0], so we can just run the updated model on the training set and confirm that your trained model is better than the original you were provided with to start from. This will need upper bounds, too - if only to catch people trying to overfit the model to guarantee they get credit. As for privacy and security... the answer is to not train on private data or things that people do not want to be trained. Period. This isn't even a problem solely with distributed computing. All AI training should be limited to either data provided with consent, or data that's so old that training on it would not cause harm. Availability is a problem, but not necessarily one that most distributed computing projects actually have to deal with. There is a minor incentive to participate with the credit system; there's a leaderboard for the fastest/highest credit users and teams. And people do compete for those leaderboard slots, because that's effectively ad space. [0] Model execution.