5 ms·
An introduction to zero-knowledge machine learning
- BiasRegularizer 4y agoA 17 million parameter model (~Resnet50) takes more than 50s proof time. Is this on top of the inference time? I can see some niche applications for this system, but I am very skeptical it's ability to handle larger models (100M+) and the ability to and it's scalability when there are increased demand.
- iskander 4y agoZK is currently stuck with arithmetic circuit representations which are predictably very expensive to use as a representation for tensor data. The matrix based formulations are still limited and don't play nicely with the parts of ML models which go beyond simple matrix multiplication. I suspect someone will unify the two threads of research eventually, but it doesn't seem like it's there yet. (FHE ML is even further away)
- wslh 4y agoJust posted another thread from an article of a16z and a Zcash tweet: https://news.ycombinator.com/item?id=35457720 https://news.ycombinator.com/item?id=35457720
- deleted 4y ago[deleted]
- lukeschlather 4y agoI'm very confused by the use case here, and this doesn't make sense to me: > A good example of this would be applying a machine learning model on some sensitive data where a user would be able to know the result of model inference on their data without revealing their input to any third party (e.g., in the medical industry). I don't get why I would care that the answer was generated specifically by GPT4. It sounds like they're billing this as some sort of "run a model on input with homomorphic encryption" but that doesn't really sound possible, and to the extent that it is I don't think you could ever convince me that the people managing the model on the GPU couldn't get access to both the plaintext input and plaintext output. The way to get this kind of security is both simple and hard: make models that can run on consumer hardware.
- reaperman 4y ago> make models that can run on consumer hardware. This will not be hard at all in 10-20 years given the pace of semiconductor FLOPS per watt improvement. https://en.wikipedia.org/wiki/Koomey%27s_law https://en.wikipedia.org/wiki/Koomey%27s_law The neural engine in the A16 bionic on the latest iPhones can perform 17 TOPS. The A100 is about 1250 TOPS. Both these performance metrics are very subject to how you measure them, and I'm absolutely not sure I'm comparing apples to bananas properly. However, we'd expect the iPhone has reached its maximum thermal load. So without increasing power use, it should match the A100 in about 6 to 7 doublings, which would be about 11 years. In 20 years the iPhone would be expected to reach the performance of approximately 1000 A100's. At which point anyone will be able to train a GPT-4 in their pocket in a matter of days.
- MacsHeadroom 4y agoYou're assuming no algorithmic enhancements and missing the currently happening shift from 16bit to 4bit operations which will soon give ML hardware a 4x improvement on top of everything else. We could be training GPT-4s in our pockets by the end of this decade.
- vlovich123 4y agoTo be fair, they’re also being extremely generous about HW scaling. There’s no way we’re going to see doublings every 18 months for the next 6+ years when we’ve already stopped doing that for the past 5-10.
- reaperman 4y agoI haven't seen evidence in a slowdown of Koomey's Law. Would be very interested in those sources!
- vlovich123 4y agoHave you read the Wikipedia page? Moore’s law started ending ~23 years ago followed by Denmark Scaling ~18 years ago. It’s not necessarily fully stopped because there are other architectural improvements that have been delivered along the way, but we simply have reached nearly the end of the road for scaling this due to a combination of heat dissipation challenges and inability to shrink transistors further. 3D packaging might increase things further but it’s difficult and an area of active research (+ once you do that afaik you’ve unlocked the “last” major architectural improvement). I think the current estimates put the complete end to further HW improvements at ~2050 or so. You can still improve software or build dedicated ASICS/accelerators for expensive software algorithms, but that’s the world pre-Moore which saw most accelerators die off because the exponential growth of CPU compute obviated the need for most of them (except for GPUs). We’re coming back to it with things like Tensor cores. Reversible computing is the way forward after we hit the wall but no one knows how to do this yet. > But in 2011, Koomey re-examined this data[2] and found that after 2000, the doubling slowed to about once every 2.6 years. This is related to the slowing[3] of Moore's law, the ability to build smaller transistors; and the end around 2005 of Dennard scaling, the ability to build smaller transistors with constant power density.
- satoshiwasme 4y ago[dead]
- rasengan 4y agoIf Facebook releases Llama, and updated models thereafter, for purchase or as freeware, there will not really be as much need for this since everything will happen safely, locally, no? It would be cool to see Meta release a 7B parameter as shareware, and subsequent larger models for a fee. Edit: To be clear, I'm all for ZK, generally!