4 ms·
If solving systems of linear equations is a stronger priority they'd produce hardware with more precision. Unfortunately this is just an offshoot from a machine
by dchftcs 4y ago
If solving systems of linear equations is a stronger priority they'd produce hardware with more precision. Unfortunately this is just an offshoot from a machine-learning driven project. But at least the timings could tell you how much value people can gain by implementing their own hardware solution.
- dekhn 4y agogoogle's entire quest in machine learning has simply recapitulated the previous development of high performance computing. that is, ML is just the same kind of problems the physicists were doing on supercomputers, and google is slowly getting to the point of recognizing their ML problems are just supercomputer problems. Expect larger floats in TPU hardware eventually.
- gwern 4y agoTPUs are for DL, not matrix factorization, even if you might want to try that out since you have TPUs sitting around. The trend in DL has for a decade now been to ever lower precision (and many expect binary at some point and a shift away completely to spiking); why are you convinced that it will not just stop but completely reverse and go all the way to like FP64 or FP128 or something?
- dekhn 4y agoBecause I used to work at Google on TPUs with all the teams (from ads to Physics), and have about 30 years experience in high performance computing. Neural nets and DL are just one of the reasons to use TPUs; in the future, they might not even be the preeminent approach. DL is just one part of the much larger world of HPC, it's been very effective, but I expect that as things evolve people will rediscover the value of deep precision. https://arxiv.org/abs/2009.06489 https://arxiv.org/abs/2009.06489
- gwern 4y agoIf anyone is curious, https://www.gwern.net/docs/ai/2021-jouppi.pdf#page=5 https://www.gwern.net/docs/ai/2021-jouppi.pdf#page=5 provides a breakdown of TPU workload by type. I disagree with Hooker's hardware lottery thesis for the simple reason that if it was true, she would be able to point to examples of DL-competitive methods which use the same amount of FLOPS to get much superior results (because everyone is so badly neglecting them due to the 'lottery'), or at least show better scaling curves so that they will at some point surpass DL, but at much worse wallclock due to whatever hardware specializations favor DL and penalize those alternatives. This is how DL operated: they ran on CPUs, very slowly compared to GPUs, but they did run, enabling eventual exploitation of GPUs. And that is how they could start a virtuous circle: the success of DL, because it's the right thing, starting from hardware not even remotely designed for DL (like 2010-era GPUs were not) and designed to favor other tasks (classic GPGU stuff) has pulled hardware in its wake. So, what are the DL-like things running on CPU or contemporary GPUs? The only actual example Hooker gives is capsule networks - which I thought were doomed when they were unveiled, a poor attempt at stuff soft attention did better already, and have not impressed anyone in the 5 years since, particularly as larger NNs (whether CNN, Transformer, or MLP) continue to deliver what capsnets promised. Nor have any more compelling examples arisen in the years since. With no examples or comparisons, it boils down to nothing much.
- dekhn 4y agoNorm worked just down the hall and he consulted me about physics codes that scientists were running on TPUs. Next generation of TPUs will include higher precision accelerated operations. Note that many scientists tried the "DL will solve physics problems" approach only to find out that DL mostly just did slightly better interpolation than existing systems and happened to run faster mainly because there was a larger pool of experts tuning the engine. I appreciate your passion but it's clear that outside the world of DL there is a lot of HPC that wants TPUs with higher precision. I don't particularly agree with Sara either, as I have always just moved to whatever resource was most available (IE, finding a better lottery).