4 ms·
> that blog runs on a TPU. You can just SSH into them. By “runs on a TPU,” you mean it runs on a CPU in a VM somewhere that also has access to TPU hardware, ri
by rrss 5y ago
> that blog runs on a TPU. You can just SSH into them.
By “runs on a TPU,” you mean it runs on a CPU in a VM somewhere that also has access to TPU hardware, right?
If this is the new TPU phraseology, that’s pretty confusing IMO.
- sillysaurusx 5y agoYou’re right, I spoke carelessly. The proper term is that the blog is running on a “TPU VM”, which is exactly what you describe: a box with /dev/accel0 that libtpu.so uses to communicate directly with the TPU hardware. The difference is, every TPU VM has 96 cpu cores and 350GB of RAM. (Pods only get 48 cores per host, but they have 1 host per 8 TPU cores, and the smallest pod has 32 TPU cores, for a whopping 1.2TB of RAM and 196 CPU cores.) Which is to say, I still think of them as “a TPU”, because nowhere in the world have I ever been able to access that amount of raw horsepower. Not on x86_64 Ubuntu that you can pip install things on, at least. Like a blog. :) Wanna see a magic trick? Clone tensorflow onto a TPU VM and start building it. htop will light up like a Christmas tree (https://twitter.com/theshawwn/status/1400771262901854214 https://twitter.com/theshawwn/status/1400771262901854214) and it'll finish in about 25 minutes flat, if I remember correctly. So yeah, Google is throwing around compute like Israel doling out vacations to Tel Aviv for distant relatives: it’s totally free. Partake! (I’m really looking forward to seeing Israel someday. I never realized how beautiful the beaches are...) 96 core TPU VMs, free as in beer! It’s so exciting that I just can’t shut up about it. TRC gives you access to 100 separate VMs the moment you sign up. Having access to 100 VM-yachts totally rules. SSHing into 100 of them feels like commanding a WW2 carrier division, or something. It’s quite literally too fun: I have to force myself not to spend all day unlocking their secrets. There’s so much new territory to explore — every day feels like a tremendous adventure. Our computing ancestors could only dream of exploiting as much hardware as our M1’s take for granted, let alone one of these behemoths. Let alone one hundred! I went back to look at my old notes. Christmas in 2019 was magical, because in January of 2020 I managed to fire up santa's 300 TPUs, while a colleague fired up santa's other 100 TPUs. Then we swarmed them together into a tornado of compute so powerful that even connecting to all 400 TPUs required special configuration settings ("too many open files" aka sockets): https://twitter.com/theshawwn/status/1221241517626445826 https://twitter.com/theshawwn/status/1221241517626445826 We were training models so fast that I bet even Goku in the hyperbolic time chamber would have a hard time training faster.
- rrss 5y agoThanks for the clarification. It’s cool that google has a ton of money and can give people free access to lots of big servers, but it sounds like what you are excited about (lots of cpu cores, RAM, many VMs) seems to be mostly unrelated to the actual new TPU hardware, which is sorta disappointing.