4 ms·
How these chips work is pretty reliant on Google's specific infrastructure and there's a stark difference between the work needed to offer a service and selling
by kajecounterhack 8y ago
How these chips work is pretty reliant on Google's specific infrastructure and there's a stark difference between the work needed to offer a service and selling chips. It isn't a matter of bad will or trying to slow ML progress but a matter of what's practical.
- polskibus 8y agoMaybe Google could just license the tech to parties that would be interested in doing the hardware for the masses? Or maybe we just have to wait to see what Nvidia is brewing?
- ori_b 8y ago> How these chips work is pretty reliant on Google's specific infrastructure. I find that hard to believe. Can you be more specific?
- lawrenceyan 8y agoAt Berkeley, we had a guest lecture where one of the major contributors in developing Google’s cloud infrastructure gave a talk on building scalability and robustness into their systems. The majority of the public facing products/services that you find on GCP actually started much earlier as internal tools for employees. There is a huge amount of work that is put into ensuring that everything works reliably, with significant underlying infrastructure needed to achieve it. The problems you face at the level of scale Google works in on a day to day basis turns even the most mundane tasks into extraordinarily difficult algorithmic challenges that have to be solved.
- ori_b 8y ago...er. So, by that reasoning, there is nothing that can exist outside of Google, because Google does things internally at scale.
- lawrenceyan 8y agoI think the point is that a lot of the products and services wouldn't work outside the Google Cloud platform without a significant amount of underlying infrastructure to support it, such that it would be very difficult to just provide some ready to go package for people to buy and implement themselves.
- ori_b 8y agoWhich is largely bullshit. There's complexity in managing failures at scale -- but strip away the scale, and you strip away the complexity in managing failures. You just get normal, bog standard server-grade hardware, with the usual downtime.
- kajecounterhack 8y agoNo, think about all the various services and how they're implemented. For example, for ACLs if you use an internal-facing service that itself relies on proprietary data formats, protocols, assumptions about underlying hardware, and how the API will be used (maybe there's no way to sign up users because you don't need to on a corporate network), then externalizing the ACL service itself would be a huge task in SWE-hours. Now think about how many different internal services like that a single complex ecosystem has to touch. Saying "software is just a binary that needs to be made compatible with X platforms" is naive. It's like saying "facebook is just a bunch of UIs with form boxes, I could write facebook." Yeah, good luck.
- puzzle 8y agoI haven't been at Google for years and never worked on TPUs, but off the top of my head: They control the system attached to the chips. That includes the kernel and userland (glibc, libstdc++, etc.). Very limited number of configurations (CPU, RAM, network, motherboard, any RDMA use) to qualify. Monitoring, profiling, diagnostics, firmware, networking, security and automation follow the internal Google standards. They can tune cooling to accommodate Google's motherboards and racks, whether it's air (v2) or liquid cooling (v3). You can bet that they talk directly to the GCS backends through Stubby/gRPC, rather than sending HTTP traffic through the outside network and traversing GFEs. Then there's all the other "mundane" stuff they don't need to worry about: packaging, manuals, warranties and user-facing RMA, multi-tenancy, etc.
- ori_b 8y agoAs far as I'm aware, they're still on OCP, which standardizes the bulk of this. Basically, this comes down to not being willing to be bothered with testing on diverse hardware. Which, fair enough, is a pain.
- btian 8y agoNot just hardware. TPU probably was only tested to work on one kernel version, one compiler, one version of GRPC etc. It probably takes a lot of effort to develop Windows driver, libraries, support for many kernel version etc.
- deleted 8y ago[deleted]
- ori_b 8y ago> It probably takes a lot of effort to develop Windows driver, libraries, support for many kernel version etc. That's easy to solve: "We only support Linux. Testing has only been done on kernel 4.13."
- kajecounterhack 8y agoGoogle isn't on OCP afaik. Their hardware layer is pretty custom. Maybe you were thinking of Facebook? Google's hardware layer is one thing; as btian mentioned, TPU also is deeply integrated with software infra e.g borg https://ai.google/research/pubs/pub43438 https://ai.google/research/pubs/pub43438
- psds2 8y agoAgreed. TPU will come on premise when Google launches a similar offering to AWS Outposts and ships you an appliance.