3 ms·
> I think this could be more about the competency of the hardware team that manages your accelerators rather than the underlying chip. Google's systems are rel
by mleonhard 3y ago
> I think this could be more about the competency of the hardware team that manages your accelerators rather than the underlying chip.
Google's systems are reliable because of the tens of billions of dollars that Google has invested into developing datacenter hardware, software, and processes over 25 years. Highly-competent teams at smaller and less-mature organizations will always deliver a much worse product.
Another thing to consider is priorities. Google prioritizes reliability. They retire parts that fail repeatedly, even if the failures are relatively infrequent. Smaller and less-sophisticated datacenters keep parts in service even with frequent failures, or don't even monitor failure rates of certain parts. Smaller datacenters buy and use Google's old parts and unreliable parts.
Therefore unreliable machines does not imply anything about the competency of the hardware team.
If the low reliability of the hardware is making your work slow, then how about improving the software so it can tolerate the unreliable hardware, or switching to a more reliable (more expensive) hardware provider?