4 ms·
here is an older take on this same topic.. https://www.yitay.net/blog/training-great-llms-entirely-from-ground-zero-in-the-wilderness https://www.yitay.net/blo
by ai4ever 2y ago
here is an older take on this same topic..
https://www.yitay.net/blog/training-great-llms-entirely-from-ground-zero-in-the-wilderness https://www.yitay.net/blog/training-great-llms-entirely-from...
GPU vs TPU, and good software managing large clusters of them across all sorts of failure.
the funny bit from the above article is the incident when someone forgot about a training job at google, and month later had the model fully trained without an alert of any kind. "outrageously good infra"