2 ms·
Sure, that makes sense. With this many cores it seems like the probability that a core dies during a multi-hour job (or in case it's used for inference, during
by groundlogic 7y ago
Sure, that makes sense.
With this many cores it seems like the probability that a core dies during a multi-hour job (or in case it's used for inference, during a very long-lived realtime job) is pretty high, so the software in all layers would need to handle this kind of exception. They probably don't, today, since we haven't seen a 400k core chip before.