5 ms·
Application-specific metrics are the way to go. For ML training this is one example: https://cloud.google.com/blog/products/ai-machine-learning/goodput-metric-a
by sundalia 2y ago
Application-specific metrics are the way to go. For ML training this is one example:
https://cloud.google.com/blog/products/ai-machine-learning/goodput-metric-as-measure-of-ml-productivity https://cloud.google.com/blog/products/ai-machine-learning/g...
- roanakb 2y agoNice, seems like ML Productivity Goodput is a pretty well thought-out metric to understand the overall efficiency of your cluster. I'll consider adding this into our cluster management platform. Only potential drawbacks I'd guess are it being somewhat difficult to compute since it relies on metrics like MFUs, and not something we can observe layer-by-layer to understand inefficient kernels, but I'll take a deeper look. Thanks!