5 ms·
That sounds exactly like the new autopilot system they describe in their USENIX paper. Talk: https://eventsonair.withgoogle.com/events/autopilot-research-talk-
by o- 6y ago
That sounds exactly like the new autopilot system they describe in their USENIX paper.
Talk: https://eventsonair.withgoogle.com/events/autopilot-research-talk-series/watch?talk=autopilot https://eventsonair.withgoogle.com/events/autopilot-research...
Paper: https://research.google/pubs/pub49174/ https://research.google/pubs/pub49174/
- kyrra 6y agoGoogler, opinions are my own. The paper and the system that caused this incident are different. Google has a ton of different automated systems for maintaining production. The paper you linked is for a system called autopilot, which can scale a specific jobs memory and CPU usage up or down depending on historical load of that job. Think of it as having different instance sizes on GCP or AWS, then you have a tool that monitors how much actual CPU and memory your job uses on the instance, in will upsize or downsize your machine depending on load over time. The system that broke based on the incident report above had to do with quota that a given production system could use as a maximum. Similar resource management automation, but very different systems. So one is about fine-tuning, while the other is about preventing excessive resource usage.