3 ms·
Running the largest workload I’ve ever personally launched (4,200 nodes, 144 cores each, 16,000 simultaneous jobs with a mix of one 1,000 node cpu job, 1 AI sel
by trws 3y ago
Running the largest workload I’ve ever personally launched (4,200 nodes, 144 cores each, 16,000 simultaneous jobs with a mix of one 1,000 node cpu job, 1 AI selection service on one node, about 3,200 nodes worth of 4-core cpu input pre-processing jobs and 4 GPU jobs per node co-located with all the CPU stuff) at 2am the day of the deadline for something about 50 of us had been working on for most of a year, realizing something was set wrong in the launcher. It was our last chance for a full run, we couldn’t start it over or try again, and it was going to fail because of a single runtime value.
I attached gdb to the launcher, “print <var>=<value>” and detached. The run started going about 10 times faster and we got the whole thing done. Crazy, dangerous, but it worked.
Question also made me think of the last-minute change we needed to make to a database’s structure to avoid about 6 million lost user updates when all DB admins were out with no password. That was a fun one too. Not sure I should admit how we managed it.