3 ms·
It could just be that improving the models does not parallelize that well, so you need to wait for the massive months long training runs to finish to see what w
by tsurba 3y ago
It could just be that improving the models does not parallelize that well, so you need to wait for the massive months long training runs to finish to see what worked and what to try next.
OpenAI got started earlier going full on with scaling up the transformer architecture (even if Googlers came up with it first).
Of course if you are smarter or can run more experiments simultaneously, you can catch up at some point. But it could still take a while even with just one year head start.