3 ms·
Certainly is a possible outcome although there are a few problems with the current paper as I see it: * Quite slow to execute (somewhat inherit to diffusion mo
by f_devd 4y ago
Certainly is a possible outcome although there are a few problems with the current paper as I see it:
* Quite slow to execute (somewhat inherit to diffusion models)
* Requires a lot of human data which increases dev time, since it needs to be done near the end of gamedev to get a consistent enviroment
* The current paper doesn't consider previous observations (I can't find the reference so I could be wrong)
I believe issue 1 & 3 can be overcome quite easily with changes in model architecture, and 2 can probably be overcome with RLHF to pretrain on self-play and fine-tune on human input.