4 ms·
Being an agent isn't about how complex the model is, nor about how much RL is used. Bacteria is an agent. Sponges and Jellyfish are agents. It's about if the mo
by Vetch 4y ago
Being an agent isn't about how complex the model is, nor about how much RL is used. Bacteria is an agent. Sponges and Jellyfish are agents. It's about if the model engages in exploration and or generates inferences in service of non-trivial control problems.
> On the other hand, if you strap GPT-3 on an agent in a 3D Home Environment
GPT-3, just as any probability distribution, can inform an agent's actions but that doesn't make it one. WebGPT3 is an agent though.
- astrange 4y agoYes, WebGPT3 is more like intelligence research than GPT is. So is DeepMind's Perceiver. WebGPT does need safeguards because eg it could start querying websites and end up hitting their /delete APIs. But that's not really an "AI alignment" problem; same thing happens with GoogleBot.
- visarga 4y agoYes, RL is mostly about learning from rewards (Reward is Enough - https://www.deepmind.com/publications/reward-is-enough https://www.deepmind.com/publications/reward-is-enough) and GPT-3 has been used so far only to augment agents, not to do reward based learning. The current obstacle is speed - it needs to run in real time. Maybe we need more efficient algoritms, maybe we'll have better hardware. It will come soon enough.