4 ms·
How do you see this specifically creating AGI? How would you calculate the loss on a reinforcement learning model that's autonomous?
by icipiracy 7y ago
How do you see this specifically creating AGI?
How would you calculate the loss on a reinforcement learning model that's autonomous?
- arthurcolle 7y agoAllegory: Running entity A performs action at time t_sub(0) that costs n_sub(t_sub(0)) currency. Response to this action by the counterparty creates a cascading tree of potential new actions, each of which require individual “reconciliations” (new actions, which each require a new response), defined by the probability distribution of potential responses->new actions. We either know these distributions a priori based on our initial conditions, or we can create them based on an initialization function. The net present value of the expected action to these new responses can be evaluated with respect to the NPV of the current holdings of the running entity’s portfolio, and that difference can be treated as the loss function. I’m not an ML researcher, so I apologize if my lack of terminology makes this sound stupid to you but that’s my initial thinking. Feel free to email me if you’d like to discuss further, I’ve been tangentially working in this area for a while but this really gets my sparkplugs going.
- uoaei 7y agoI don't think trying to shove this model into a gradient-descent framework makes the most sense here. I'm an ML (industry) researcher and I highly doubt that AGI will be achieved with gradient descent on neural networks alone. Those may play a small role somewhere in the stack but the orchestration and reasoning will be managed by something else entirely. Neural networks today are fancy MLE machines -- nowhere close to reasoning machines, which require an "understanding" (whatever that means in this context) of dynamics with respect to the environment. Seems more appropriate to start with a population of agents who reproduce at a rate proportional to their recent rewards, and allow them to die off at a rate inversely proportional to the same, a la a continuously-evolving genetic algorithm setup. You may have to modify the reward function to disincentivize behaviors which cause systemwide collapse, but that goes without saying.
- arthurcolle 7y ago> You may have to modify the reward function to disincentivize behaviors which cause systemwide collapse, but that goes without saying. Hopefully humans will figure this out one day too.