4 ms·
How can AlphaGo advance capabilities by using the output of a model of the same level? It's not as simple in this domain but there are ways.
by Tenoke 2y ago
How can AlphaGo advance capabilities by using the output of a model of the same level? It's not as simple in this domain but there are ways.
- rightbyte 2y agoGo can be modeled with perfect accuracy. The feedbackloop is perfect. LLMs have no feedback loop with reality if it is short circuited with itself. It needs fresh training data.
- willj 2y agoI think that’s different. AlphaGo is using reinforcement learning in a context in which there is a clear evaluation function— did a strategy lead to a win or loss.
- Tenoke 2y agoI said it's not as simple but there are ways - e.g. you can generate more of your best quality of data, you can try to model the direction of quality or an objective, you can have minor human input at some points to judge which direction performs better, you can objectively verify some of your input to use as a partial objective - code, math, logic etc.