5 ms·
This feels exactly like the loop I'm stuck in right now, making me rely on sudden waves of motivation to complete such projects rather than curiosity driving me
by reallymental 3y ago
This feels exactly like the loop I'm stuck in right now, making me rely on sudden waves of motivation to complete such projects rather than curiosity driving me. Have you ever come out of this cycle before? I.e. Become more effective in planning a project so it doesn't lead one directly into a death-loop like above?
On a side-note, love your writing style, it's dry and sprinkled with the right amount of humour.
- Buttons840 3y agoThanks. I'm about to give up on this project. Solving Slay the Spire on a single computer using model free methods was always quite ambitious. I've seen enough learning to believe that it works, I know I've implemented the algorithms correctly enough, that feels good at least. I've learned a lot along the way. I might try to get everything running in a container, then I can just ship the containers out to run on the cheapest hosting I can find and I won't have to babysit the experiments so much. I was recently watching https://karpathy.ai/zero-to-hero.html https://karpathy.ai/zero-to-hero.html and at one point he basically says "stop, it's not time to train yet, first we need an experimental harness", and I've been thinking a lot about that. That is to say, I haven't actually followed the advice yet, because who the hell can resist trying as soon as there's the slimmest chance it might work.
- jiggawatts 3y ago> "stop, it's not time to train yet, first we need an experimental harness" I was going to say that the model is probably stuck on the boss because it doesn't get sufficient success feedback for incremental improvement. The Google Alpha team used simplified games that allowed the AI to get over humps like this.
- reaperman 3y agoSo it's tricky. I believe the solution is to break the gameplay down into basic tasks - moving to a targeted point, deciding which point to move to, avoiding attacks, attacking, etc. But the more "tasks" you train the RL AI on, its like you put more guardrails on it and it becomes closer and closer to "supervised" learning -- it's only learning what you imagined for it. Which misses out on the fun of a lot of the potential for RL AI, which is that it should come up with novel techniques and strategies to games that humans can learn from. So, yeah. Personally I feel like the answer is doing lots of different versions of mini-training for tons of different games -- but just the sub-part training for the games. Then once it's trained on the "universal mechanics" of gameplay, allow it to start playing more games. The trick is going to be super-generalization. IMHO. But in AI, opinions are a dime a dozen.