4 ms·
There's one aspect I don't understand of how the gel "learns" . What are the factors for positive reinforcement and negative reinforcement? How does the gel "kn
by poikroequ 2y ago
There's one aspect I don't understand of how the gel "learns" . What are the factors for positive reinforcement and negative reinforcement? How does the gel "know" what is a good result and what is a bad result?
- BaculumMeumEst 2y agoI would encourage you to download the linked paper and run it thru Claude. I would post results of doing so but that is not allowed here.
- poikroequ 2y agoSo it can hallucinate some bullshit? No thanks.
- BaculumMeumEst 2y agoPeople here hallucinate bullshit all the time. I just came from a thread with some dude saying microwaves all leak dangerous amounts of radiation. Either you read it and reason for yourself or you let jesus take the wheel. OP was not interested in reasoning.
- poikroequ 2y agoSo? What's your point? It doesn't change the fact that LLMs hallucinate bullshit all the time. It doesn't change the fact that I don't trust anything coming out of an LLM.
- BaculumMeumEst 2y agoThen we agree, both LLMs and hackernews commenters hallucinate all the time. You can see my point in my original reply to OP. Not interested in arguing further.
- hasbot 2y agoI've never used Claude only ChatGPT. What does running a paper through Claude do?
- BaculumMeumEst 2y agoIt provides answers to your questions about the document.
- outlace 2y agoIt looks like the gel “learns” to get better at the game because if it correctly positions the Pong paddle over time then the dynamics of stimulation become predictable rather than random. since the system tries to naturally minimize its free energy, it will eventually start to model the ball dynamics enough to better control the paddle, all making the input dynamics more predictable and thus minimizing the energy of the system.
- scilaaverkie 2y agoLove this point. Willingness to spend less energy = need to find predictable patterns
- amelius 2y ago> making the input dynamics more predictable and thus minimizing the energy of the system. How does this follow?
- techjamie 2y agoThe traditional way that you train a biological network like this is that you give it stimulations based on what the game is doing and you give it an ability to control the game. But when something bad happens it's given a lot of random stimulation that it doesn't understand. So the biological network tries to minimize the amount of random stimulation it gets and it learns to play the game better because the stimulation is consistent and predictable. I didn't go into the paper to see if that's exactly what they're doing, and I'm no expert. But from what I've read before, that's how this usually works, and I'm sure they're doing something similar to that.