5 ms·
A while back I saw a post where people ran a model over and over to accomplish a code base port from one language to another. In their prompt, they told it to
by ddingus 1y ago
A while back I saw a post where people ran a model over and over to accomplish a code base port from one language to another.
In their prompt, they told it to leave itself a note and to accomplish something each time.
Then they put the model in a loop and it worked. In one instance, a model removed itself from the loop by editing a file or some other basic means.
To me, iterative tasks like like multiply and long divide, look an awful lot like the code port experiment.
Putting models into loops so they get more than one bite at the task seems to be a logical progression to improve capability.
- CaveTech 1y agoThe amount of paths in the wrong direction are infinitely more than then number in the right direction. You'll quickly realize this doesn't actually scale.
- hodgehog11 1y agoI'm a bit confused by this; are you referring to vanishing/exploding gradients during training or iteration at inference? If the former, this is only true if you take too many steps. If the latter, we already know this works and scales well.
- CaveTech 1y agoThe latter, and I would disagree that “this works and scales well” in the general sense. It clearly has very finite bounds by the fact we haven’t achieved agi by running an llm in a loop.. The approach of “try a few more things before stopping” is a great strategy akin to taking a few more stabs at RNG. It’s not the same as saying keep trying until you get there - you won’t.
- hodgehog11 1y ago> It clearly has very finite bounds by the fact we haven’t achieved agi by running an llm in a loop.. That's one hell of a criterion. Test-time inference undergoes a similar scaling law to pretraining, and has resulted in dramatically improved performance on many complex tasks. Law of diminishing returns kicks in of course, but this doesn't mean it's ineffective. > akin to taking a few more stabs at RNG Assuming I understand you correctly, I disagree. Scaling laws cannot appear with glassy optimisation procedures (essentially iid trials until you succeed, the mental model you seem to be implying here). They only appear if the underlying optimisation is globally connected and roughly convex. It's no different than gradient descent in this regard.
- CaveTech 1y agoI never made a claim that it's ineffective, just that it's of limited effectiveness. The diminishing returns kick in quickly, and it's not applicable in more domains than it is applicable.
- razodactyl 1y agoBut test-time inference leads to better data to train better models that can generate better test-time inference data. There's an obvious trend going on here, of course we're still just growing these systems and going with whatever works. It's worked well so far, even if it's more convoluted than elegant... What puts my mind at ease is that the current state of these AI systems isn't going to go backwards because of the data they generate which contributes to the pool of possible knowledge for more advanced systems.
- ddingus 1y agoAchieving agi is not a requirement to working well.
- malfist 1y agoHow do you know if you've taken too many steps beforehand?
- hodgehog11 1y agoIt's a hyperparameter much like learning rate. If the learning rate is too high, the training process would not work either. Addressing this is just a matter of a grid search.
- ddingus 1y agoI am not sure it needs to scale.
- razodactyl 1y agoThe feedback from compilation tools / linters fed into the training loops is an example of this. What we end up with however is a model good at coding for example but bad at something else. And without enough general coding, good at one language over another. And we're back to square one. The problem of being able to achieve true intelligence by distilling the essence of it not just knowing the answers to specific problems. Given enough time, we'll plug the gaps and maybe get good enough but it's not true intelligence until it can learn in a way that excels at all fields in a cross-disciplinary way - much better than the side-effect way it's doing now where some other knowledge does actually contribute to achieving goals in other domains.