3 ms·
I have become increasingly doubtful when it comes to using existing applications as a measure for coding agents. If the task in question is difficult enough tha
by Orien_18 26d ago
I have become increasingly doubtful when it comes to using existing applications as a measure for coding agents.
If the task in question is difficult enough that it cannot be completed in just one try, then the result's quality largely depends on the model's supporting framework, i.e. how the problem is divided in the prompt, how many iterations are made, what tools are employed, and how many instructions it gets from a human operator.
I would also propose to take into account the documentation and testing processes while constructing the framework. The action of an agent that preserves its history, records unsuccessful attempts and known pitfalls, and constantly makes and executes tests can differ a lot from the same model in the absence of such history. In a lengthy task the ability to remember what was learned from previous failures can be as significant as producing the next piece of code.
Another issue is related to contamination. When it comes to such game as Minecraft, the model is not starting from scratch. Many elements, such as game mechanics, design, tutorials, and many lines of code are available online for a long time and probably formed a basis