3 ms·
Changing my old coding behavior aside, biggest limiting factor for me is understanding how and why the coding agent is doing this a certain way, so that I have
by SafeDusk 1y ago
Changing my old coding behavior aside, biggest limiting factor for me is understanding how and why the coding agent is doing this a certain way, so that I have the confidence to continually sharpen my tools.
I want something simple that I have full control on, if not just to understand how they work. So I made a minimal coding agent (with edit capability) that is fully functional using only seven tools: read, write, diff, browse, command, ask, and think.
As an example, I can just disable `ask` tool to have it easily go full autonomous on certain tasks. Or, ask it to `think` for refactoring.
Have a look at https://github.com/aperoc/toolkami https://github.com/aperoc/toolkami to see if it might be useful for you.
- eschaton 1y agoIt’s producing statically likely next tokens based on its training corpus and prompts. That’s what it’s doing. It’s not analyzing your code, it’s not considering approaches, it’s not theorizing about behavior, it’s just producing statistically likely tokens. This is why I don’t touch the shit—it’s fucking snake oil.
- ekidd 1y agoIf you "don't touch the stuff", you are probably relying on your initial impressions from 2022 to evaluate models in 2025. Also, predicting the next token with high accuracy may demand very high degrees of knowledge and reasoning. As an extreme example, please predict the next 1,000 tokens in this sequence: "ABSTRACT. In this paper, we show a mathematically elegant unification of quantum mechanics and gravity, which makes testable predictions. Testing these predictions shows that the quantum gravity model provides previously unexpected results accurate to 1 part in..." To predict the rest of that paper, it helps to actually come up with a workable model of quantum gravity. Which no current LLM can do, happily. But lots of current models are good enough at "predicting the next token" to solve high school honors math problems that they've never seen before. They can apply the chain rule, factor polynomials, double-check their work, backtrack, etc. Similarly, current-generation coding models are perfectly capable of reading compiler error messages, and generating diffs that fix the underlying problem. Current-generation summarization models are capable of reading several scientific papers, extracting the key concepts, and turning them into a fairly serviceable podcast. All of this happens because a (1) in order to predict the next token better, LLMs actually build thousands of specialized models that do things like "keep track of the state of a chess board" or "recognize lions in photos", and (2) transformer models are sufficiently powerful to model many problems well. So a big current-generation LLM is basically an ensemble of thousands of domain models glued together with a language model and a bunch of feed-forward layers. Now, none of the outputs from LLMs are of super high quality, compared to experienced humans. And there are deep reasons for that. But if you happen to have problem where medium-competence AI "slop" is actually beneficial, then yes, LLMs can actually be of real-world use.