4 ms·
This is the same "tension" I keep seeing in my day job. Some people approach LLMs like they're writing code. They give a long list of detailed instructions for
by YZF 2mo ago
This is the same "tension" I keep seeing in my day job. Some people approach LLMs like they're writing code. They give a long list of detailed instructions for specific scenarios. When I use LLMs I leave things as open as possible. I just give them the information they need and my ask.
As you say frontier models are very good at figuring things out. Being too prescriptive is counterproductive, it over-constrains the model, it fills the context with conflicting instructions, it reduces the ability of the agent to respond to novel situations (and really in real life most situations are going to be novel). If you want to follow a process or a checklist you probably shouldn't use an LLM, or you should use it for some sub-tasks in the checklist/process but something more deterministic to work through the list.
- wonnage 2mo agoThat works for well trod paths, e.g “fix ci” works exceedingly well. “why app slow” obviously doesn’t work because the task is underspecified. But in order to properly specify you either need an experienced engineer who knows how to narrow the problem domain, or you have to provide some template instructions/output formats (e.g, skills) which will invariably never fit the problem perfectly
- gmadsen 2mo agoIt really doesn’t need to be that much more specified, give it context to the tools and level of analysis you expect then “why app slow” is a reasonable prompt
- hombre_fatal 2mo agoI wouldn't agree. Sota models can do self-directed sampling, profiling, benchmarking, read call trees, etc. to give you a report of the app's bottlenecks and then recommend solutions that can be vetted. I do this constantly. As the upstream comment points you, you don't need to specify. Sota models are that good. And by being overprescriptive you can accidentally shut off branches that they would've taken, downgrading the quality of their work.
- wonnage 2mo agoIn my experience if you’re at the point where you have something to sample then the hard part is already done. In a perfect world everything is covered by distributed tracing and the problems are only in your application code and the agent just needs to find the data In reality the data is often missing or misleading. “Your observability sucks”? Yeah, but that’s life
- YZF 2mo agoI use skills. The skills are not typically "how to perform a task in detail" they are more about what relevant tools and knowledge are required to work in a domain. That is I give the LLM the information it needs about the system but not a sequence of how to accomplish a task. I treat it more like a human and less like a computer.
- 0x457 2mo ago> . “why app slow” obviously doesn’t work because the task is underspecified. Not always. In my case LLM goes to grafana mcp, pulls metrics/traces/cpu profiles. Figures out what is slow and proposes a solution.
- rurban 2mo agoIn my cases it always used linux perf to sample the calls, because that's the best tool for my jobs. Never had to tell it to use instrumentation.
- pests 2mo ago> “why app slow” obviously doesn’t work because the task is underspecified Definitely not true and like everyone else is saying, shows how people still underestimate these models. I have been working on a simple vite + react app lately and commonly ask Gemini/Antigravity to just "improve speeds", "x is running slow, check it out" and have no complaints.
- wonnage 2mo agoI’m not surprised it works on a simple app.
- jimjimjim 2mo agodisturbingly, when I was using antigravity with gemini pro it was actually quite good at working out 'why app slow' types of problems. Maybe I've been lucky but it seems really good at determining why something might be wrong. It may ask for more logging or diagnostics and run for a long time but it was really digging in and making changes or suggestions to solve the problems.
- theptip 2mo agoHonestly I have had great success with “I’m worried here about cpu and latency, please rigorously profile and propose fixes”. The models can build micro-benchmarks with a level of rigor that few could muster for a new feature. I agree that if the issue is architectural they will struggle to understand that scope.
- msdz 2mo ago> When I use LLMs I leave things as open as possible. I just give them the information they need and my ask. How do you handle security? Both “internally” against e.g. data loss, I’m assuming via limiting the harness, and “externally”, i.e. stuff like prompt injection risks?
- YZF 2mo agoSandboxing and reviewing the output. I don't have any incredible insight to add here- that's the same process I think most of us are doing.
- avadodin 2mo agoThis vibe people sentiment is not wrong per se. If you want outlier performance from these models it is best to just ask in the most high level prompt of the most minimal harness and let them loose. Any extra information reduces their performance. However, as often as these models output masterpieces, they also produce utter garbage so our current choice is for them to have a process to follow that can be reviewed by humans and LLMs.
- visarga 2mo ago> If you want to follow a process or a checklist you probably shouldn't use an LLM I like to externalize tasks as markdown files with checklists, they are still planned by agents but I can pass the plan around to judge agents and fix some errors before implementing. I also have the coding agents comment on each closed checklist item, so the same file becomes a log of what happened. This goes to the implementation judge. I can also switch agents anytime, or resume a task days later no problem. I am avoiding internally provided tools for todo lists and planning because they do not leave the same artifact trail which makes judging with separate agents easy.
- stymaar 2mo agoThe problem is that even Fable still make trivial yet high impact mistake when let on their own, and then you'd need to read the whole code to catch them… Meanwhile they are very good at implementating an explicit algorithm that you feed it to them.
- theptip 2mo agoThe trick is to set up the harness so that the solution is easy to verify - you’ve profited as long as verification is cheaper than building, but ideally verification is close to automatic (not always achievable of course). Generally you want to include objective/repeatable outputs as citations. An example would be, if you invest in an awesome layered test rig (browser test, fuzz/property tests, very well reviewed unit/integration tests, etc.) then you should be able to add features by just reading the acceptance test and scanning unit tests.
- stymaar 2mo ago> then you should be able to add features by just reading the acceptance test and scanning unit tests. That “just” is bearing a lot of weight though as tests are often even longer than the code itself, in addition to being excruciating to review.
- pelasaco 1mo ago> Some people approach LLMs like they're writing code. They give a long list of detailed instructions for specific scenarios. When I use LLMs I leave things as open as possible. I just give them the information they need and my ask. Hm, but thats ok right? I mean some people like to code with LLM and other people like to let LLM code for them.. no?