3 ms·
Repeat the test like 5 times for each model and see the results.
by computerex 2mo ago
Repeat the test like 5 times for each model and see the results.
- epolanski 2mo ago+1, a single test means little.
- jklmnopqrstuvw 2mo agoI don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.
- seunosewa 2mo agoDo it a second time at least.
- computerex 2mo agoThey are causal autoregressive models, the output is sensitive even to the implementation nuances in inference. Even 1 token that's badly selected could throw off the entire answer.
- segmondy 2mo agoyou're thinking of one shot. if they are running an agentic loop then they don't need multiple passes. an agentic loop is multiple passes with tool calls and tools could fail and agent would correct from seeing the failure. a bad model will compound on error and fail, a good model will correct. 1 test is fine to gauge the quality of the model.
- computerex 2mo agoAn agent doing a task even with multiple back to back calls like normal without an example is zero shot. An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms. The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session out of whack, because LLM's build on the previous context. It's actually meaningless to argue, one could simply sample more than 1 times and let the numbers speak for themselves.
- gpt5 2mo agoThat's not true. An agent in a loop can test itself, review, verify and iterate as much as needed. That's one of the primary reasons more capable models tend to have a higher success rate. I don't disagree that multiple tests increase confidence, but it's not correct to argue that an agent in a loop harness is equivalent to oneshotting
- techpression 2mo agoIt takes more than that. I had Claude do a feature, it took five minutes, then I had Claude do a code review of its changes, that took 65(!) separate agents and 40minutes. I don’t think single agent loops are good enough.
- nl 2mo ago> An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms This is a different thing. Yes, giving multiple example is called "few-shot prompting". But one-shot vs few-shot benchmarking is different. In this context "one-shot" means "pass at 1 effort" as opposed to "multi-shot". In the literature this is called "pass@k". Anthropic has a good explanation here: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents https://www.anthropic.com/engineering/demystifying-evals-for... (search for "pass@k"). In this discussion we are discussing pass@1 (single shot) vs pass@(k>1) (multi shot). > The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session out of whack, because LLM's build on the previous context. This isn't really true. In an agentic loop the LLM can correct itself via in-context learning.