4 ms·
Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project. Deepseek 4 pro: Worked for
by jklmnopqrstuvw 2mo ago
Tested both DS v4 pro 0813 and Grok 4.6 (all from openrouter) on Codex cli. Worked on a same new feature development on my project.
Deepseek 4 pro: Worked for 12m 02s - cost $0.12 - has bug.
Grok 4.6: Worked for 3m 18s - cost $ 1.41 - no bug.
- hugmynutus 2mo agoNullius in verba
- computerex 2mo agoRepeat the test like 5 times for each model and see the results.
- epolanski 2mo ago+1, a single test means little.
- jklmnopqrstuvw 2mo agoI don't think so. I specifically kept this PR to test model capabilities, and I've already tested a bunch of models. Current test results show that the more advanced the model is, the easier it passes. For example, GPT-5.5 Medium fails the test(has bug), but High passed.
- seunosewa 2mo agoDo it a second time at least.
- computerex 2mo agoThey are causal autoregressive models, the output is sensitive even to the implementation nuances in inference. Even 1 token that's badly selected could throw off the entire answer.
- segmondy 2mo agoyou're thinking of one shot. if they are running an agentic loop then they don't need multiple passes. an agentic loop is multiple passes with tool calls and tools could fail and agent would correct from seeing the failure. a bad model will compound on error and fail, a good model will correct. 1 test is fine to gauge the quality of the model.
- computerex 2mo agoAn agent doing a task even with multiple back to back calls like normal without an example is zero shot. An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms. The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session out of whack, because LLM's build on the previous context. It's actually meaningless to argue, one could simply sample more than 1 times and let the numbers speak for themselves.
- gpt5 2mo agoThat's not true. An agent in a loop can test itself, review, verify and iterate as much as needed. That's one of the primary reasons more capable models tend to have a higher success rate. I don't disagree that multiple tests increase confidence, but it's not correct to argue that an agent in a loop harness is equivalent to oneshotting
- techpression 2mo agoIt takes more than that. I had Claude do a feature, it took five minutes, then I had Claude do a code review of its changes, that took 65(!) separate agents and 40minutes. I don’t think single agent loops are good enough.
- nl 2mo ago> An agent doing a task with 1 example is one shot. An agent doing a task with a few examples is few shot. I don't think you are correctly using these terms This is a different thing. Yes, giving multiple example is called "few-shot prompting". But one-shot vs few-shot benchmarking is different. In this context "one-shot" means "pass at 1 effort" as opposed to "multi-shot". In the literature this is called "pass@k". Anthropic has a good explanation here: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents https://www.anthropic.com/engineering/demystifying-evals-for... (search for "pass@k"). In this discussion we are discussing pass@1 (single shot) vs pass@(k>1) (multi shot). > The multiple back to back LLM calls are done on accumulating context, so if there is a sampling error it could throw the entire session out of whack, because LLM's build on the previous context. This isn't really true. In an agentic loop the LLM can correct itself via in-context learning.
- ferongr 2mo ago[flagged]
- nozzlegear 2mo agoThis but unironically
- deleted 2mo ago[deleted]
- SV_BubbleTime 2mo ago[flagged]
- aliasxneo 2mo agoThere are still sane people here, we just don't talk about "rocket man" because it enrages the particular subgroup on display here and usually goes no where actually productive (and has like a 50% of getting flagged to death anyways). I'm not pro Elon by any means, but the standard HN profile of him is pretty bat shit crazy.
- sergiotapia 2mo agoCorrect
- dgellow 2mo agoHe is honestly batshit insane
- nozzlegear 2mo agoIs this bait? > So.. flesh this out and don’t be a coward about it. No.
- SV_BubbleTime 2mo agoTracks. The position seems to be faith based, so I understand why you are uncomfortable with its shape.
- NooneAtAll3 2mo agoI thought it was impossible to downvote posts?
- numpad0 2mo agoMaybe a tug of war between flags and vouches might work like downvotes?
- benjiro29 2mo agoI thought it was impossible to downvote posts? User Posts can be downvoted but you need over 500 karma to have access to the downvote button. A Submission can not be downvoted.
- Barbing 2mo agoYes and- Submissions can be flagged by anyone and mods/admins can downweight them. (If I’m not mistaken this is common for, say, Flock posts at the moment.) Curiosity & repetition are two key factors.
- NooneAtAll3 2mo agocomments can be downvoted, posts can't
- deleted 2mo ago[deleted]
- Zetaphor 2mo agoIt's the third link on the front page right now?
- bigmadshoe 2mo agoWhy are people giving these n=1 comparisons like they mean anything? The worst offender is that pelican guy. These are non-deterministic systems and a single trial should not update your priors much at all. Of course it's significant that your response had a bug and took four times longer, but if you're only going to try once, this isn't real science, it's just vibes.
- jklmnopqrstuvw 2mo agoMonths ago I start making this kind of test for my own reference. At beginning I I test each model multiple times, and results always same(pass or fail). Later I test only once for new models, I trust the results.
- kees99 2mo ago> multiple times, and results always same Not my experience at all. With smaller models, whenever I see a response that is going into wrong direction, I would just redo that step, and more often that not that brings improvement. This effect is less pronounced with SOTA, but still there.
- shunia_huang 2mo agoYes not my experience either. I've tried or sometimes be stupid to work on bugs/features and ask with almost identical prompts with same modal and harness set, and yes, they generate totally different results. Sometimes the output is unusable and even with extended guidance it will still drift away from what I was expecting. Sometimes the output is just one shot and follows almost whatever I want. I then be used to work like this, if the model and harness set does not work for one time, I just start a new session and do it again. And currently there is one of my task working like this.
- gnunez 2mo agoI don’t understand why people are calling these transformers models non-deterministic? Are you referring to the temperature parameter? I haven’t played with transformer internals in a while but my understanding is that if the temperature is fixed at a value where the top logit is always picked, then because they weights are fixed, the exact same input should produce the exact same output. Am I missing something?