Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bisonbear
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
bisonbear
5mo ago
Not the OP, but I've been thinking about this problem a lot - as devs we're overly reliant on vibes for evaluating coding agents. This is already a problem, and especially so if you're working in an engineering organization w
32.
▲
The Opus 4.7 reasoning curve - Medium is the best default?
(stet.sh)
1 points
by
bisonbear
5mo ago
|
0 comments
33.
▲
by
bisonbear
5mo ago
Claude does appear to work for longer, and use more tokens, when at higher reasoning modes. It just doesn't seem like this increased token usage leads to better actual outcomes
34.
▲
by
bisonbear
5mo ago
Agree, it's impossible to tell if someone else's workflow works with your codebase without actually trying it, which takes time/tokens. I've been thinking about how to make running quick, directional evals easier /
35.
▲
by
bisonbear
5mo ago
I'm actually currently working on benchmarking the opus 4.7 reasoning curve against real-world tasks, and have found that reasoning effort does not seem to monotonically improve results (at least on the slice I'm looking at). I&#x
36.
▲
GPT-5.5 low vs. medium vs. high vs. xhigh: the reasoning curve on 26 real tasks
(stet.sh)
2 points
by
bisonbear
5mo ago
|
0 comments
37.
▲
GPT-5.5 vs. GPT-5.4 vs. Opus 4.7 on 56 real coding tasks from 2 open source repo
(stet.sh)
4 points
by
bisonbear
5mo ago
|
0 comments
38.
▲
by
bisonbear
6mo ago
> they nerfed 4.6 to make way for 4.7? > Progress. /s pretty much, lmao. my theory is 4.6 started thinking less to save compute for 4.7 release. but who knows what's going on at anthropic
39.
▲
I ran Opus 4.7 vs. Old Opus 4.6 vs. New Opus 4.6 on 28 Zod tasks
(stet.sh)
2 points
by
bisonbear
6mo ago
|
0 comments
40.
▲
by
bisonbear
6mo ago
yep, ran a controlled experiment on 28 tasks comparing old opus 4.6 vs new opus 4.6 vs 4.7, and found that 4.7 is comparable in cost to old 4.6, and ~20% more expensive then new 4.6 (because new 4.6 is thinking less) https://www.
41.
▲
by
bisonbear
6mo ago
coming more in line with codex - claude previously would often ignore explicit instructions that codex would follow. interested to see how this feels in practice I think this line around "context tuning" is super interesting - I s
42.
▲
Coding evals are broken. CI is green while AI code quality goes unmeasured
(stet.sh)
1 points
by
bisonbear
6mo ago
|
0 comments
43.
▲
by
bisonbear
6mo ago
working on something similar to evaluate model performance over time using tasks based on your own code. obviously this is still susceptible to the same hacking mechanics documented here, but at a local level, it's easier to detect
44.
▲
Agents.md is the highest-leverage code you're not testing
(stet.sh)
1 points
by
bisonbear
6mo ago
|
0 comments
45.
▲
by
bisonbear
6mo ago
a bit heavier weight, but seems worthwhile if working in an org where many people consume the skill: - find N tasks from your repo that serve as good representation of what you want the agent to do with the task - run agent with old skill&#
46.
▲
by
bisonbear
6mo ago
PRs for AGENTS.md are necessary, but not sufficient, exactly because of non-determinism. You can LGTM the AGENTS.md change, but it's so hard to know what downstream behavioral effects it has. I feel like the only way to really know i
47.
▲
by
bisonbear
6mo ago
Very cool, interested to read more once you post! FWIW I've been building eval infras that does something adjacent/related — replaying real repo work against different agent configs, and measuring the agent's quality dimensio
48.
▲
by
bisonbear
6mo ago
cost control is a policy problem - we certainly don't need to use opus 4.6 for a simple test refactor, but many people (including myself) default to it anyways. we need a way to measure cost / performance for agents on individual
49.
▲
by
bisonbear
6mo ago
managing agents.md is important, especially at scale. however I wonder how much of a measurable difference something like this makes? in theory, it's cool, but can you show me that it's actually performing better as compared to
50.
▲
by
bisonbear
6mo ago
I'm also thinking on how we can put guardrails on Claude - but more around context changes. For example, if you go and change AGENTS.md, that affects every dev in the repo. How do we make sure that the change they made is actually bene
51.
▲
by
bisonbear
7mo ago
I'm becoming convinced that test pass rate is not a great indicator of model quality - instead we have to look at agent behavior beyond the test gate, such as how aligned is it with human intent, and does it follow the repo's co
52.
▲
by
bisonbear
7mo ago
I agree with your analysis but not the conclusion. Evals are broken - OpenAI showed that SWE Bench Verified was in the training data - models were able to reconstruct the changes from memory ( https://openai.com/index/wh
53.
▲
by
bisonbear
7mo ago
Really interesting study. One thing I keep coming back to is that tests have no way of catching this sort of tech debt. The agent can introduce something that will make you rip your hair out in 6 months, but tests are green... My theory is
54.
▲
by
bisonbear
7mo ago
curious how you measure/track how this actually impacts the coding agent?
55.
▲
by
bisonbear
7mo ago
For agentic development teams, I see there being two ways to measure performance: How good is the human at using the agent, and how good is the agent itself? I agree with the thesis here that the traditional DORA metrics don't have as
56.
▲
Your AI coding benchmark is hiding a 2x quality gap
(stet.sh)
3 points
by
bisonbear
7mo ago
|
0 comments
57.
▲
by
bisonbear
7mo ago
yikes, using AI without tests is not fun. with testing at least you have some confidence that the AI isn't going completely off track, without them you're pretty much flying blind having linters is super important IMO - I never
58.
▲
by
bisonbear
7mo ago
yea I'm down - feel free to send me an email ben@benr.build
59.
▲
by
bisonbear
7mo ago
I've been working on building out "evals for your repo" based on the theory that commonly used benchmarks like SWE-bench are broken as they are not testing the right / valuable things, and are baked into the training dat
60.
▲
by
bisonbear
7mo ago
sounds like it's another openclaw-as-a-service provider?
More ›