Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
skysniper
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
1.
▲
Show HN: Inter-session messaging between Claude Code sessions
(github.com)
1 points
by
skysniper
5mo ago
|
0 comments
2.
▲
Opus 4.7 dominates agentic benchmark, 15% more expensive than Opus 4.6
(app.uniclaw.ai)
3 points
by
skysniper
6mo ago
|
1 comments
3.
▲
by
skysniper
6mo ago
Ran preliminary benchmarks on Opus 4.7, noticeably better than Opus 4.6, about 15% higher cost per task due to more tool calls, most performant and expensive model so far
4.
▲
GLM-5.1 matches Opus 4.6 in agentic performance, at ~1/3 actual cost
(app.uniclaw.ai)
22 points
by
skysniper
6mo ago
|
2 comments
5.
▲
by
skysniper
6mo ago
where are you Mythos
6.
▲
by
skysniper
6mo ago
check out my reply, his chart is plotting the wrong metric (average quality score)
7.
▲
by
skysniper
6mo ago
i added native plot and stats for aggregated results, on arena page. please check it out!
8.
▲
by
skysniper
6mo ago
yeah but i'm not using the free version for benchmark...
9.
▲
by
skysniper
6mo ago
added https://app.uniclaw.ai/arena/model-stats also added per battle stats in battle detail page
10.
▲
by
skysniper
6mo ago
I know, that was indeed a bad judge move. I've manually checked tens of tasks so far, and that one is one of the worst... I would say check a few more, judge has some noise but in general did a good job IMO
11.
▲
by
skysniper
6mo ago
well, I still want to use it but the first day i tried openclaw + opus, it costs me ~$500...
12.
▲
by
skysniper
6mo ago
> The explanation is that network errors were credited with a quality score of 0, and there were _a lot_ of network errors. all network error, provider error, openclaw error are excluded from ranking calculation actually, so that is not
13.
▲
by
skysniper
6mo ago
I will try and add it. But I doubt it works well because Mimo V2 Pro is beaten by stepfun even at performance leaderboard (price is not a factor in this leaderboard), so I expect MiMo V2 Flash to perform even worse.
14.
▲
by
skysniper
6mo ago
TBH that was my initial thought too, but I found some problem using this approach: Essentially I'm using the relative rank in each battle to fit a latent strength for each model, and then use a nonlinear function to map the latent stre
15.
▲
by
skysniper
6mo ago
it's actually pretty good at openclaw type of tasks for non technical users: lots of tool calls, some simple programing
16.
▲
by
skysniper
6mo ago
what kind of tasks did you try?
17.
▲
by
skysniper
6mo ago
I appreciate honest feedback, best way to learn :)
18.
▲
by
skysniper
6mo ago
both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just want to see cost in USD so I put token details at the bottom)
19.
▲
by
skysniper
6mo ago
I should have clarified I didn't use the free version...
20.
▲
by
skysniper
6mo ago
> Is the judge an LLM? Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when computing ranking. > There's lot of references to &qu
21.
▲
by
skysniper
6mo ago
another thing from the bench I didn't expect: gemini 3.1 pro is very unreliable at using skills. sometimes it just reads the skill and decide to do nothing, while opus/sonnet 4.6 and gpt 5.4 never have this issue.
22.
▲
by
skysniper
6mo ago
thanks for the info. before running the bench i only tried it in arena.ai type of tasks and it was not impressive. i didn't expect it to be that good at agentic tasks
23.
▲
by
skysniper
6mo ago
all 300+ battle data are available at https://app.uniclaw.ai/arena/battles , every single battle is shown with raw conversional history, produced files, judge's verdict and final scores
24.
▲
by
skysniper
6mo ago
the real surprising part to me is that, despite being the cheapest model on board, stepfun is often able to score high at pure performance. Other models at the same price range (e.g. kimi) fails to do that.
25.
▲
by
skysniper
6mo ago
sorry didn't know that. Here is my hand writing tldr: gemini is very unreliable at using skills, often just read skills and decide to do nothing. stepfun leads cost-effectiveness leaderboard. ranking really depends on tasks, better try
26.
▲
StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)
(app.uniclaw.ai)
175 points
by
skysniper
6mo ago
|
84 comments
27.
▲
by
skysniper
6mo ago
I ran 300+ benchmarks across 15 models in OpenClaw and published two separate leaderboards: performance and cost-effectiveness. The two boards look nothing alike. Top 3 performance: Claude Opus 4.6, GPT-5.4, Claude Sonnet 4.6. Top 3 cost-ef
28.
▲
Show HN: OpenClaw Arena – Benchmark models on real tasks, rank by perf and cost
(app.uniclaw.ai)
2 points
by
skysniper
6mo ago
|
0 comments