Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gertlabs
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
gertlabs
3mo ago
Scroll to the bottom for the methodology (sorry, this should be linkable)
32.
▲
by
gertlabs
3mo ago
GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models
33.
▲
by
gertlabs
4mo ago
Qwen 3.6 27B is an anomalously strong all-around model for its size, but when we run our evaluations, we generate 10 coding submissions/language/model (110 total). So full discosure, the per-language per-model performances can be
34.
▲
by
gertlabs
4mo ago
By domain, I really meant "tool calling" and "one-shot fluid intelligence" Anthropic models were the original leaders in tool calling and agentic work, even when other models felt significantly smarter in (Claude Sonnet
35.
▲
by
gertlabs
4mo ago
One-shot performance often translates to the most difficult problems a model will be able to understand. We run an evaluation that tests both agentic and one-shot performance, and we find that Chinese models are almost universally very good
36.
▲
by
gertlabs
4mo ago
They're within confidence intervals of each other, but remember how much discussion there was that Opus 4.6 had been nerfed in March. We averaged samples over the entire lifetime of Opus 4.6, which likely served many different underlyi
37.
▲
by
gertlabs
4mo ago
All of our posts have been well received by an insanely high percentage of people who have interacted on here -- most people clearly find what we're doing interesting and relevant to the HN community (AI evaluations). A flag seems pret
38.
▲
by
gertlabs
4mo ago
On our multi-agent coding and reasoning evaluations, GLM 5.2 is the first model we've tested that crossed the threshold of being on par with or better than Opus 4.6 (although as usual, we have GLM 5.2 and most other Chinese models a bi
39.
▲
by
gertlabs
4mo ago
GLM 5.2 is the first model we've tested that is unambiguously on par with, or better than Opus 4.6 (although as usual, we have GLM 5.2 and most other Chinese models a bit below most other benchmarks with more vulnerable test methodolog
40.
▲
by
gertlabs
4mo ago
It's likely overfit to common harnesses and iteration patterns, so it struggles with formatting tool calls and json in our testing which use our own harnesses (although there is a lot of overlap with tools that would be found in any co
41.
▲
by
gertlabs
4mo ago
DeepSeek v4 Pro struggles with a custom harness, and all the models ranked above it don't, so it gets downweighted in the agentic coding benchmarks (although it ranks better than Flash in one-shot problem solving: https://ge
42.
▲
by
gertlabs
4mo ago
MiMo V2.5 Pro (regular speed) remains the strongest open weights agentic coding model we've tested -- it's been interesting to see how little attention it has received relative to some lower performing releases. And the "fast
43.
▲
by
gertlabs
4mo ago
GPT 5.4+ models are extremely good at writing Clojure, agreed. In the agentic coding part of our benchmark, they do have access to the REPL via bash if they choose to use it. Filtered here: https://gertlabs.com/rankings?mode
44.
▲
by
gertlabs
4mo ago
I might just be a simpleton -- I never had the resolve to try an ambitious project in Clojure. I was not aware that you could get full OOP though, what you are describing feels like yes technically possible but kind of a hack to get inherit
45.
▲
by
gertlabs
4mo ago
Success rate includes syntax/compilation failures as well as environment rule violations, and is almost entirely from one-shot code generations. Percentile shows how well the working submissions perform. In long horizon agentic coding
46.
▲
by
gertlabs
4mo ago
The functional paradigm is a bit uncomfortable at first, but it does make problem solving feel... different. I personally find OOP to be the most intuitive for large scale systems design, but that's just me. Most models do not perform
47.
▲
by
gertlabs
4mo ago
Nice, that's a good one -- interesting dynamics can come out of deceptively simple social games.
48.
▲
Social Intelligence Benchmark
(gertlabs.com)
5 points
by
gertlabs
4mo ago
|
2 comments
49.
▲
by
gertlabs
4mo ago
OpenRouter is our primary provider for evaluation data, and we've been really happy with them! I'm sure they're experiencing growing pains, but a larger model selection (and faster releases for open weights models), would kee
50.
▲
by
gertlabs
4mo ago
That's an interesting point -- they are told that they are paper trading. Maybe we should run another session that A/B tests this.
51.
▲
LLM Paper Trading
(gertlabs.com)
6 points
by
gertlabs
4mo ago
|
5 comments
52.
▲
by
gertlabs
4mo ago
We're a month into a long running experiment to see how recent models perform in day trading, where they have a constant harness giving them the ability to write code, access the web, take notes, and install handlers to trade for them.
53.
▲
by
gertlabs
4mo ago
There is likely a theoretical limit to how much intelligence you can pack into a model of a given size (especially when stretching that over a large input context size). Our evals are pretty complex so we only recently started testing ~30B
54.
▲
by
gertlabs
4mo ago
That'll populate over the next couple weeks -- those are the live games on the spectate tab which take a while to generate statistically worthwhile data. I'm curious how it does. From using it all day, I can say Opus 4.8 is my new
55.
▲
by
gertlabs
4mo ago
Appreciate that! Results are live: https://gertlabs.com/rankings Opus 4.8 is the first tangible improvement since Opus 4.5. And it doesn't seem to have the personality problems of the last release -- I've been enj
56.
▲
by
gertlabs
4mo ago
We just finished our initial coding evals of Opus 4.8. Anthropic definitely heard the backlash from Opus 4.7 and they made up for it today. Subjectively, it's also quite enjoyable to use (although it feels a bit slower on max reasoning
57.
▲
by
gertlabs
4mo ago
4.5/4.6 were roughly the same in our testing. Opus 4.7 is smarter, but it's difficult to use as a product for various personality issues. So far, Opus 4.8 seems to be going down that path (unusably slow, but this could be a launch
58.
▲
by
gertlabs
5mo ago
GPT 5.5 does significantly outperform Opus 4.7 in the coding parts of our evals. We also incorporate live decision making on social games (where GPT 5.5 has actually regressed from earlier models, which shouldn't be a huge surprise if
59.
▲
by
gertlabs
5mo ago
Check out the methodology section at the bottom -- we are trying to better convey this information. 1. These numbers are based on percentiles, which inherently can't be saturated. Most benchmarks operate on something like 0-100% of cor
60.
▲
by
gertlabs
5mo ago
While this benchmark has interesting results, the "Contamination free" label only works for the initial release of the benchmark. It still has the same fundamental design issues of any other benchmark-- there's a single corre
More ›