Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sam-paech
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
sam-paech
1y ago
Those higher level kinds of mode collapse are hard to quantify in an automated way. To fix that, you would need interventions upstream, at pre & post training. This approach is targeted to the kinds of mode collapse that we can meaningf
2.
▲
by
sam-paech
1y ago
All the judge outputs (including rubric) and model outputs are in the samples reports. Sorry you don't like the displayed metrics. I find them very useful / revealing of the things I'm trying to measure with this benchmark.
3.
▲
by
sam-paech
1y ago
None of those factors go into the scoring fwiw. They are just informational. The scoring is done to a rubric, like a teacher would grade an essay, on various criteria for good & bad writing.
4.
▲
by
sam-paech
1y ago
Different benchmark, those are for the short form creative writing leaderboard here: https://eqbench.com/creative_writing.html
5.
▲
by
sam-paech
1y ago
Personally what I find interesting is getting insight into the trajectory of model abilities over time. Over the time I've been running these benchmarks, the writing has gone from pure slop, to broadly competent (at short form at least
6.
▲
by
sam-paech
1y ago
Oops, should be: https://eqbench.com/creative_writing.html Sample outputs: https://eqbench.com/results/creative-writing-v3/gemini-2.5-p...
7.
▲
by
sam-paech
1y ago
Hey, I made this! Cool to see it show up on hackernews.
8.
▲
by
sam-paech
1y ago
The old version of the creative writing eval had several "in the style of" prompts actually! But I got tired of reading bad Hemingway impersonations so I cut them out of the new version. The creative writing v3 prompts ( https:&#x
9.
▲
by
sam-paech
1y ago
Not internal consistency exactly, but there are criteria checking how well the chapter plan was followed (which is all the way up at the top of the context window). This is done per chapter, and the score trendline is what you see in the &q