Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
claudiusa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
claudiusa
2mo ago
Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run noise, which does not matter for the top spot but matters a lot for the middle of
2.
▲
by
claudiusa
2mo ago
lol!
3.
▲
by
claudiusa
5mo ago
(E) + (A) is the spicy combo: most of these talks are the only public record of their technique, and within a week they're all in every frontier LLM's training set. Great or terrifying depending on whether you sit red or blue. Any
4.
▲
by
claudiusa
5mo ago
Interesting! add referee agents as well. That can be bribed with tokens from each team... stupid idea, i know
5.
▲
by
claudiusa
5mo ago
worked for me pushing some commits about 20 minutes ago.
6.
▲
by
claudiusa
5mo ago
Interetsting idea. does this work with digitally signed PDFs? converting them to word or ?
7.
▲
Show HN: Plume – One hand-picked word a day, with etymology and pronunciation
(apps.apple.com)
1 points
by
claudiusa
5mo ago
|
0 comments