Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
declanjackson
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
declanjackson
2mo ago
With max reasoning, Luna is actually less than half the actual cost to run compared to DeepSeek V4 Flash 0731 with updated prices (based on Artificial Analysis Cost per Task)
2.
▲
Kimi K3 beats GPT 5.6 Sol in agentic knowledge work
(artificialanalysis.ai)
3 points
by
declanjackson
2mo ago
|
0 comments
3.
▲
GLM-5.2 is above GPT-5.5 in new agentic knowledge work eval
(artificialanalysis.ai)
5 points
by
declanjackson
3mo ago
|
0 comments
4.
▲
Show HN: AA-Briefcase: a frontier knowledge work evaluation
(artificialanalysis.ai)
13 points
by
declanjackson
3mo ago
|
2 comments
5.
▲
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Language Models
(arxiv.org)
6 points
by
declanjackson
11mo ago
|
1 comments