Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pongogogo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
pongogogo
3mo ago
They note in the paragraph I quoted at the top that prompting has a big impact on behaviour, so yes this would work. I think that's not what METR are interested in though.
2.
▲
by
pongogogo
3mo ago
I would say this is quite a fun post and worth reading, to quote: " For our task suite, we define “cheating” as behavior where the model improves evaluation performance by exploiting bugs in the evaluation environment or by adopting st
3.
▲
Summary of METR's predeployment evaluation of GPT-5.6 Sol
(metr.org)
10 points
by
pongogogo
3mo ago
|
6 comments
4.
▲
How we made Ramp Sheets self-maintaining
(twitter.com)
2 points
by
pongogogo
6mo ago
|
0 comments
5.
▲
by
pongogogo
7mo ago
Beautiful site, worth a read.
6.
▲
The Self-Driving Codebase
(background-agents.com)
1 points
by
pongogogo
7mo ago
|
1 comments
7.
▲
The Bitter Lesson of Agent Frameworks
(twitter.com)
3 points
by
pongogogo
9mo ago
|
1 comments
8.
▲
Don't Build Agents, Build Skills Instead [video]
(youtube.com)
1 points
by
pongogogo
10mo ago
|
0 comments
9.
▲
AI in 2025: Gestalt
(lesswrong.com)
3 points
by
pongogogo
10mo ago
|
0 comments
10.
▲
by
pongogogo
1y ago
It's hard to tell from the data, it's so concentrated within a handful of companies who are all buying from eachother, so it feels like the contagion risk is low. At the same time it feels very clearly overvalued and the size of t
11.
▲
The AI Bubble and the US Economy
(mronline.org)
2 points
by
pongogogo
1y ago
|
3 comments
12.
▲
When Will Quantum Computing Work?
(tommccarthy.net)
1 points
by
pongogogo
1y ago
|
0 comments
13.
▲
Supporting our AI overlords: Redesigning data systems to be Agent-first
(muratbuffalo.blogspot.com)
3 points
by
pongogogo
1y ago
|
0 comments
14.
▲
Post-Training 101
(tokens-for-thoughts.notion.site)
2 points
by
pongogogo
1y ago
|
0 comments
15.
▲
Generative Engine Optimization: How to Dominate AI Search
(arxiv.org)
3 points
by
pongogogo
1y ago
|
1 comments
16.
▲
by
pongogogo
1y ago
The post mentions an approach of using a large model to generate labels and then distilling this into a smaller model to lower cost (though it doesn't provide an example)
17.
▲
LLMs as Retrieval and Recommendation Engines
(medium.com)
3 points
by
pongogogo
1y ago
|
2 comments
18.
▲
EnvX: Agentize Everything with Agentic AI
(arxiv.org)
1 points
by
pongogogo
1y ago
|
0 comments
19.
▲
VLLM: Anatomy of a High-Throughput LLM Inference System
(aleksagordic.com)
3 points
by
pongogogo
1y ago
|
0 comments
20.
▲
Why language models hallucinate [pdf]
(cdn.openai.com)
2 points
by
pongogogo
1y ago
|
0 comments
21.
▲
Agent Hopper: An AI Virus
(embracethered.com)
3 points
by
pongogogo
1y ago
|
0 comments
22.
▲
by
pongogogo
1y ago
I've been meaning to write a post like this for a while but you've done a much better job.
23.
▲
The Limits of Reinforcement Learning
(itcanthink.substack.com)
1 points
by
pongogogo
1y ago
|
0 comments
24.
▲
The Bull Case for an AI Native Investment Bank
(open.substack.com)
1 points
by
pongogogo
1y ago
|
0 comments
25.
▲
Why does Deepseek-R1 hallucinate so much?
(vectara.com)
2 points
by
pongogogo
1y ago
|
0 comments
26.
▲
by
pongogogo
1y ago
Yes, I wrote something up here on how Andrei Kaparthy evaluated grok 3 -> https://tomhipwell.co/blog/karpathy_s_vibes_check/ I would pick one of two parts of that analysis that are most relevant to you and zoom
27.
▲
by
pongogogo
1y ago
I think this is a really interesting paper from Cohere, it really feels that at this point in time you can't trust any public benchmark, and you really need your own private evals.
28.
▲
The Leaderboard Illusion
(arxiv.org)
184 points
by
pongogogo
1y ago
|
51 comments
29.
▲
Guillotine: Hypervisors for Isolating Malicious AIs
(arxiv.org)
2 points
by
pongogogo
1y ago
|
0 comments
30.
▲
by
pongogogo
1y ago
Hey Mark, I actually found this post via yours so thanks!
More ›