Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
uberdavid
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
uberdavid
5mo ago
Author here. DeepSeek-V4 replaced multi-domain RL with a specialist-then-distill pipeline: train domain experts independently, merge through on-policy distillation. This post connects that production decision to three recent papers (Neural
2.
▲
Random Machines: Why the Optimizer Is the Least Important Part of Deep Learning
(sotaverified.org)
1 points
by
uberdavid
6mo ago
|
1 comments
3.
▲
by
uberdavid
6mo ago
Author here. The core idea is that when you train the same model with different random seeds, both reach the same accuracy but disagree on ~10% of predictions. The reason connects three well-established results (loss landscape geometry, the
4.
▲
by
uberdavid
6mo ago
The system card directly compares to Opus 4.6 and other frontier models on the same evals. Cybench went from ~75% to 100%, Firefox exploitation from 1 bug unreliably to 4 bugs reliably. It's true there are many capable coding models ou
5.
▲
How RL Reward Hacking Made Claude Mythos a Zero-Day Hunter
(uberdavid.substack.com)
2 points
by
uberdavid
6mo ago
|
2 comments
6.
▲
The Dark Factory Harness: From Autonomous Hill-Climbing to Autonomous Research
(sotaverified.org)
2 points
by
uberdavid
6mo ago
|
1 comments
7.
▲
by
uberdavid
6mo ago
Author here. This is a synthesis of Karpathy's autoresearch (the experiment loop) and OpenAI's harness engineering post (the environment design) applied to ML research with 5 practical design principles. The core idea is that afte
8.
▲
by
uberdavid
6mo ago
Thank you! The progress on research agents is exciting, but understanding what papers are reproducible on different datasets and architectures is often the bottleneck.
9.
▲
SOTAVerified the open verification layer for ML research
(sotaverified.org)
2 points
by
uberdavid
6mo ago
|
3 comments
10.
▲
by
uberdavid
6mo ago
Hi HN, I'm David, an ML researcher at Meta. I built SOTAVerified as an independent project after Papers with Code shut down last year and took 575k papers worth of benchmark data with it. SOTAVerified inherits that dataset (658k papers