Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
natolambert
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
Learning to solve hard problems in RL for LLMs by never giving up
(mnoukhov.github.io)
52 points
by
natolambert
16d ago
|
0 comments
2.
▲
Show HN: Colloquium – a Markdown-native slide tool for academics
(github.com)
2 points
by
natolambert
7mo ago
|
0 comments
3.
▲
The Atom Project
(atomproject.ai)
13 points
by
natolambert
1y ago
|
4 comments
4.
▲
by
natolambert
1y ago
ATOM: American Truly Open Models
5.
▲
by
natolambert
2y ago
Author here! Just wanted to say that this is indeed in a good place to share, some very useful stuff, but is also very work in progress. I'm may 60% or so to my first draft. Said progress is coming every day and I happily welcome fixes
6.
▲
by
natolambert
2y ago
As the other commenter said, R1 required very standard RLHF techniques too. But a fun way to think about it is that reasoning models are going to be bigger and uplift the RLHF boat. But we need a few years to establish basics before I can w
7.
▲
by
natolambert
2y ago
thx, you're right. I've fixed it!
8.
▲
by
natolambert
3y ago
Author here, you can find more on my website: https://natolambert.com/cv Have been building RLHF systems at HuggingFace since ChatGPT, with some other experience before.
9.
▲
How RLHF Works
(interconnects.ai)
165 points
by
natolambert
3y ago
|
32 comments
10.
▲
Different Development Paths of LLMs
(interconnects.ai)
1 points
by
natolambert
3y ago
|
0 comments
11.
▲
Evaluating and Uncovering Open LLMs
(interconnects.ai)
1 points
by
natolambert
3y ago
|
0 comments
12.
▲
Unfortunately, OpenAI and Google have moats
(interconnects.ai)
2 points
by
natolambert
3y ago
|
0 comments
13.
▲
Specifying Hallucinations of LLMs
(interconnects.ai)
1 points
by
natolambert
3y ago
|
0 comments
14.
▲
The last reliable path into AI research
(natolambert.com)
1 points
by
natolambert
5y ago
|
0 comments
15.
▲
Remote Robotic-Data Farms
(robotic.substack.com)
1 points
by
natolambert
5y ago
|
0 comments
16.
▲
Reward Is Not Enough
(robotic.substack.com)
2 points
by
natolambert
5y ago
|
0 comments
17.
▲
All machine learning becomes reinforcement learning
(robotic.substack.com)
1 points
by
natolambert
5y ago
|
0 comments
18.
▲
Debugging model-based reinforcement learning systems
(natolambert.com)
2 points
by
natolambert
5y ago
|
0 comments
19.
▲
Setting ourselves up for exploitation: RL in the wild
(robotic.substack.com)
1 points
by
natolambert
6y ago
|
0 comments
20.
▲
Decoupling AI from the latent variable of spoken languages
(robotic.substack.com)
2 points
by
natolambert
6y ago
|
0 comments
21.
▲
Covid didn’t give us personal robots, it gave us Woebot
(robotic.substack.com)
1 points
by
natolambert
6y ago
|
0 comments
22.
▲
Boston Dynamics: Studying Athletic Intelligence
(robotic.substack.com)
1 points
by
natolambert
6y ago
|
0 comments
23.
▲
Robotic Startups 2.0: Horizontal Modularity
(democraticrobots.substack.com)
1 points
by
natolambert
6y ago
|
0 comments
24.
▲
Social Networks and Degradation to the Public Square of Discourse
(democraticrobots.substack.com)
2 points
by
natolambert
6y ago
|
1 comments
25.
▲
Our idea of Free Will can bias AIs we create
(democraticrobots.substack.com)
2 points
by
natolambert
6y ago
|
0 comments
26.
▲
The Ubiquity and Future of Model-Based Reinforcement Learning
(democraticrobots.substack.com)
2 points
by
natolambert
6y ago
|
0 comments
27.
▲
Constructing Axes for (Legal) Reinforcement Learning Policy
(democraticrobots.substack.com)
2 points
by
natolambert
6y ago
|
0 comments
28.
▲
The Collingridge Dilemma and Current Policy on Robots
(democraticrobots.substack.com)
2 points
by
natolambert
6y ago
|
0 comments
29.
▲
EE's See the World in Models
(democraticrobots.substack.com)
2 points
by
natolambert
6y ago
|
0 comments
30.
▲
Automated: The levers tech companies pull to direct our lives
(democraticrobots.substack.com)
1 points
by
natolambert
6y ago
|
0 comments
More ›