Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
danielmarkbruce
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
danielmarkbruce
27d ago
No, it won't. How do you verify some causal claim in biology? The reason AI is doing so well in math proof writing is that it can verify every idea it has, quickly.
32.
▲
by
danielmarkbruce
27d ago
Maybe. Maybe not. Look at AI drug design - it's not really speeding up the important part - drug trials. There isn't really a coherent plan to use AI for the most complex part of drug discovery at all.
33.
▲
by
danielmarkbruce
27d ago
If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.
34.
▲
by
danielmarkbruce
27d ago
Highly capable of writing math proofs, no doubt. It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue
35.
▲
by
danielmarkbruce
27d ago
If you look closely at the gains in math, it's largely in proof writing. The reason is Lean, it's not some general intelligence jump, and the number of people actually working on proofs in life rounds to zero.
36.
▲
by
danielmarkbruce
27d ago
With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't sp
37.
▲
by
danielmarkbruce
28d ago
Ok. This all makes sense I guess. Good going, hope it goes well.
38.
▲
by
danielmarkbruce
28d ago
It's more responsible than the use of electricity for a messaging board for people to argue minutiae.
39.
▲
by
danielmarkbruce
28d ago
Why not specify 2-3 problems? Won't you have people show up having already spent a bunch of time on their self chosen problem? Which sort of defeats the point of seeing what you can do in a short period of time?
40.
▲
by
danielmarkbruce
29d ago
But the next move is self evident if you have a prediction of the value of being in each of the states possible.
41.
▲
by
danielmarkbruce
29d ago
Your initial comment says "reward". Reward and return are not the same thing. The policy is choosing the moves based off of returns at the next state, not the immediate rewards. One choice might have reward 0 and expected return 1
42.
▲
by
danielmarkbruce
1mo ago
I'm not the one hiding behind a fake name. If you want to understand how this stuff works, there are totally decent books about building them from scratch. It's not that hard, and you'll likely find it interesting. Sebastian
43.
▲
by
danielmarkbruce
1mo ago
When you are done with the section on RLVR, consider whether the model is predicting tokens, or making moves. There is a reason the word "policy" is used in RL.
44.
▲
by
danielmarkbruce
1mo ago
Many people are in fact claiming the thing you are saying they are not - even if you are not. The reason they are claiming it is that it was true at one point, and most intro courses/blog posts/videos still describe them that way
45.
▲
by
danielmarkbruce
1mo ago
Brush up :) The policy is optimized to maximize the the total reward, defined as the sum of the reward at each step, discounted by some factor.
46.
▲
by
danielmarkbruce
1mo ago
Mine was sarcasm. People who actually understand cars have built them. Until you build something, you don't understand it.
47.
▲
by
danielmarkbruce
1mo ago
Nathan Lambert wrote a good book recently, and he and his team wrote the paper below about Tulu 3 (Allen Institute). Both are good reads. https://arxiv.org/pdf/2411.15124
48.
▲
by
danielmarkbruce
1mo ago
The policy is not learned token by token during RLHF and RLVR. The reward model doesn't score token by token.
49.
▲
by
danielmarkbruce
1mo ago
A guess at what though? One guesses at truths they don't know, or events that haven't happened yet. What is the model guessing?
50.
▲
by
danielmarkbruce
1mo ago
So, this is the cause of the problem.... People take an intro to LLMs course, follow happily along, and don't realize there is more to it than the next token prediction. And those courses teach how LLMs were built in 2017-2020 maybe. T
51.
▲
by
danielmarkbruce
1mo ago
Assuming you are saying that RL is changing the model from doing one thing to another, yes. RL is changing the nature of the model.
52.
▲
by
danielmarkbruce
1mo ago
The discussion is basically: what is a model trying to do? One may reasonably assert it isn't trying to do anything. But, in practice, if you give it an objective function and optimize it, the model is basically trained to "do&quo
53.
▲
by
danielmarkbruce
1mo ago
No one is arguing about the architecture of the model. It's the objective function and optimizer.
54.
▲
by
danielmarkbruce
1mo ago
This is pedantic, but, actually RL has improved the quality of sentence construction in LLMs quite dramatically... And once you do some RL on that model, it aint a next token prediction machine any longer.
55.
▲
by
danielmarkbruce
1mo ago
Probably the easiest way to describe an LLM that it's a policy. There is a reason that word has stuck in RL. And it's not just RLVR. RLHF has been going on for years and years. LLMs have not been "next token predictors"
56.
▲
by
danielmarkbruce
1mo ago
It's not an estimation of something. It's a policy.
57.
▲
by
danielmarkbruce
1mo ago
You are conflating "half built" with "a piece of a system". The model weights change as the model goes through the training process. They aren't stored after pre-training is done and other weights are put somewhere
58.
▲
by
danielmarkbruce
1mo ago
Yes, it was. Nobody says a system or product works a certain way and means the system while it's half built. "Bridges drop cars in the water!". Right. You aren't in this field. You are clearly wrong and just can't h
59.
▲
by
danielmarkbruce
1mo ago
Your mistake was assuming people would be bothered to understand the details of how things work. Most people are lazy and don't know the details of how anything works.
60.
▲
by
danielmarkbruce
1mo ago
They aren't predicting the next token. It's quite literally not a prediction.
More ›