Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
danielmarkbruce
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
61.
▲
by
danielmarkbruce
1mo ago
All modern LLMs that actually get used go through post-training. The finished product is something which has been through post training. So they are not next token prediction machines.
62.
▲
by
danielmarkbruce
1mo ago
No, it's not philsophical. Because if you optimize to predict, you are doing something different to optimizing for a reward. It's a different process - different objective function, different optimization, different set up.
63.
▲
by
danielmarkbruce
1mo ago
Read through the article and comments. You are talking solely about pre-training. I'm talking about post training. Respectfully, you are miles out of your depth. GPT-2 didn't use any reinforcement learning and is often given as a
64.
▲
by
danielmarkbruce
1mo ago
Nope. This isn't right.
65.
▲
by
danielmarkbruce
1mo ago
Yup, you are mostly right. I guess the people in my camp find the "it's just a next token predictor" stupid in that it's like saying "it's just a bunch of carbon and hydrogen", but it's also one of th
66.
▲
by
danielmarkbruce
1mo ago
Lol, sure, just read a blog post and you'll understand how a car works....It's very simple....
67.
▲
by
danielmarkbruce
1mo ago
Predict implies you don't control a situation. That's the difference.
68.
▲
by
danielmarkbruce
1mo ago
People have and are trying things. Lots and lots of things. They just don't go around promoting failed ideas.
69.
▲
by
danielmarkbruce
1mo ago
The fight is about the predictor language in some cases. Because it's only a trivial difference to those who don't understand the details of how these things are made. In pre-training the model really is trained to predict the nex
70.
▲
by
danielmarkbruce
1mo ago
It's not a prediction of the next move though, and that is the point. It's a prediction of what will happen if you make that move. So, it's not a next move predictor. It's a game result predictor.
71.
▲
by
danielmarkbruce
1mo ago
Emitting and predicting are different things though. Prediction implies there is some "truth" or event or something that you can test against. Prediction implies the model just learns from existing text, and optimizes to predict t
72.
▲
by
danielmarkbruce
1mo ago
If you are going to say "literally", then what is your literal definition for the word "prediction" ?
73.
▲
by
danielmarkbruce
1mo ago
If you haven't built one, and don't understand how they work, why comment?
74.
▲
by
danielmarkbruce
1mo ago
Respectfully, go build one, including doing RLHF and RLVR. Those phases generate lots of tokens, then get scored on the entirety of the output, then optimize based on a scoring of that output. It doesn't check a "prediction"
75.
▲
by
danielmarkbruce
1mo ago
Nope, it doesn't. No logic required, you can just build an LLM yourself, including post training. You'll see that predicting the next token isn't something the model does or is optimized for in RLHF or RLVR. You can hand wave
76.
▲
by
danielmarkbruce
1mo ago
Post train a model, you'll be able to determine it is not.
77.
▲
by
danielmarkbruce
1mo ago
There isn't a truth to test against. If I predict the next word in a sequence is "sat", we can check against the sequence. If I predict the roll of a die will be 4, we can check against it. Whether i give 100% or give a proba
78.
▲
by
danielmarkbruce
1mo ago
There is no truth for RLHF or RLVR. You can't predict against something if you can't check against the truth. It's not pedantry. The objective function changes. The optimization changes. THese are real things when training a
79.
▲
by
danielmarkbruce
1mo ago
Nope. Post training means even the raw model isn't predicting.
80.
▲
by
danielmarkbruce
1mo ago
The word "predict" has a meaning. I don't "predict" my next move in chess. I might predict what someone elses first move is.
81.
▲
by
danielmarkbruce
1mo ago
yes, it's exactly this. And it's not a trivial distinction.
82.
▲
by
danielmarkbruce
1mo ago
The biggest problem is the word "predictor". Once you get into post training with RLHF and RLVR, it simply isn't doing that. It is not predicting anything. It's producing tokens, but it isn't predicting them. The ch
83.
▲
by
danielmarkbruce
1mo ago
yes, it is. It's incorrect, misleading and douchey. If they didn't say "revenue tripled" and you asked people to speculate how much Uber had grown in 5 years given the statement, you'd get a lot of guesses around 10
84.
▲
by
danielmarkbruce
1mo ago
If you build an application which uses AI, you have many providers and models rigged up for various different parts of the application, and various fallback mechanisms. When one model is down, you route traffic to another model which is sim
85.
▲
by
danielmarkbruce
1mo ago
If you spend any time trying to find such companies you'll see that for every invisible profitable company you'll find 20 invisible unprofitable disaster companies. Search funds have existed for a long time now, and it's hard
86.
▲
by
danielmarkbruce
1mo ago
touche.
87.
▲
by
danielmarkbruce
1mo ago
Even if all the big ideas are gone and we are entering a new part of the curve, there is still an enormous amount of improvement possible. Just iterating on data mix/quality etc, training pipelines, reward functions, specific ways of r
88.
▲
by
danielmarkbruce
1mo ago
Sure, but the use of "decimated" is used approximately correctly in most cases, even if not precisely right. And it is a cool sounding word, and it's one word. "Orders of magnitude" is three words when one would do,
89.
▲
by
danielmarkbruce
1mo ago
You can't give a concrete and nonsensical way Uber has grown 100x (minimum number of "orders") while revenue tripled. Cannot be miles. Can't be riders. Can't be drivers. Revenue is top line, by definition. This stat
90.
▲
by
danielmarkbruce
1mo ago
"Over the last 5+ years, Uber has grown by orders of magnitude, with our top line nearly tripling" Orders of magnitude.
More ›