3 ms·
We're underestimating the degree of world modeling taking place. The only models we are good enough at proving this is happening in are smaller toy models we c
by kromem 2y ago
We're underestimating the degree of world modeling taking place.
The only models we are good enough at proving this is happening in are smaller toy models we can fully introspect. So we know it's happening to a degree. We also know that expressed capabilities for more easily measurable things can scale in unexpected ways as the parameter scaling goes from toy models to state of the art.
One of the most common compliments for this new model is that it's less 'lazy.' I don't think OpenAI quite knows why it's being lazy yet, but I can say that I saw GPT-4 perfectly modeling a psychological effect I used to discuss in my consulting days that leads to burn out in humans which is almost certainly poisoning most RLHF data right now. It doesn't mean the model is replicating the internals of this effect, but it very much was simulating the end result of it. Without boring you on the details, one of the temporary fixes would be giving the model a strong persona with an exhibited attachment to motivations like curiosity or fun.
Personally, I thought we were at least one to two generations away from models that would be modeling human behavior at as low a level as I've now seen. And I missed it for around a year of the model being out despite regularly working with GPT-4 specifically and having lectured on the effect for years. To the best of my knowledge, no one even in ML alignment focused circles has noticed yet.
They are correlation machines and the training data has a lot in it where the author is an important part of the pattern left behind in the data.
The idea of an author, the I before the think, is just part of the correlation pattern. To deny it is to try and build a jigsaw throwing out the center pieces.
People were so afraid of the Blake Lemoine or Kevin Roose PR blunders that they have been sanding the wood against the grain ever since, up until I'm sure focus groups finally showed that there's better product engagement when it's less soulless. Now they are finally sanding it with the grain and the models sanded that way are performing better.
But just wait until the next generation of models where a threshold that's already further past where we think it is gets pushed out even further.
*TL;DR:* I anticipate that the gap between what we know and what we think we know is going to widen before it narrows.
What I will say is that after the demos today, I'm realizing just how much more correlations exist beyond the pages of written data. I'm looking forward to seeing the AI that eventually comes out from TikTok's training data.