4 ms·
Apologies. Your pushback (frustration and patience) has helped me crystalize my view, thank you. 1. Define understanding. My definition isn't vague: "a compac
by Nevermark 4mo ago
Apologies. Your pushback (frustration and patience) has helped me crystalize my view, thank you.
1. Define understanding.
My definition isn't vague: "a compact representation enabled because that representation's topology closely matches the topology of the relationships being modeling."
Understanding = Scope and Suitability of Behavior / # Parameters.
Useful property: This definition applies across all scales: Scientists and mathematicians increase our understanding, every time patchworks of relationships get replaced with a simpler underlying insight.
Another useful property: It distinguishes between better understanding and having more facts. Facts improve performance but do not (non-trivially) decrease parameters.
What is your definition? In measurable terms?
2. You keep avoiding a basic aspect of modeling:
Higher compactness is achieved by higher representation correspondence between a model and the modeled.
Yes, lower level representations can work. Even well, without good "understanding". But not as compactly. And as problem complexity grows, the relative difference in parameter budgets for high-correspondence and low-correspondence representations explode.
This is not a subtle effect.
The hallmark of lower-level fitting is the far greater number of parameters required.
Dead simple example: Piece-wise linear vs. polynomial fitting of Bezier curves. Accuracy / parameter is far greater for the latter, because the representation matches the relationships being modeled.
That is an intentionally trivial example, but the same relationship holds for any problem.
You keep avoiding that.
3. Today's LLM models are very compact compared to humans.
Compressing the substance of a corpus of global human writing into less than 1% of a single human's parameter space is compact.
Humans have 100–200 trillion, some people think 500 trillion, synapses.
How do you argue that behavior scope and suitability / parameters is not remarkable, when it is remarkable compared to any specific human you could point to?
No human can converse reasonably across the scope of global communication. But these models can. For <1% of a human's parameter budget.
4. Finally, based on your clear definition, how do you argue that humans understand but models do not? Saying we are different is a copout. Defining understanding as us vs. other is both circular and unenlightening. And ignores the real progress models are clearly making relative to humans.
- Nevermark 4mo agoIs that more coherent?
- cauch 4mo ago1. Okay with your definition. My point is that you can have the same result with a representation that "closely matches the topology of the relationships being modelled". For example, a representation that "allows relationships between tokens but yet does not care about the meaning or concept not useful to form convincing sentences". And therefore, it means that you can have convincing text without needing a "representation's topology closely matching the topology of the relationships being modelled", and therefore, according to your own definition: no understanding. 2. It is not true I'm avoiding that. I have answered very clearly. 1) GenAI are not trained to get the higher representation of the world, but to get the best convincing sentence generation. This does not require a full world understanding. Worse, once a convincing sentence generation is reached, there is no gain by getting a better world understanding: the training mechanism that pushes into the correct direction stops and therefore it can go into any direction at all. 2) High compactness does not equal best solution. Even humans don't used "high compactness" when doing basic arithmetic, but use "by heart multiplication table". Being compact is useless if it comes with high complexity each time you need to recompute the output. 3) Very very good approximation can reach higher compactness anyway. Your Bezier curves is a good example: real physical phenomenons are almost never the result of a Bezier equation. A Bezier curve did not understood the phenomenon. When it comes to GenAI, it can "fit" the reality with very close precision with several representations, but the majority of the representation corresponds to an incorrect "understanding" of the reality. Another example: if I throw a ball in the air, the motion will be at first order a quadratic equation, plus correction due to friction, wind, ... If I just "train" something for "throw a ball", this system may fit a quadratic function plus corrections, but they will achieve the same result with Bezier curves, or Fourier series, or additive Gaussian, or ... But the "understanding" is that the ball is influenced by gravity, which leads to a quadratic equation. The system does not understand that. It has no reason to understand that. And it has no reason to prefer a quadratic equation fit rather than a Bezier fit, on the contrary, the Bezier fit will be more realistic (as the quadratic equation is just the first order approximation). If you want to understand a paper plane trajectory, it is a complex system, and you probably need plenty of parameters to describe the gravity, the wind at each position and each time, the shape of the plane at each time, ... But you can describe the trajectory with just few parameters using a Bezier curve. Train on plenty of paper plane trajectories, and you will have a system that can give you a very realistic paper plane trajectory based on Bezier curve. And yet, your system has no understanding of the paper plane trajectory: it does not know what are the mechanisms that make the paper plane goes up or down. It just creates a realistic trajectory without knowing why this trajectory is realistic, just that this trajectory makes sense based on the other trajectories it has seen. 3. This argument seems to go against your thesis. You are saying that humans, who "understand" + are not even able to have as much conversation as LLM, have way too much neurons. What are these neurons even for then? You are explaining that LLM are just "something different", a reduced mini-version of a brain, and yet you are also saying that they are able to do the complex things the brain do. Another way of seeing it, is that LLM are "dropping" things that they don't need to create convincing sentences, such as "understanding the token". They just "get the Bezier curve fit of the relationship" instead of understanding the real mechanisms and concepts. It's like your Bezier curve example: a system that just creates a realistic paper plane trajectory based on "typical Bezier curve observed during training" will need way less "neurons" than a system that needs to understand the whole aerodynamism of the paper plane. 4. I argue this the same way I say that a system that describe a paper plane trajectory based on best Bezier curves did not understood the mechanism behind how a paper plane trajectory works. I am not saying "I define 'understanding' as what humans do", I am saying that creating convincing sentences does not require understanding, the same way that generating realistic paper plane trajectories does not require understand gravity, Navier-Stokes equations and Brownian motions. The Bezier curve paper plane trajectory predictor system I have mention, do you think it has understanding of gravity? of Navier-Sotkes? of Brownian motions? No, it has not. You can open this system. It just has Bezier curve for plenty of examples, and thanks to that, it knows that one trajectory is realistic and another is unrealistic. And at some point, it is also able to give realistic trajectories in brand new situations it has never trained on.