6 ms·
This is humorous (and well-written), but I think its more than that. I'm always making the joke (observation) that ML (AI) is just curve-fitting. Whether "just
by elijahbenizzy 2y ago
This is humorous (and well-written), but I think its more than that.
I'm always making the joke (observation) that ML (AI) is just curve-fitting. Whether "just curve-fitting" is enough to produce something "intelligent" is, IMO, currently unanswered, largely due to differing viewpoints on the meaning of "intelligent".
In this case they're demonstrating some very clean, easy-to-understand curve-fitting, but it's really the same process -- come up with a target, optimize over a loss function, and hope that it generalizes, (this one, obviously, does not. But the elephant is cute.)
This raises the question Neumann was asking -- why have so many parameters? Ironically (or maybe just interestingly), we've done a lot with a ton of parameters recently, answering it with "well, with a lot of parameters you can do cool things".
- visarga 2y ago> Whether "just curve fitting" is enough to produce something "intelligent" is, IMO, currently unanswered Continual "curve fitting" to the real world can create intelligence. What is missing is not something inside the model. It's missing a mechanism to explore, search and expand its experience. Our current crop of LLMs ride on human experience, they have not largely participated in creating their own experiences. That's why people call it imitation learning or parroting. But once models become more agentic they can start creating useful experiences on their own. AlphaZero did it.
- soist 2y agoAlphaZero did not create any experiences. AlphaZero was software written by people to play board games and that's all it ever did.
- visarga 2y agoAZ trained in self-play mode for millions of games, over multiple generations of a player pool.
- soist 2y agoI am familiar with the literature on reinforcement learning.
- pharrington 2y agoThey're saying the board games AlphaZero played with itself are experiences.
- soist 2y agoAnd I am saying they are confused because they are attributing personal characteristics to computers and software. By spelling out what computers are doing it becomes very obvious that there is nothing that can be aware of any experiences in computers as it is all simply a sequence of arithmetic operations. If you can explain which sequence of arithmetic operations corresponds to "experiences" in computers then you might be less confused than all the people who keep claiming computers can think and feel.
- nyssos 2y ago> By spelling out what computers are doing it becomes very obvious that there is nothing that can be aware of any experiences in computers as it is all simply a sequence of arithmetic operations. By spelling out what brains are doing it becomes very obvious that it's all simply a sequence of chemical reactions - and yet here we are, having experiences. Software will never have a human experience - but neither will a chimp, or an octopus, or a Zeta-Reticulan. Mammalian neurons are not the only possible substrate for intelligence; if they're the only possible substrate for consciousness, then the fact that we're conscious is an inexplicable miracle.
- soist 2y agoThis is a common retort. You can read my other comments if you want to understand why you're not really addressing my points because I have already addressed how reductionism does not apply to living organisms but it does apply to computers.
- rocqua 2y agoAre you launching into a semantic argument about the word 'experience'? If so, it might help to state what essential properties alphago was missing that makes it 'not having an experience'. Otherwise this can quickly devolve into the common useless semantic discussion.
- soist 2y agoJust making sure no one is confused by common computationalist sophistry and how they attribute personal characteristics to computers and software. People can have and can create experiences, computers can only execute their programmed instructions.
- HeatrayEnjoyer 2y agoOn what priors are you making that statement?
- gowld 2y agoUsername `soist is an abbreviation for `solipsist, then?
- elijahbenizzy 2y agoThere are a whole bunch of assumptions here. But sure, if you view the world as a closed system, then you have a decision as a function of inputs: 1. The world around you 2. The experiences within your (really, the past view of the world around you) 3. Innateness of you (sure, this could be 2 but I think it's also something else) 4. The experience you find + the way you change yourself to impact (1), (2), and (3) If you think of intelligence as all of these, then you're making the assumption that all that's required for (2), (3), and (4) is "agentic systems", which I think skips a few steps (as the author of an agent framework myself...). All this is to say that "what makes intelligence" is largely unsolved, and nobody really knows, because we actually don't understand this ourselves.
- godelski 2y ago> Continual "curve fitting" to the real world can create intelligence. I'm going to need a citation on this bold claim. And by that I mean in the same vein as what Carl Sagan would say Extraordinary claims require extraordinary evidence
- Grimblewald 2y agoI'd simply argue what we do is precisley that, so either we are intelligent, or we are not, however we might define that intelligence.
- kgeist 2y ago>It's missing a mechanism to explore, search and expand its experience. Can't we create an agent system which can search the internet and choose what data to train itself with?
- Xcode23 2y agoyou need to define what the utility function of the agent is so it can know what to actually use to train itself. If we knew that this whole debate about human intelligence in computers would either be solved already or well on its way to being solved.
- luplex 2y agoI mean the devil is in the details. In Reinforcement Learning, the target moves! In deep learning, you often do things like early stopping to prevent too much optimization.
- soist 2y agoThere is no such thing as too much optimization. Early stopping is to prevent overfitting to the training set. It's a trick just like most advances in deep learning because the underlying mathematics is fundamentally not suited for creating intelligent agents.
- rocqua 2y agoIs over fitting different from 'too much optimization'? Optimization still needs a value that is optimized. Over fitting is the result of too much optimization for not quite the right value (i.e. training error when you want to reduce prediction error)
- soist 2y agoWhat value is being optimized and how do you know it is too much or not enough?
- godelski 2y agoI think the miscommunication is due to the proxy nature of our modeling. From one perspective, yes you're right because it's just on your optimization function and objectives. But if we're in the context where we recognize the practical usage of our model replies on it being an inexact representation (proxy) then certainly there is too much optimization. I mean most of what we try to model in ML is intractable. In fact, that entire notion of early stopping is due to this. We use a validation set as a pseudo test set to inject information into our optimization products without leaking information from the test set (why you shouldn't choose parameters based on test results. That is spoilage. Doesn't matter if it's status quo, it's spoilage) But we also need to consider that a lack of divergence between train/val does not mean there isn't overfittng. Divergence implies overfittng but the inverse statement is not true. I state this because it's both relevant here and an extremely common mistake.
- maitola 2y agoIn the case of AI, the more parameters, the better! In Physics is the opposite.
- elijahbenizzy 2y agoOne of the hardest parts of training models is avoiding overfitting, so "more parameters are better" should be more like "more parameters are better given you're using those parameters in the right way, which can get hard and complicated". Also LLMs just straight up do overfit, which makes them function as a database, but a really bad one. So while more parameters might just be better, that feels like a cop-out to the real problem. TBD what scaling issues we hit in the future.
- will1am 2y agoA dichotomy between these fields
- will1am 2y agoYour humorous observation captures a fundamental truth to some extent