4 ms·
ML is one of the easiest fields out there. When I learned it I was actually turned off by how simplistic the concept was. Of course let me preface to say that i
by corethree 3y ago
ML is one of the easiest fields out there. When I learned it I was actually turned off by how simplistic the concept was. Of course let me preface to say that it's hard to develop the intuition and skill in the same way learning to skateboard is hard. But conceptually it's easy and very possible for almost anyone.
The whole thing is just curve fitting. Literally finding some best fit curve across a series of points. This is very very easy for any software engineer to understand. I literally lost interest when I found out that the entire field was just all about messing with the data and the curve to try to get things to fit.
Literally it's just about eyeballing the data and qualitatively picking and training the thing that looks like it's the best fit. But because the data is N-dimensional and in the millions it's impossible to "eye-ball" it with your physical eyes, you have to come up with other techniques equivalent to "eye-balling" it.
Douglas Hofstadter had this whole theory of consciousness and when he found out that an LLM was a simple feed forward network with no feedback loops he went into a crisis. Basically his whole theory in GEB was wrong, according to him.
This stuff is NOT quantum physics. It's startling how simple it is and that's one of the big mysteries about it.
We only understand and build these things at a high level. At the very low level we don't actually understand what's going on. As I stated earlier we understand ML the same way a person understand data from an "eye-ball" perspective so it's impossible to even justify what exactly specifically went on with chatGPT when he answered a specific question correctly.
- rokkitmensch 3y agoCan I get a link to anything about this crisis? My evening popcorn lulzsession demands...
- corethree 3y agohttps://www.nytimes.com/2023/07/13/opinion/ai-chatgpt-consciousness-hofstadter.html https://www.nytimes.com/2023/07/13/opinion/ai-chatgpt-consci... Did you read his book Godel Escher Bach? It's a good book. CS people love it as it's all about recursion and puzzles.
- moandcompany 3y ago"It is difficult to get a man to understand something, when his salary depends on his not understanding it." (Upton Sinclair)
- graphe 3y agohttps://en.wikipedia.org/wiki/Moravec%27s_paradox https://en.wikipedia.org/wiki/Moravec%27s_paradox Moravec's paradox is the observation in artificial intelligence and robotics that, contrary to traditional assumptions, reasoning requires very little computation, but sensorimotor and perception skills require enormous computational resources. The principle was articulated by Hans Moravec, Rodney Brooks, Marvin Minsky and others in the 1980s. Moravec wrote in 1988, "it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility".
- anon291 3y ago> Moravec wrote in 1988, "it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility". The resolution to the paradox is so simple I must be missing something. The amount of data in datasets for 'mobility' is basically zero. You would have to manually construct such a dataset. Whereas, humans have for thousands of years been trained to symbolically encode their reasoning processes in a way that has been incredibly accessible to computers (prose). If I understand correctly, the scaling laws for mobility are the same for language and reasoning. We need more data.
- graphe 3y agoThe training data for chess is very easy in comparison to walking. If you had the same amounts of data for both and the ability to get it, understand it and use it you wouldn't have a problem. Basically it's hard to make a machine use and understand how to use it's physical form and in 1988 it was even harder. For chess it was easy. It's easy to get, understand and use chess data.
- beoberha 3y agoI probably shouldn’t take the bait here, but this reads like someone took intro to ML and thinks that’s all there is to know. “Just” fitting a curve couldn’t be more reductionist and discount the work of a ton of incredibly intelligent people. I tend to agree that I don’t really find ML work all that interesting (much more interested in making it go fast :)), but simple it is not.
- corethree 3y agoNo bait. Apologies if it sounded insulting. Put it this way, it's extremely challenging and not simple at all to walk and balance on a wire. Tight rope walking is not simple at all because very few people can do it. But tightrope walking is different from something like Quantum physics. I may not be able to tightrope walk but I can understand the concept in it's entirety. For Quantum Physics, many people will never truly understand it. What I'm saying is this, ML is tightrope walking. Challenging, not simple, but NOT quantum physics. It only seems like quantum physics.
- calebh 3y agoWhat are ML jobs about? I have this vague notion that you spend a lot of time gathering/cleaning data and throwing things at the wall, but maybe that's not accurate. I've always been stronger at discrete type math/programming, which is why I tend to shy away from statistics-based stuff like ML. One thing to note is that LLMs are indeed feed forward, however the generation of the text (from my understanding) is recursive in that you feed each output token to another forward pass of the neural network.
- anon291 3y ago> I've always been stronger at discrete type math/programming, which is why I tend to shy away from statistics-based stuff like ML. I think there's a major misconception that ML in the form of deep learning is about statistics. There's no statistics in deep learning models. There are some statistical measurements made of final models, much in the same way a good computer science paper covering implementations of discrete data structures might make statistical statements showing the performance of the author's implementation, but like transformers and traditional neural nets and backprop have nothing to do with statistics.
- golly_ned 3y agoWhat parts of the ML pipeline have you focused your work on?
- anon291 3y ago> Douglas Hofstadter had this whole theory of consciousness and when he found out that an LLM was a simple feed forward network with no feedback loops he went into a crisis. Basically his whole theory in GEB was wrong, according to him. LLMs are not conscious. The training process for an LLM is not a feed forward network. If we were going to try to fit the idea of consciousness a la humanity (which is really the only fully 'conscious' creature we know of) into LLMs, then 'running' an LLM is identical to cloning a frozen human, thawing it, firing some neurons, reading the result and then destroying the clone. A better argument for actual consciousness would come from the training process, but that itself is also dubious. It's unlikely consciousness is an emergent phenomenon. Or rather, such a claim is extraordinary and would require an extraordinary amount of proof, which GEB does not provide, sorry.
- corethree 3y agoDouglas agrees with you. He thinks his own book, GEB is wrong. Essentially chatGPT basically made him do a 180. He sees his life work in shambles. See here: https://www.nytimes.com/2023/07/13/opinion/ai-chatgpt-consciousness-hofstadter.html https://www.nytimes.com/2023/07/13/opinion/ai-chatgpt-consci...
- zozbot234 3y agoLLM's are all about "let's think this through step by step". That's literally a loop in the programming/software engineering sense, it's just expressed via natural language.
- wait_a_minute 3y ago> The whole thing is just curve fitting. > We only understand and build these things at a high level. At the very low level we don't actually understand what's going on. Contradictory..
- corethree 3y agoCurve fitting a set of data is essentially the same thing as understanding something at a high level. We have data points but no way to extract the low level exact equation that generated that data... So we create a curve and estimate it. We will never know the true equation. Additionally the curve has hundreds of dimensions and is essentially something that can't be visualized or understood cohesively. We have this neural network that represents the curve but the neural network is a black box.
- deleted 3y ago[deleted]