5 ms·
While reading the article I enjoyed pretending that a powerful LLM just crash landed on our planet and researchers at Anthropic are now investigating this fasci
by indigoabstract 2y ago
While reading the article I enjoyed pretending that a powerful LLM just crash landed on our planet and researchers at Anthropic are now investigating this fascinating piece of alien technology and writing about their discoveries. It's a black box, nobody knows how its inhuman brain works, but with each step, we're finding out more and more.
It seems like quite a paradox to build something but to not know how it actually works and yet it works. This doesn't seem to happen very often in classical programming, does it?
- BoredPositron 2y agoThe bigger problem is that nobody knows how a human brain works that’s the real crux with the analogy.
- richardatlarge 2y agoI would say that nobody agrees, not that nobody knows. And it’s reductionist to think that the brain works one way. Different cultures produce different brains, possible because of the utter plasticity of the learning nodes. Chess has a few rules, maybe the brain has just a few as well. How else can the same brain of 50k years ago still function today? I think we do understand the learning part of the brain, but we don’t like the image it casts, so we reject it
- wat10000 2y agoThat gets down to what it means to “know” something. Nobody agrees because there isn’t enough information available. Some people might have the right idea by luck, but do you really know something if you don’t have a solid basis for your belief but it happens to be correct?
- richardatlarge 2y agoPotentially true, but I don’t think so. I believe it is understood and unless you’re familiar with every neuro/behavioral literature, you can’t know. Science paradigms are driven by many factors and being powerfully correct does not necessarily rank high when the paradigms implications are unpopular
- absolutelastone 2y agoWell there are some people who think they know. I personally agree with the above poster that such people are probably wrong.
- deleted 2y ago[deleted]
- cma256 2y agoIn my experience, that's how most code is written... /s
- jfarlow 2y ago>to build something but to not know how it actually works and yet it works. Welcome to Biology!
- oniony 2y agoAt least, now, we know what it means to be a god.
- School-Cotton 2y ago> This doesn't seem to happen very often in classical programming, does it? Not really, no. The only counterexample I can think of is chess programs (before they started using ML/AI themselves), where the search tree was so deep that it was generally impossible to explain "why" a program made a given move, even though every part of it had been programmed conventionally by hand. But I don't think it's particularly unusual for technology in general. Humans could make fires for thousands of years before we could explain how they work.
- deleted 2y ago[deleted]
- woah 2y ago> It seems like quite a paradox to build something but to not know how it actually works and yet it works. This doesn't seem to happen very often in classical programming, does it? I have worked on many large codebases where this has happened
- worldsayshi 2y agoI wonder if in the future we will rely less or more on technology that we don't understand. Large code bases will be inherited by people who will only understand parts of it (and large parts probably "just works") unless things eventually get replaced or rediscovered. Things will increasingly be written by AI which can produce lots of code in little time. Will it find simpler solutions or continue building on existing things? And finally, our ability to analyse and explain the technology we have will also increase.
- Sharlin 2y agoSee: Vinge’s “programmer-archeologists” in A Deepness in the Sky. https://en.m.wikipedia.org/wiki/Software_archaeology https://en.m.wikipedia.org/wiki/Software_archaeology
- bob1029 2y agoI think this is a weird case where we know precisely how something works, but we can't explain why.
- k__ 2y agoI've seen things you wouldn't believe. Infinite loops spiraling out of control in bloated DOM parsers. I’ve watched mutexes rage across the Linux kernel, spawned by hands that no longer fathom their own design. I’ve stared into SAP’s tangled web of modules, a monument to minds that built what they cannot comprehend. All those lines of code… lost to us now, like tears in the rain.
- baq 2y agoDo LLMs dream of electric sheep while matmuling the context window?
- timschmidt 2y agoHow else would you describe endless counting before sleep(); ?
- FeepingCreature 2y agowhile (!condition && tick() - start < 30) __idle(); // Baaa.
- indigoabstract 2y agoHmm, better start preparing those Voight-Kampff tests while there is still time.
- qingcharles 2y agoI can't understand my own code a week after writing it if I forget to comment it.
- resource0x 2y agoIn technology in general, this is a typical state of affairs. No one knows how electric current works, which doesn't stop anyone from using electric devices. In programming... it depends. You can run some simulation of a complex system no one understands (like the ecosystem, financial system) and get something interesting. Sometimes it agrees with reality, sometimes it doesn't. :-)
- Vox_Leone 2y ago>>It seems like quite a paradox to build something but to not know how it actually works and yet it works. This doesn't seem to happen very often in classical programming, does it? Well, it is meant to be "unknowable" -- and all the people involved are certainly aware of that -- since it is known that one is dealing with the *emergent behavior* computing 'paradigm', where complex behaviors arise from simple interactions among components [data], often in nonlinear or unpredictable ways. In these systems, the behavior of the whole system cannot always be predicted from the behavior of individual parts, as opposed to the Traditional Approach, based on well-defined algorithms and deterministic steps. I think the Anthropic piece is illustrating it for the sake of the general discussion.
- indigoabstract 2y agoCorrect me if I'm wrong, but my feeling is this all started with the GPUs and the fact that unlike on a CPU, you can't really step by step debug the process by which a pixel acquires its final value (and there are millions of them). The best you can do is reason about it and tweak some colors in the shader to see how the changes reflect on screen. It's still quite manageable though, since the steps involved are usually not that overwhelmingly many or complex. But I guess it all went downhill from there with the advent of AI since the magnitude of data and the steps involved there make traditional/step by step debugging impractical. Yet somehow people still seem to 'wing it' until it works.
- IngoBlechschmid 2y ago> It seems like quite a paradox to build something but to not know how it actually works and yet it works. This doesn't seem to happen very often in classical programming, does it? I agree. Here is a remote example where it exceptionally does, but it is mostly practically irrelevant: In mathematics, we distinguish between "constructive" and "nonconstructive" proofs. Intertwined with logical arguments, constructive proofs contain an algorithm for witnessing the claim. Nonconstructive proofs do not. Nonconstructive proofs instead merely establish that it is impossible for the claim to be false. For instance, the following proof of the claim that beyond every number n, there is a prime number, is constructive: "Let n be an arbitrary number. Form the number 1*2*...*n + 1. Like every number greater than 1, this number has at least one prime factor. This factor is necessarily a prime numbers larger than n." In contrast, nonconstructive proofs may contain case distinctions which we cannot decide by an algorithm, like "either set X is infinite, in which case foo, or it is not, in which case bar". Hence such proofs do not contain descriptions of algorithms. So far so good. Amazingly, there are techniques which can sometimes constructivize given nonconstructive proofs, even though the intermediate steps of the given nonconstructive proofs are simply out of reach of finitary algorithms. In my research, it happened several times that using these techniques, I obtained an algorithm which worked; and for which I had a proof that it worked; but whose workings I was not able to decipher for an extended amount of time. Crazy! (For references, see notes at rt.quasicoherent.io for a relevant master's course in mathematics/computer science.)
- deleted 2y ago[deleted]
- gwd 2y ago> It seems like quite a paradox to build something but to not know how it actually works and yet it works. That's because of the "magic" of gradient descent. You fill your neural network with completely random weights. But because of the way you've defined the math, you can tell how each individual weight will affect the value output at the other end; and specifically, you an take the derivative. So when the output is "wrong", you say, "would increasing this weight or decreasing have gotten me closer to the correct answer"? If increasing the node would have gotten you closer, you increase it a bit; if decreasing it would have gotten you closer you decrease it a bit. The result is that although we program the gradient descent algorithm, we don't directly program the actual circuits that the weights contain. Rather, the nodes "converge" into weights which end up implementing complex circuitry that was not explicitly programmed.
- gwd 2y agoIn a sense, the neural network structure is the "hardware" of the LLM; and the weights are the "software". But rather than explicitly writing a program, as we do with normal computers, we use the magic of gradient descent to summon a program from the mathematical ether. Put that way, it should be clearer why the AI doomers are so worried: if you don't know how it works, how do you know it doesn't have malign, or at least incompatible, intentions? Understanding how these "summoned" programs work is critical to trusting them; which is a major reason why Anthropic has been investing so much time in this research.
- miraculixx 2y agoWe know how they work. It's just that it works better than expected. Which of course doesn't mean we don't know, it just means there are second-order effects that are non-obvious. Intelligence is not implied.
- miraculixx 2y ago> This doesn't seem to happen very often in classical programming, does it? Try concurrent programming. It happens all the time.