5 ms·
> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sente
by Borealid 6mo ago
> No refusal fires, no warning appears — the probability just moves
I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text.
"The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds nothing that couldn't be better said by "the model neither refuses nor equivocates".
I wish I better understood how ingesting and averaging large amounts of text produced such a success in building syntactically-valid clauses and such a failure in building semantically-sensible ones. These LLM sentences are junk food, high in caloric word count and devoid of the nutrition of meaning.
- dvt 6mo ago> I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. People keep saying this, but it's quite literally fancy autocorrect. LLMs traverse optimized paths along multi-dimensional manifolds and trick our wrinkly grey matter into thinking we're being talked to. Super powerful and very fun to work with, but assuming a ghost in the shell would be illusory.
- Borealid 6mo agoIf all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. But we don't appear to have entirely done that yet. It's just curious to me that the linguistic structure is there while the "intelligence", as you call it, is not.
- staticassertion 6mo agoSentences only have semantic meaning because you have experiences that they map to. The LLM isn't training on the experiences, just the characters. At least, that seems about right to me.
- pixl97 6mo agoWhat does an experience map to?
- staticassertion 6mo agoThere are things that happen in the world that are external to us. We observe those things, and that observation is what I'm calling an experience. We can say things about the experience, but those words are not the experience. As to what the experience maps to, I think the simplest answer is that our phenomenal experiences are encoded as structures in our brain, but that's not necessary to understanding the difference between words that describe experiences and experiences themselves.
- pixl97 6mo agoOk, what kind of information structure is that experience encoded in? This is where it's really easy to start thinking the brain is some kind of interesting magic rather than encoded information.
- staticassertion 6mo agoI'm not sure I understand your question but I'll try to answer as best I can - also keep in mind that this is simply one view. The structure of the brain encodes information based on experience in the same way that the force of gravity encodes information on two rocks that collide, or other physical forces encode information into chemical structures, etc. In the case of the brain the encoding is such that various functions "fall out of" it, like being able to relate experiences, etc. There's no magic proposed here, this is a physicalist functionalist view. Nothing about this prevents a computer from being sentient. As I said, none of this even matters. The key premise is that LLMs are trained on language and not experiences. Unless you believe that a description of an experience is identical to the phenomenal experience, then we agree on the key premise. Do you think that they are identical?
- dvt 6mo ago> If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. Not necessarily. You can check this yourself by building a very simple Markov Chain. You can then use the weights generated by feeding it Moby Dick or whatever, and this gap will be way more obvious. Generated sentences will be "grammatically" correct, but semantically often very wrong. Clearly LLMs are way more sophisticated than a home-made Markov Chain, but I think it's helpful to see the probabilities kind of "leak through."
- WarmWash 6mo agoBut there is a very good chance that is what intelligence is. Nobody knows what they are saying either, the brain is just (some form) of a neural net that produces output which we claim as our own. In fact most people go their entire life without noticing this. The words I am typing right now are just as mysterious to me as the words that pop on screen when an LLM is outputting. I feel confident enough to disregard duelists (people who believe in brain magic), that it only leaves a neural net architecture as the explanation for intelligence, and the only two tools that that neural net can have is deterministic and random processes. The same ingredients that all software/hardware has to work with.
- dvt 6mo ago> I feel confident enough to disregard duelists I'm a dualist, but I promise no to duel you :) We might just have some elementary disagreements, then. I feel like I'm pretty confident in my position, but I do know most philosophers generally aren't dualists (though there's been a resurgence since Chalmers). > the brain is just (some form) of a neural net that produces output We have no idea how our brain functions, so I think claiming it's "like X" or "like Y" is reaching.
- WarmWash 6mo agoAgain, unless you are a dualist, we can put comfortable bounds on what the brain is. We know it's made from neurons linked together. We know it uses mediators and signals. We know it converts inputs to outputs. We know it can only be using deterministic and random processes. We don't know the architecture or algorithms, but we know it abides by physics and through that know it also abides by computational theory.
- codebje 6mo agoWhy would that be curious? The network is trained on the linguistic structure, not the "intelligence." It's a difficult thing to produce a body of text that conveys a particular meaning, even for simple concepts, especially if you're seeking brevity. The editing process is not in the training set, so we're hoping to replicate it simply by looking at the final output. How effectively do you suppose model training differentiates between low quality verbiage and high quality prose? I think that itself would be a fascinatingly hard problem that, if we could train a machine to do, would deliver plenty of value simply as a classifier.
- thrownthatway 6mo agoI’m not up with what all the training data is exactly. If it contains the entire corpus of recorded human knowledge… And most of everything is shit…
- joquarky 6mo agohttps://en.wikipedia.org/wiki/Sturgeon%27s_law https://en.wikipedia.org/wiki/Sturgeon%27s_law
- CamperBob2 6mo agoBecause AI is not intelligent, it doesn't "know" what it previously output even a token ago. You have no idea what you're talking about. I mean, literally no idea, if you truly believe that.
- codebje 6mo agoThat's only true if you consider the process the LLM is undergoing to be a faithful replica of the processes in the brain, right?
- CamperBob2 6mo agoNo.
- Tossrock 6mo ago> Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. Of course it knows what it output a token ago, that's the whole point of attention and the whole basis of the quadratic curse.
- dvt 6mo ago> Of course it knows what it output a token ago... It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.
- Borealid 6mo agoI don't know anything (or even much) about how our brains function, but the idea of a neuron sending an electrical output when the sum of the strengths of its inputs exceeds some value seems to be me like "a bunch of weights" getting repeatedly updated by stimulus. To you it might be obvious our brains are different from a network of weights being reconfigured as new information comes in; to me it's not so clear how they differ. And I do not feel I know the meaning of the word "know" clearly enough to establish whether something that can emit fluent text about a topic is somehow excluded from "knowing" about it through its means of construction.
- 8note 6mo agoi dont think this is a meaningful distinction. it knows the past tokens because theyre part of the input for predicting the next token. its part of the model architecture that it knows it. if that isnt knowing, people dont know how to walk, only how to move limbs, and not even that, just a bunch of neurons firing
- Jensson 6mo agoIt doesn't know if it produced that token itself or if someone else did.
- gopher_space 6mo ago
- kybernetikos 6mo agoNeural networks are universal approximators. The function being approximated in an LLM is the mental process required to write like a human. Thinking of it as an averaging devoid of meaning is not really correct.
- Borealid 6mo agoI don't think of it as "devoid of meaning". It's just curious to me that minimizing a loss function somehow results in sentences that look right but still... aren't. Like the one I quoted.
- kybernetikos 6mo agoA human in school might try to minimise the difference between their grades and the best possible grades. If they're a poor student they might start using more advanced vocabulary, sometimes with an inadequate grasp of when it is appropriate. Because the training process of LLMs is so thoroughly mathematicalised, it feels very different from the world of humans, but in many ways it's just a model of the same kinds of things we're used to.
- deleted 6mo ago[deleted]
- Terr_ 6mo ago> The function being approximated in an LLM is the mental process required to write like a human. Quibble: That can be read as "it's approximating the process humans use to make data", which I think is a bit reaching compared to "it's approximating the data humans emit... using its own process which might turn out to be extremely alien."
- TeMPOraL 6mo agoGood point. Then again, whatever process we're using, evolution found it in the solution space, using even more constrained search than we did, in that every intermediary step had to be non-negative on the margin in terms of organism survival. Yet find it did, so one has to wonder: if it was so easy for a blind, greedy optimizer to random-walk into human intelligence, perhaps there are attractors in this solution space. If that's the case, then LLMs may be approximating more than merely outcomes - perhaps the process, too.
- WarmWash 6mo agoSurely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.
- throwanem 6mo ago[flagged]
- Schiendelman 6mo agoI doubt you've ever thrown a drink in anyone's face, and I hope I'm right. This kind of thing isn't appropriate for HN.
- throwanem 6mo agoOh, good grief. Flag my comment, then. Per the HN guidelines that is the preferable action: > Don't feed egregious comments by replying; flag them instead. If you flag, please don't also comment that you did. Of course I disagree with "egregious," did it need saying. After an insult like that, I promise you, no one in my bar would consider I had acted egregiously at all. But I admit it is a surprise to see you violate the site's discussion guidelines, in the very effort to enforce them.
- nandomrumber 6mo ago> After an insult like that Did I miss something?
- throwanem 6mo ago> "real" AI being much more like a charismatic extroverted socialite As I said in my opening clause here, I fit that description exactly, and "'real' AI," as my original interlocutor would have it, sounds nothing like me. The insult arises from the fact that "'real' AI" sounds nothing particularly like anyone, because it isn't any one: if it had eyes there would be nothing happening behind them. This is why it keeps driving people insane: there are cognitive vulnerabilities here which, for most humans, have until a couple of years ago been about as realistic to need to worry about as a literal alien invasion. To a human, being compared with something which can only pretend to humanity - and that not at all well! - is an insult. It should be an insult, too. Anyone is welcome to try and fail to convince me otherwise.
- Natsu 6mo ago> I wish I better understood how ingesting and averaging large amounts of text produced such a success in building syntactically-valid clauses and such a failure in building semantically-sensible ones. These LLM sentences are junk food, high in caloric word count and devoid of the nutrition of meaning. I suspect that's because human language is selected for meaningful phrases due to being part of a process that's related to predicting future states of the world. Though it might be interesting to compare domains of thought with less precision to those like engineering where making accurate predictions is necessary.
- Jblx2 6mo ago>I wish I better understood how ingesting and averaging large amounts of text produced such a success in building syntactically-valid clauses I wonder if these LLMs are succumbing to the precocious teacher's pet syndrome, where a student gets rewarded for using big words and certain styles that they think will get better grades (rather than working on trying to convey ideas better, etc).
- coppsilgold 6mo agoThis is more or less what happens. These models are tuned with reinforcement learning from human feedback (RLHF). Humans give them feedback that this type of language is good. The notorious "it's not X, it's Y" pattern is somewhat rare from actual humans, but it's catnip for the humans providing the feedback.
- hexaga 6mo agoIt's really simple. RL on human evaluators selects for this kind of 'rhetorical structure with nonsensical content'. Train on a thousand tasks with a thousand human evaluators and you have trained a thousand times on 'affect a human' and only once on any given task. By necessity, you will get outputs that make lots of sense in the space of general patterns that affect people, but don't in the object level reality of what's actually being said. The model has been trained 1000x more on the former. Put another way: the framing is hyper-sensical while the content is gibberish. This is a very reliable tell for AI generated content (well, highly RL'd content, anyway).
- coppsilgold 6mo ago<https://en.wikipedia.org/wiki/Supernormal_stimulus https://en.wikipedia.org/wiki/Supernormal_stimulus>