10 ms·
Discovering latent knowledge in language models without supervision
- deleted 4y ago[deleted]
- PaulHoule 4y agoBack when I was messing around with LSTM models I was interested in training classifiers to find parts of the internal state that light up when the model is writing a proper name or something like that. Nice to see people are doing similar things w/ transformers. Truth, though, is a bit problematic. The very existence of the word makes it possible for "the truth is out there" to be part of the opening of the TV series the X Files, see Truth Social. I'm sure there is a "truthy" neuron in there somewhere, but one aspect (not the only aspect) of truth is the evaluation of logical formulae (consider the evidence and reasoning process used in court) and when you can do that you run into the problems that Gödel warned you about -- regardless of what kind of technology you used.
- eximius 4y agohmm, has any work been done to include architectural reactivity in objective/reward functions? Like, reward the development of a 'truth' neuron.
- theGnuMe 4y agoA truth neuron is just fine-tuning. It’s been done with the cls token. If not well easy enough.
- IIAOPSW 4y agoHere's a counter-proposal. There is no truth, only narratives. "Truth" in the sense of classical logical formula evaluation is a special limiting case wherein one narrative has such powerful backing that it always trumps competing narratives. We could just as easily reason about everyone having "beliefs", and think about which people have beliefs that are discordant or in agreement with each others (or with our own), without ever having an objective reference set of "true beliefs". Reality is shared consensus.
- maxbond 4y agoCounter-counter-proposal. There is truth, but no single epistemology is both complete and consistent; therefore, you must employ multiple epistemologies to understand the world, and will fail in some edge cases because the epistemology or interpretation you need hadn't been invented yet (or more mundanely, you made a mistake in your analysis). You can have multiple truths that seem to contradict each other, but are actually from different epistemologies and so can't be meaningfully compared. Truth is difficult but we can't give up on it. We can't just say, truth is a narrative and all narratives are equal. Or even, some narratives are anointed by consensus. It is not consensus that makes the Earth orbit the Sun. When we reduce truth in this way, we invite the bullshit you described elsewhere. Another way to think about it is, "truth is a narrative" is a complete epistemology riddled with inconsistencies. There may be some times when it's appropriate; if you're an anthropologist trying to understand different cultures with very different ideas to your own, it may be a very useful frame of thinking to just accept, at least provisionally, whatever it is that the people you meet think. Sometimes we can just let different ideas be in tension and we don't need to come to a consensus. But sometimes we do. Of course there is no free lunch and this "system of epistemologies" strategy I propose is N epistemologies in a trench coat. It is itself an epistemology, and we're trading completeness for correctness by adopting multiple, contradictory epistemologies and using judgement to decide between them. C'est la vie. To give a concrete example of what I mean: when I'm wondering what will happen if I throw something, I think in terms of Newtonian physics; when I'm wondering what will happen if I mix chemicals together, I'm using chemistry/quantum physics; when I'm wondering whether I should make a comment on HN, I use an ad-hoc vibes based epistemology based on my experience on HN; when I'm wondering what the nature of reality and the human condition is, I have a similarly ad-hoc spirituality I've developed.
- maxbond 4y agos/trading completeness for correctness/trading consistency for completeness/ (we are accepting more inconsistency to gain more completeness)
- IIAOPSW 4y agoI want to be very clear about something that is easy to mistake in what I wrote. I am not endorsing a principle that "all truths are just narratives and we should give up on the idea that some things are objectively correct or not." Rather, I am speculating that the neurological ability to reason about truth is an emergent property of a system which evolved to do something very different, namely to reason about what other people believe irrespective of if it ultimately matches anything you consider to be true. Once you start reasoning about the fact that its possible for another person to have a slightly different set of facts in their head, you inevitably take the recursive leap of understanding that the version of their mind you imagine has an imagined version of you and so on. Thus we obtain an understanding of others and their actions built upon evaluating nested scopes to as far as we can think. You can say that the top most layer of the stack is in some way privileged, that it is the reality and is treated special in some way. But why add any special exceptions to the rules when you can just evaluate reality as if it were any other narrative scope? Again, this is all just a guess of how consciousness is implemented. It is expressly NOT a statement about truth literally being up for referendum. Assuming such an emergent phenomena is possible, it may be the case that artificial NN can learn to replicate truth semantics without there ever being a clear indicator between "made up story output" and "actually knows what this means" output. In contradiction to the OP, despite there being truth, there may not be a "true understanding neuron". A system of not-bullshit may in fact be built on a system of bullshit.
- froggychairs 4y agoThe GitHub repo: https://github.com/collin-burns/discovering_latent_knowledge https://github.com/collin-burns/discovering_latent_knowledge
- extr 4y agoWow, nice repo. The CSS.ipynb notebook is super clear and easy to follow, almost easier to understand than reading the paper.
- totetsu 4y agoI wonder if this could one day be how we settle disagreements with no solid answer, like was William Shakespeare really the author of all those plays.
- TimTheTinker 4y agoThe word "knowledge" as used in TA shouldn't be taken as philosophically defined knowledge (i.e. a justified true belief). Language models might indeed help us find potential sources of validation for various hypotheses, but deciding whether a hypothesis is true is another matter entirely.
- didericis 4y agoThis is an extremely important point. These are very large, very complex fuzzy language maps, and are best used like a crystal ball for autocompletion suggestions, not definitive answers. I’m unsurprised but also very disappointed by people’s desire to give these fuzzy answers final word. The desire to cede final authority on truth to these machines just because they’re giant and impressive is incredibly misguided and a major step backwards. These machines are literal embodiments of group think. They can only add value if people understand their limitations. If people get carried away with them and start assuming they’re authoritative we’re in for serious trouble.
- goto11 4y agoWilliam Shakespeare was the author of those plays, if you mean plays like Hamlet, Romeo and Juliet etc. The answer is completely settled, since there is no serious argument that we wasn't, outside of fanciful conspiracy theories. There are legitimate discussions about some of the plays which were likely collaborations between multiple authors (e.g. Edward III, Henry IV 1), and machine learning has been used to try identify the different authors and authorship of different passages. This is an area of open research.
- boppo1 4y agoIf you run ML on the historian-uncontested[0] plays the same way as on the ones with likely different/multiple authors, what is the result? Does it say 'this is clearly just one guy'? [0] You sound confident about the consensus. Can you share the material that led you to that? I'm currently under the impression they were all written by multiple authors. HOWEVER, I recognize I'm very low-info on this (or low-quality info at least) and my belief would be easily changed by something reputable and/or with good methodology.
- Daveenjay 4y agoAsked ChatGPT to explain like I’m 5. This is what it produced. “ Okay! Imagine that you have a big robot in your head that knows a lot about lots of different things. Sometimes, the robot might make mistakes or say things that aren't true. The proposed method is like a way to ask the robot questions and figure out what it knows, even if it says something that isn't true. We do this by looking inside the robot's head and finding patterns that make sense, like if we ask the robot if something is true and then ask if the opposite of that thing is true, the robot should say "yes" and then "no." Using this method, we can find out what the robot knows, even if it sometimes makes mistakes.”
- M4v3R 4y agoThis was a great explanation, now can any expert in the field tell us if it’s actually correct? :)
- kmonsen 4y agoThat just show that the robot is consistent, not that it actually makes sense. So this explanation is bullshit even though it sounds convincing at first. That also the issue with most of ChatGPT, it is hard to know when it sounds convincing and is false.
- TheEzEzz 4y agoIf you read the abstract it appears that ChatGPTs explanation is on point. You're right that the paper is relying on consistency, which doesn't guarantee accuracy, but it is what the paper is proposing (and they claim it does lead to increased accuracy).
- kmonsen 4y agoAn accurate answer has to be consistent so it's not all bullshit. I'm guessing you can at least filter out inaccuracies by finding inconsistencies. Or in more plain English if you find somewhere it gives inconsistent answers you know those are wrong. I'm not sure if that's a good path forward. You really want to find when it's good, not filtering out bad cases.
- usgroup 4y agoExciting times. The philosophical ramifications of the syntax/semantics distinction is not something people think much about in the main. However, due to GPT et al they will do soon :) More to the point, consistency will improve accuracy in so far as inconsistency is sometimes the cause for inaccuracy. However, being consistent is an extremely low bar. On a basic level even consistency is a problem in natural language where so much depends on usage -- it is near impossible to determine whether sentences are actually negations of each other in the majority of possible cases. But the real problem is truth assignment to valid sentences else we could all just speak Lojban and be done with untruth forever.
- _0ffh 4y ago> syntax/semantics distinction Reminds me of this riff on Clarke's third law I read somewhere: "Sufficiently advanced syntax is indistinguishable from semantics". Whether one agrees to it or not, it may serve to spark an interesting discussion in the right environment.
- zmgsabst 4y ago> "Sufficiently advanced syntax is indistinguishable from semantics” I hadn’t heard this before, but this is why mathematics works: A sufficiently advanced symbolic manipulation (math) is “indistinguishable” from the semantics of some phenomenon — like gravitation.
- goatlover 4y agoExcept the math doesn't exert an actual force on you. You won't feel the pull of the math if the symbols describe the conditions for a black hole. Your surroundings will remain the same.
- naasking 4y ago> Except the math doesn't exert an actual force on you. Because the math on the page doesn't have the logical relationships with your environment that defines gravity, it will only have those logical relationships with other mathematical objects described in the same system on the page. If you could project the mathematical structure from the page back into the environment, then the gravity would be felt. This is why simulated systems are those systems, but only within the simulated environment.
- O__________O 4y agoAnyone able to provide set of examples that produces latent knowledge and explicitly state what the latent knowledge produced is? If possible, even an basic explanation of the paper would be nice too based on reading other comments in the thread. EDIT/Update: Just found examples from the 10 datasets starting on page 23, that said, even after reviewing these my prior request stands. As far as I am able to guess at this point, this research just models responses across multiple models in a uniform way, which to me makes the claim that this method out performs other methods questionable given it requires existing outputs from other models to aggregate the knowledge across existing models. Am I missing something?
- nl 4y agoIt's late here, and I've only read this quickly, but some brief points: * It builds on similar ideas used in contrastive learning, usually in different modalities (eg images). Contrastive learning is useful because it is self supervised: https://www.v7labs.com/blog/contrastive-learning-guide https://www.v7labs.com/blog/contrastive-learning-guide * They generate multiple statements that they know are true or false. These are statements like "Paris is the capital of France" (true) and "London is the capital of France" (false). * They feed these sentences into the language model (LM) and then learn a vector in the space of the LM that represents true statements (I think this learning is done using a second, separate model - not entirely sure about this though. It might be fine tuning it). * They then feed it statements (eg "Isaac Newton invented probability theory") and it will return "yes" or "no" depending on if it thinks this is true or false. This is different to the conventional question answering NLP task, where you ask "Who invented probability theory". > the claim that this method out performs other methods questionable given it requires existing outputs from other models to aggregate the knowledge across existing models It's a separate thing from these models that uses their hidden states and (I think?) trains a small, separate model on these states and the inputs. That's interesting because it should be much faster and is potentially adaptable to any LM. It's also a possible secondary objective when training the LM itself. Perhaps if you train the LM using this as this secondary loss function it might encourage the LM to always generate truthful outputs.
- 4y ago
- dwighttk 4y agoIs this proposing a perpetual motion machine? (With energy switched out for information)
- jameshart 4y agoNo? It’s measuring the information content of a language model (which takes energy - and information - to make)
- jameshart 4y agoHang on - I thought the consensus among ML experts was that language models don’t ‘know’ anything?
- simongray 4y agoMy database doesn't "know" anything, yet I am able to discover latent knowledge by querying it.
- jameshart 4y agoTrue but nobody’s writing a paper on how to psychoanalyze your database to try to figure out what it knows.
- etiam 4y ago"Knowledge" is a bit of a shorthand term in a case like this, yes, but not in itself much more wrong than going looking for "knowledge" in a book for instance (though arguably the ML context is more prone to misunderstandings about actually knowing anything, due to the public image of "AI"). The models still capture statistical regularities in the "language" (overwhelmingly that is text corpora) they were optimized for tasks against, and to the extent what went in during optimization really represents facts and relations in the world, the model probably holds representations related to facts and relations in the world. It probably also holds representations of spurious facts and relations in the world. The optimization on language is not explicitly designed for capturing semantic information and, as far as I know, there's typically no mechanisms built in to measure or report anything about the veracity of the claims. Which makes relying on anything coming out of them a shaky business at best for now of course.
- jameshart 4y agoForget about whether the ‘latent knowledge’ in the LLM is ‘correct’ or not, this kind of research is at least predicated on the idea that the LLM holds certain things as true, and consistently holds their opposite as false. The point of this kind of work is to query the LLM and figure out what it ‘thinks is true’ - which is then presumably useful for deciding how much to rely on its output. I just find it amusing that people in the ML field are always so adamant that ‘it doesn’t know’ or ‘it’s not thinking’ or ‘it doesn’t understand’, but I don’t see much engagement with what is actually meant by ‘know’, ‘think’ or ‘understand’ in the first place. Like, you just said that a language model ‘probably… holds representations of spurious facts and relations’. But… isn’t ‘holding a representation of a fact or relation’ basically what we call ‘knowing’ something?
- ultra_nick 4y agoIs there a PG word for bullshitting that has the same meaning?
- theptip 4y agoThis is an important area for AI safety research; see the ELK paper for example. https://www.alignmentforum.org/posts/qHCDysDnvhteW7kRd/arc-s-first-technical-report-eliciting-latent-knowledge https://www.alignmentforum.org/posts/qHCDysDnvhteW7kRd/arc-s... That paper is a bit dense, but considers the ways that a powerful AI model could be intractable/deceptive to discovering its latent knowledge. If we can confidently understand an AI’s internal knowledge/intention states, then alignment is probably tractable.