4 ms·
Their lack of self reference is a core problem that undergirds a lot of faults that do occur during inference, but their breadth + the agent harness successfull
by thatjoeoverthr 1mo ago
Their lack of self reference is a core problem that undergirds a lot of faults that do occur during inference, but their breadth + the agent harness successfully covers it well, so it requires a bit of poking to witness. The “hallucination” phenomenon is exactly this. They don’t know the scope of their own knowledge, and they just say stuff, so if you go out of band, it has a higher probability emitting claims that aren’t true. RAG (I don’t mean embedding indices, but any information ingest such as an agent harness executing a search) are somewhat effective in covering for it, enough to make them very useful! But when it does go wrong, it’s generally the same reasons. It has a certain nature and sometimes you run afoul of it.
But I suppose it doesn’t harm its reasoning!
- k__ 1mo agoAny signal that allows the model to see what's wrong helps. Checking code is (relatively) easy, you can use static type checks, linters, and execute it to see if it's correct. Fact checking is harder. A RAG can only check what's in the database, so you have to know what to know beforehand.
- nextaccountic 1mo ago> Fact checking is harder. A RAG can only check what's in the database, so you have to know what to know beforehand. A global database of facts would make easier for AI to stay factual Also ironically it would also make it easier to align AI to do things like consistently censor or distort some political facts
- roenxi 1mo agoThey symptoms sound a lot like humans, so I don't see how it stems from their lack of self reference. Most people you need to keep them in areas they understand or they go to pieces. The lack of self reference just means every time the context clears they reset. They are systems in a permanent state of extreme amnesia.
- thatjoeoverthr 1mo agoIt’s really not like that. If I ask you to tell me about a geographical place I just made up, you can trivially and generally instantly recognize you don’t recognize it. A child can do this. You wouldn’t be able to hold a job or generally get through life without this level of self awareness.
- roenxi 1mo agoSo can LLMs. I asked one about South Wollopop and it suggested that it'd never heard of it but maybe I meant South Wollo in Ethiopia. And that's just a local model with modest hardware and no internet access. First attempt. I mean, if our position is that a hyper advanced statistical model is going to ultimately struggle with the concept of something being unlikely to be true then the statisticians may as well give up in despair. There is no theoretical obstacle here.
- CamperBob2 1mo agoCurrent frontier models are pretty good at telling you when they don't know something, or if you're asking about something nonexistent or nonsensical. Not always, but there has been massive improvement on this front lately. If your last experience with frontier LLMs (read: not models that Google and ChatGPT are giving away for free, but models you have to pay for) was over a year ago, you may not realize that.
- heavenlyblue 1mo agoI had a great game as a kid where me and girls were playing for kisses on "words", i.e. "start the next word with the letter of the word I just said". So I invented a whole new vocabulary and then taught them that vocabulary. Apart from the fact that I got a lot of kisses from them they were so happy they learned something new they immediately went to share this with their (and my) parents. Obviously I got a slap but I just don't get how you can think "a kid can do this". You're either completely ignorant of different cultures (i.e. as an Eqstern European I still find it hard to remember Indian names, let alone tell if they are actually fucking with me) or intentionally simplifying the problem.
- fyredge 1mo agoFrom your comment, I noticed a sort of pattern that is often seen when discussing LLMs. That they are these amazing things that can run so fast they trip themselves in their attempts at achieving a task. So we resort to refining the models, creating guardrails, orchestrating harness, so as to alleviate the 'hallucination' problem. In contrast to human intelligence, there is an underlying mechanism that propels intelligent behaviour. A person is no less intelligent just because they lose sight, sound or inner voice.
- thatjoeoverthr 1mo ago“trip themselves in their attempts at achieving a task” I see this a lot in Claude Code. I assume it has to do with the training structure. Example is “fallbacks”. Claude constantly sprinkles “fallbacks” in the code, even when I ask it not to. That is, write multiple candidate implementations into the same code with some kind of switch. This is a problem because you only need one, and it would seem to have inflated the code for no reason. (The madness accelerates with code volume, so you must push back.) anyway, few people would do it this way. But I thought, what could be the benefit? If you’re being conditioned to pass evals one-shot with code that will be discarded and never read, it’s a great strategy. If you have more than one way to solve it, you can just put both. The behavior would easily be reinforced, if trained that way. But in any case, again, a certain nature and certain conditioning. I think we’ll learn to accept it as AGI but also that no intelligence is fully divorced from context, limits and conditioning.
- andai 1mo agoNot sure self-reference solves the metacognition thing; an ant can pass the mirror test but probably lacks metacognition. Though I haven't read GEB so I'm not sure how the strange loop thing ties in with either of those.
- GPerson 1mo agoYou should read his later work called I Am a Strange Loop instead of GEB, as this is the author’s preference.
- alchemism 1mo agoI think of this book often in the LLM age.
- andai 1mo agoThanks. I actually read this book in high school. I remember being very impressed with it and bringing it to school. I don't remember a single thing! (Which is unusual for me, I usually remember much of what I read.) I shall have to read it again :)