11 ms·
Not sure why everyone rates this. It’s full of very confidently made statements like “the AI has no ground truth” (obviously it does, it has ingested every pape
by pjs_ 2y ago
Not sure why everyone rates this. It’s full of very confidently made statements like “the AI has no ground truth” (obviously it does, it has ingested every paper ever), it “can’t reason logically” which seems like a stretch if you ever read the CoT of a frontier reasoning model and “can’t explain how they arrived at conclusions” where - I mean just try it yourself with o1, go as deep as you like asking how it arrived at a conclusion and see if a human can do any better.
In fact the most annoying thing about this article is that it is a string of very confidently made, black and white statements, offered with no supporting evidence, and some of which I think are actually wrong… i.e. it suffers from the same kind of unsubstantiated self confidence that we complain about with the weaker models
- bwfan123 2y agothe machine is fooling you with a mimicry of reasoning. and you are falling for it.
- ZephyrBlu 2y agoWhat is reasoning if not a chain of logically consistent thoughts?
- bwfan123 2y agofair, but "logically consistent thoughts" is a subject of deep investigation starting from the early euclidean geometry to the modern godel's theorems. ie, that logically consistent thinking starts from symbolization, axioms, proof procedures, world models. otherwise, you end up with persuasive words.
- ZephyrBlu 2y agoYou just ruled out 99% of humans from having reasoning capabilities. The beautiful thing about reasoning models is that there is no need to overcomplicate it with all the things you've mentioned, you can literally read the model's reasoning and decide for yourself if it's bullshit or not.
- quantified 2y agoThat's sort of arrogant, Most of that 99 (if that many) % could learn if inspired to and provided resources. And does use reasoning and instinct in day-to-day life even if it's as simple as "I'll take go shopping before I take my car to the shop so I have the groceries" or "hide this money in a new place so my husband doesn't drink it away". Models will get better over time, and yes humans only use models too. Humans rely in cues to tell when each other is fabricating or lying. Machines don't have those cues, and fabricate their reasoning too. So we have a complicatedly difficult time trusting them.
- llm_trw 2y ago>You just ruled out 99% of humans from having reasoning capabilities. After a conversation with humans I think you'd agree 1% of them being able to reason deeply is a vast overestimation. A good example to see how little people can reason is the following classic: > Given the following premises derive a conclusion about your poems: > 1) No interesting poems are unpopular among people of real taste. > 2) No modern poetry is free from affectation. > 3) All your poems are on the subject of soap bubbles. > 4) No affected poetry is popular among people of taste. > 5) Only a modern poem would be on the subject of soap bubbles. The average person on the street won't even know where to start, the average philosophy student will fuck up the translation to first order logic, and a logic professor would need a proof assistant to get it right consistently. Meanwhile o3-mini in 10 seconds: We can derive a conclusion about your poems by following the logical implications of the given premises. Let’s rephrase each premise into a more formal form: Premise 1: No interesting poems are unpopular among people of real taste. This can be reworded as: If a poem is interesting, then it is popular among people of real taste. Premise 2: No modern poetry is free from affectation. This tells us: If a poem is modern, then it is affected (i.e., it shows affectation). Premise 3: All your poems are on the subject of soap bubbles. In other words: Every one of your poems is about soap bubbles. Premise 4: No affected poetry is popular among people of taste. This implies: If a poem is affected, then it is not popular among people of taste. Premise 5: Only a modern poem would be on the subject of soap bubbles. This means: If a poem is about soap bubbles, then it is modern. Now, let’s connect the dots step by step: From Premise 3 and Premise 5: All your poems are on the subject of soap bubbles. Only modern poems can be about soap bubbles. Conclusion: All your poems are modern. From the conclusion above and Premise 2: Since your poems are modern, and all modern poems are affected, Conclusion: All your poems are affected. From the conclusion above and Premise 4: Since your poems are affected, and no affected poem is popular among people of taste, Conclusion: Your poems are not popular among people of taste. From Premise 1: If a poem is interesting, it must be popular among people of taste. Since your poems are not popular among people of taste (from step 3), it follows that: Conclusion: Your poems cannot be interesting. Final Conclusion: Your poems are not interesting. Thus, by logically combining the premises, we conclude that your poems are not interesting.
- Grimblewald 2y agoIf it's mimicry of reason is indistinguishable from real reasoning, how is it not reasoning? Ultimately, an LLM models language and the process behind it's creation to some degree of accuracy or another. If that model includes a way to approximate the act of reasoning, then it is reasoning to some extent. The extent I am happy to agree is open for discussion, but that reasoning is taking place at all is a little harder to attack.
- onemoresoop 2y agoNo, it is distinguishable from real reasoning. Real reasoning, while flawed in various ways, goes through personal experience of the evaluator. LLMs don't have that capability at all. They're just sifting though tokens and associate statistical parameters to it with no skin in the game so to speak.
- danenania 2y agoIt seems like an arbitrary distinction. If an LLM can accomplish a task that we’d all agree requires reasoning for a human to do, we can’t call that reasoning just because the mechanics are a bit different?
- Barrin92 2y agoYes because it isn't an arbitrary distinction. My good old TI-83 can do calculations that I can't even do in my head but unlike me it isn't reasoning about them, that's actually why it's able to do them so fast, and it has some pretty big implications about what it can't do. If you want to understand where a systems limitations are you need to understand not just what it does but how it does it, I feel like we need to start teaching classes on Behaviorism again.
- danenania 2y agoAn LLM’s mechanics are algorithmically much closer to the human brain (which the LLM is modeled on) than a TI-83, a CPU, or any other Turing machine. Which is why, like the brain, it can solve problems that no individual Turing machine can. Are you sure you aren’t just defining reasoning as something only a human can do?
- olalonde 2y agoIf it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck.
- maxdoop 2y agoWhat is reasoning? What is understanding? Do humans do either? How do you know?
- bwfan123 2y agothis is the question that the greeks wrestled with over 2000 years ago. at the time there were the sophists (modern llm equivalents) that could speak persuasively like a politician. over time this question has been debated by philosophers, scientists, and anyone who wanted to have better cognition in general.
- maxdoop 2y agoSo how can you claim what an LLM is doing if we cannot define it regardless?
- krainboltgreene 2y agoI think the third worst part of the GenAI hype era is that every other CS grad now thinks not only is a humanities/liberal arts degree meaningless but now also they're pretty sure they have a handle on the human condition and neurology enough to make judgment calls on what's sentient. If people with those backgrounds ever attempted to broach software development topics they'd be met with disgust by the same people. Somehow it always seems to end up at eugenics and white supremacy for those people.
- bwfan123 2y agomath arose firstly as a language and formalism in which statements could be made with no room for doubt. the sciences took it further and said that not only should the statements be free of doubt, but also that they should be testable in the real world via well defined actions which anyone could carry out. all of this has given us the gadgets we use today. llm, meanwhile, is putting out plausible tokens which is consistent with its training set.
- dontseethefnord 2y agoBecause we know what LLM's do. We know how they produce output. It's just good enough at mimicking human text/speech that people are mystified and stupified by it. But I disagree that "reasoning" is so poorly defined that we're unable to say an LLM doesn't do it. It doesn't need to be a perfect or complete definition. Where there is fuzziness and uncertainty is with humans. We still don't really know how the human brain works, how human consciousness and cognition works. But we can pretty confidently say that an LLM does not reason or think. Now if it quacks like a duck in 95% of cases, who cares if it's not really a duck? But Google still claims that water isn't frozen at 32 degrees Fahrenheit, so I don't think we're there yet.
- mrshadowgoose 2y agoI don't give a rat's ass about whether or not AI reasoning is "real" or a "mimicry". I care if machines are going to displace my economic value as a human-based general intelligence. If a synthetic "mimicry" can displace human thinking, we've got serious problems, regardless of whether or not you believe that it's "real".
- pembrook 2y agoSo are all the humans in this thread. Except, human mimicry of "reasoning" is usually applied in service of justifying an emotional feeling, arguably even less reliable than the non-feeling machine.
- mrbungie 2y agoIt has served us relatively fine for thousands of years. LLMs? I'm waiting for one that knows how not to say something that is clearly wrong with extreme confidence, reasoning or not.
- krainboltgreene 2y ago> obviously it does, it has ingested every paper ever Do you have a citation for such a claim?
- pjs_ 2y agohttps://www.tomshardware.com/tech-industry/artificial-intelligence/meta-staff-torrented-nearly-82tb-of-pirated-books-for-ai-training-court-records-reveal-copyright-violations https://www.tomshardware.com/tech-industry/artificial-intell... https://en.wikipedia.org/wiki/Anna's_Archive https://en.wikipedia.org/wiki/Anna's_Archive https://en.wikipedia.org/wiki/The_Pile_(dataset) https://en.wikipedia.org/wiki/The_Pile_(dataset)
- grepLeigh 2y agoLLMs that use Chain of Thought sequences have been demonstrated to misrepresent their own reasoning [1]. The CoT sequence is another dimension for hallucination. So, I would say that an LLM capable of explaining its reasoning doesn't guarantee that the reasoning is grounded in logic or some absolute ground truth. I do think it's interesting that LLMs demonstrate the same fallibility of low quality human experts (i.e. confident bullshitting), which is the whole point of the OP course. I love the goal of the course: get the audience thinking more critically, both about the output of LLMs and the content of the course. It's a humanities course, not a technical one. (Good) Humanities courses invite the students to question/argue the value and validity of course content itself. The point isn't to impart some absolute truth on the student - it's to set the student up to practice defining truth and communicating/arguing their definition to other people. [1] https://arxiv.org/abs/2305.04388 https://arxiv.org/abs/2305.04388
- ctbergstrom 2y agoYes! First, thank you for the link about CoT misrepresentation. I've written a fair bit about this on Bluesky etc but I don't think much if any of that made it into the course yet. We should add this to lesson 6, "They're Not Doing That!" Your point about humanities courses is just right and encapsulates what we are trying to do. If someone takes the course and engages in the dialectical process and decides we are much too skeptical, great! If they decide we aren't skeptical enough, also great. As we say in the instructor guide: "We view this as a course in the humanities, because it is a course about what it means to be human in a world where LLMs are becoming ubiquitous, and it is a course about how to live and thrive in such a world. This is not a how-to course for using generative AI. It's a when-to course, and perhaps more importantly a why-not-to course. "We think that the way to teach these lessons is through a dialectical approach. "Students have a first-hand appreciation for the power of AI chatbots; they use them daily. "Students also carry a lot of anxiety. Many students feel conflicted about using AI in their schoolwork. Their teachers have probably scolded them about doing so, or prohibited it entirely. Some students have an intuition that these machines don't have the integrity of human writers. "Our aim is to provide a framework in which students can explore the benefits and the harms of ChatGPT and other LLM assistants. We want to help them grapple with the contradictions inherent in this new technology, and allow them to forge their own understanding of what it means to be a student, a thinker, and a scholar in a generative AI world."
- bbor 2y agoHey, I'm definitely on your side of the Great AI Wars--and definitely share your thoughts on the overall framing--but I think you're missing the serious nature of this contribution: 1. Small correction, it's actually a whole book AFAIK, and potentially someday soon, a class! So there's a lot more thought put in then the typical hot-take blog post. I also pop into one of these guy's replies on Bluesky to disagree on stuff fairly regularly, and can vouch for his good faith, humble effort to get it right (not something to be taken for granted!) 2. RE:“the AI has no ground truth”, I'd say this is true, no matter how often they're empirically correct. Epistemological discussions (aka "how do humans think") invariably end up at an idea called Foundationalism, which is exactly what it sounds like: that all of our beliefs can be traced back to one or more "foundational" beliefs that we either do not question at all (axioms) or very rarely do (premises on steroids?). In that sense, this phrase is simply recalling the hallucination debates we're all familiar with in slightly more specific, long-standing terms; LLMs do not have a systematic/efficient way of segmenting off such fundamental beliefs and dealing with them deliberately. Which brings me to... 3. RE:“can’t reason logically”, again this is a common debate that I think is being specified more than usual here. A lot of philosophy draws a distinction between automatic and deliberate cognition. I give credit to Kant for the best version, but it's really a common insight, found in ideas like "Fast vs. Slow thinking"[1], "first order vs. recursive" thought[2], "ego vs. superego"[3], and--most relevantly--intuition vs. reason.[4] At the very least, it's not a criticism to be dismissed out of hand based on empirical success rates! 4. Finally, RE:“can’t explain how they arrived at conclusions”, that's really just another discussion of point 2 in more explicitly epistemic terms. You can certainly ask o3 to reason (hehe) about the cognitive processing likely to be behind a given transcript, but it's not actually accessing any internal state, which is a very important distinction! o3 would do just as well explaining the reasoning behind a Claude output as it would with one of its own. Sorry for the rant! I just leave a lot of comments that sound exactly like yours on "LLMs are useless" blog posts, and I wanted to do my best to share my begrudging appreciation for this work. The title is absurdly provocative, but they're not dismissing LLMs, they're characterizing their weaknesses using a colloquial term -- namely "bullshit" as used for "lying without knowing that you're lying". [1] https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow [2] https://www.mit.edu/~dxh/marvin/web.media.mit.edu/~minsky/papers/SymbolicVs.Connectionist.html https://www.mit.edu/~dxh/marvin/web.media.mit.edu/~minsky/pa... [3] https://en.wikipedia.org/wiki/Id,_ego_and_superego https://en.wikipedia.org/wiki/Id,_ego_and_superego [4] https://plato.stanford.edu/entries/intuition/ https://plato.stanford.edu/entries/intuition/ , and a flawed but interesting one from Gary Marcus: https://garymarcus.substack.com/p/llms-dont-do-formal-reasoning-and https://garymarcus.substack.com/p/llms-dont-do-formal-reason...
- eigenform 2y agoIt depends on your tolerance for error. When you have a machine that can only infer rules for reasoning from inputs [which are, more often than not, encoded in a very roundabout way within a language which is very ambiguous, like English], you have necessarily created something without "ground." That's obviously useful in certain situations (especially if you don't know the rules in some domain!), but it's categorically not capable of the same correctness guarantees as a machine that actually embodies a certain set of rules and is necessarily constrained by them.
- throwaway4aday 2y agoAre you contending that every human derives their reasoning from first principals rather than being taught rules in a natural language?
- eigenform 2y agoI'm contending that, like any good tool, there is a context where it is useful, and a context where it is not (and that we are at a stage where everything looks suspiciously like a nail).
- fmbb 2y agoTraining on all papers does not mean the model believes or knows the truth. It is just a machine that spits out words.
- lifthrasiir 2y agoI would be very careful to claim exactly that as emergent properties seem kinda crucial for artificial and human intelligences. (Not to say that they are equally functioning nor useful.)
- criley2 2y ago>Training on all papers does not mean the model believes or knows the truth. It is just a machine that spits out words. Sounds like humans at school. Cram the material. Take the test. Eject the data.
- joenot443 2y agoIt's 1994. Larry Llyod Mayer has read the entire internet, hundreds of thousands of studies across every field, and can answer queries word for word the same as modern LLMs do. He speaks every major language. He's not perfect, he does occasionally make mistakes, but the sheer breadth of his knowledge makes him among the most employable individuals in America. The Pentagon, IBM, and Deloitte are begging to hire him. Instead, he works for you, for free. Most laud him for his generosity, but his skeptics describe him as just a machine that spits out words. A stochastic parrot, useless for any real work.
- nullc 2y ago> I mean just try it yourself with o1, go as deep as you like asking how it arrived at a conclusion I don't mean to disagree overall, but on this point the LLM can post-facto rationalize its output but it has no introspection and has absolutely no idea why it made a given bit of output (except in so far as it was a result of COT which it could reiterate to you). The set of weights being activated could be nearly disjoint when answering and explaining the answer. One can also make the same argument about humans -- that they can't introspect their own minds and are just posthoc rationalizing their explanations unless their thinking was a product of an internal monolog that they can recount. But humans have a lifetime of self-interaction that gives a good reason to hope that their explanations actually relate to their reasoning. LLM's do not. And LLMs frequently give inconsistent results, it's easy to demonstrate the posthoc nature of LLM's rationalizations too: Edit the transcript to make the LLM say something it didn't say and wouldn't have said (very low probability), and then have it explain why it said that. (Though again, split brain studies show humans unknowingly rationalizing actions in a similar way)
- lanstin 2y agoI doubt people are very accurate at knowing why they made the choices they did. If you want them to recite a chain of reasoning they can but that is kind of far from most decision making most people do.
- nullc 2y agoI agree people aren't great at this either and my post said as much. However we're familiar with the human limits of this and LLMs are currently much worse. This is particularly relevant because someone suffering from the mistaken belief that LLM's could explain their reasoning might go on to attempt to use that to justify the misapplication of an LLM. E.g. fine tune some LLM using resume examples so that it almost always rejects Green-skinned people, but approve the LLMs use in hiring decisions because it is insistent that it would never base a decision on someone's skin color. Humans can lie about their biases of course, but a human at least has some experience with themselves while a LLM usually has no experience observing themself except for the output visible in their current window.
- llm_trw 2y agoI've literally build a dynamic bench mark where I test reasoning models on their performance on deriving conclusions from assumptions through sequent calculus. o3-mini high effort can derive chains that are 8 inference rules deep with >95% confidence I didn't have the money to test it further. This is better than the average professor in logic when given pen and paper. It seems like a course critiquing 5 year old technology at this point.
- radioactivist 2y agoI've had frontier reasoning models (or at least what I can access in ChatGPT+ at any given moment) give wildly inconsistent answers when asked to provide the underlying reasoning (and the CoT weren't always given). Inventing sources and then later denying them mentioned them. Backtracking on statements it claimed to be true. Hiding weasel words in the middle of a long complicated argument to arrive at whatever it decided the answer was. So I'm inclined to believe the reasoning steps here are also susceptible to all the issues discussed in the posted article.
- MichaelZuo 2y agoThis sounds similar to a median human with little scruples?
- cess11 2y agoComputers are "reasoning" in the same sense they have a "heartbeat".
- randomNumber7 2y ago> “can’t explain how they arrived at conclusions” Imagine I would tell my wife, that whenever we have a discussion, her opinion would only be valid when she can explain how she arrived at her conclusion.
- oblio 2y agoYour wife is one of the end products of cutthroat competition across several billion years so let's just say her general intelligence has a fair bit more validation than 20 years of research.
- falcor84 2y agoSexual selection applies an evolutionary pressure against men who challenge women too much about the validity of their reasoning.
- oblio 2y agoI was really, really trying to ignore the casual misogyny in OP's comment but you're really making this hard.
- falcor84 2y agoWell, for what it's worth, I believe that this evolutionary pressure works as strongly, or even more so, against women who challenge men about the validity of their reasoning.
- poulpy123 2y agoBut we know how the LLM works, and that's exactely how the authors explain it. And that explain also the weird mistakes they do, that nothing with the ability of reason or having a ground truth would do. I really do not understand how technical people can think they are sentient
- wg0 2y ago> “the AI has no ground truth” Yeah? It has? Where the irrefutable proof of that?
- yapyap 2y ago> “the AI has no ground truth” (obviously it does, it has ingested every paper ever it does not, AI is predicting the next ‘token’ based on the last ‘token’. There is no sentience, it’s machine learning except the machines are really strong. It’d be illogical to say an AI has a ground truth just because it ‘ingested’ every paper ever.
- pjs_ 2y agoWhat does sentience have to do with truth? I didn’t make that connection, you did. Wikipedia isn’t sentient but it contains a lot of truth. Raw data isn’t sentient but it definitely “has ground truth”.
- deleted 2y ago[deleted]
- MantisShrimp90 2y agoThe writer is speaking from the perspective of the traditional philosophical understanding of a thinking being. No, LLMs are not thinking beings with internal state. Even these "reasoning" models are just prompting the same LLM over and over again which is not true "logic" the way you and I think when we are presented with a new problem. The key difference is they do not have actual logic, they rely on statistical calculations and heuristics to come up with the next set of words. This works surprisingly well if the thing has seen all text written, but there will always be new scenarios, new ideas it has not encountered and no these are not better than a human at those tasks and likely never will be. However, what is happening is that our understanding of intelligence is being expanded, and our belief that we are going to be the only intelligent beings ever is under threat and that makes us fundamentally anxious.
- eqqn 2y ago>“the AI has no ground truth” (obviously it does, it has ingested every paper ever) It also ingested every reddit thread and tweets of every politician ever.