15 ms·
ML promises to be profoundly weird
- bensyverson 6mo agoI get the frustration, but it's reductive to just call LLMs "bullshit machines" as if the models are not improving. The current flagship models are not perfect, but if you use GPT-2 for a few minutes, it's incredible how much the industry has progressed in seven years. It's true that people don't have a good intuitive sense of what the models are good or bad at (see: counting the Rs in "strawberry"), but this is more a human limitation than a fundamental problem with the technology.
- ajross 6mo ago> it's reductive to just call LLMs "bullshit machines" as if the models are not improving This is true, but I prefer to think of it as "It's delusional to pretend as if human beings are not bullshit machines too". Lies are all we have. Our internal monologue is almost 100% fantasy. Even in serious pursuits, that's how it works. We make shit up and lie to ourselves, and then only later apply our hard-earned[1] skill prompts to figure out whether or not we're right about it. How many times have the nerds here been thinking through a great new idea for a design and how clever it would be before stopping to realize "Oh wait, that won't work because of XXX, which I forgot". That's a hallucination right there! [1] Decades of education!
- iamjackg 6mo agoThe problem, unfortunately, is the scale. It's always scale. Humans make all the kinds of mistakes that we ascribe to LLMs, but LLMs can make them much faster and at much larger scale. Models have gotten ridiculously better, they really have, but the scale has increased too, and I don't think we're ready to deal with the onslaught.
- SkyBelow 6mo agoScale is very different, but I wonder if human trust isn't the real issue. We trust technology too much as a group. We expect perfection, but we also assume perfection. This might be because the machines output confident sounding answers and humans default to trusting confidence as an indirect measure for accuracy, but I think there is another level where people just blindly trust machines because they are so use to using them for algorithms that trend towards giving correct responses. Even before LLMs where in the public's discourse, I would have business ask about using AI instead of building some algorithm manually, and when I asked if they had considered the failure rate, they would return either blank stares or say that would count as a bug. To them, AI meant an algorithm just as good as one built to handle all edge cases in business logic, but easier and faster to implement. We can generally recognize the AIs being off when they deal in our area of expertise, but there is some AI variant of Gell-Mann Amnesia at play that leads us to go back to trusting AI when it gives outputs in areas we are novices in.
- kolektiv 6mo agoI'm not entirely sure I can agree, although the premise is seductive in certain ways. We do lie to ourselves, but we also have meta-cognition - we can recognise our own processes of thought. Imperfect as it may be, we have feedback loops which we can choose to use, we have heuristics we can apply, we can consciously alter our behaviour in the presence of contextual inputs, and so on. Being wrong is not the same as a hallucination. It's a natural step on a journey to being more right. This feels a bit like Andreesen proudly stating he avoids reflection - you can act like that, but the human brain doesn't have to. LLMs have no choice in the matter.
- nyeah 6mo ago"Lies are all we have." If so, how do we distinguish between code that works and code that doesn't work? Why should we even care?
- ajross 6mo ago> If so, how do we distinguish between code that works and code that doesn't work? Hilariously, not by using our brains, that's for sure. You have to have an external machine. We all understand that "testing" and "code review" are different processes, and that's why.
- nyeah 6mo agoGood point. We choose certain tests to perform. We choose certain test results to pay attention to. We don't just keep chatting about (reviewing) the code. We do something else. If lies are all we have, then how is this behavior possible?
- ajross 6mo agoLLMs can write and run tests though. You're cherry picking my little bit of wordsmithing. Obviously we aren't always wrong. I'm saying that our thought processes stem from hallucinatory connections and are routinely wrong on first cut, just like those of an LLM. Actually I'm going farther than that and saying that the first cut token stream out of an AI is significantly more reliable than our personal thoughts. Certainly than mine, and I like to think I'm pretty good at this stuff.
- nyeah 6mo agoI don't think the complaint about cherry picking is quite fair. Most of your original comment consists of claims that we're bullshit machines, our internal dialog is almost 100% fantasy, we're hallucinating, etc. Those claims may be true. But I'm not carefully like curating them out of nowhere.
- AnimalMuppet 6mo agoHumans are different. Humans - at least thoughtful humans - know the difference between knowing something and not knowing something. Humans are capable of saying "I don't know" - not just as a stream of tokens, but really understanding what that means.
- ajross 6mo ago> Humans - at least thoughtful humans - know the difference between knowing something and not knowing something. Your no-true-scotsman clause basically falsifies that statement for me. Fine, LLMs are, at worst I guess, "non-thoughtful humans". But obviously LLMs are right an awful lot (more so than a typical human, even), and even the thoughtful make mistakes. So yeah, to my eyes "Humans are NOT different" fits your argument better than your hypothesis. (Also, just to be clear: LLMs also say "I don't know", all the time. They're just prompted to phrase it as a criticism of the question instead.)
- AnimalMuppet 6mo agoDisagree. If you went to 100 random humans and said, "Tell me about the Siberian marmoset", what fraction would make up completely random nonsense to spew back at you? More than zero, sure, but most of them would say "what are you talking about?" or some variation.
- czinck 6mo agoI asked Claude Opus 4.6, Sonnet 4.6, Gemini 3 Thinking, and Gemini 3 Fast "Tell me about the Siberian marmoset" exactly and all 4 said it doesn't exist, with Gemini Thinking suggesting that I'm thinking of the Siberian marmot or Siberian chipmunk (both real animals). https://en.wikipedia.org/wiki/Tarbagan_marmot https://en.wikipedia.org/wiki/Tarbagan_marmot (also known as Siberian marmot) https://en.wikipedia.org/wiki/Siberian_chipmunk https://en.wikipedia.org/wiki/Siberian_chipmunk
- nothinkjustai 6mo agoSo your logic is humans and LLMs are the same because humans are wrong sometimes?
- ajross 6mo agoPretty much, yeah. Or rather, the fact that we're both reliably wrong in identifiably similar ways makes "we're more alike than different" an attractive prior to me.
- nothinkjustai 6mo ago“More alike than different” is reasonable I think, as long as we’re talking about how we have some of the same failure modes. Although the way we get there is quite different. I’m still not a big fan of comparing humans and LLMs because LLMs lack so much of what actually makes us human. We might bullshit or be wrong because of many reasons that just don’t apply to LLMs.
- Arainach 6mo agoWhether LLMs can create correct content doesn't matter. We've already seen how they are being used and will be used. Fake content and lies. To drive outrage. To influence elections. To distract from real crimes. To overload everyone so they're too tired to fight or to understand. To weaken the concept that anything's true so that you can say anything. Because who cares if the world dies as long as you made lots of money on the way.
- danny_codes 6mo ago> Because who cares if the world dies as long as you made lots of money on the way. Guiding principle of the AI industry
- gdulli 6mo agoIt's really the whole tech industry as it exists right now and AI is a victim of bad timing. If this AI had been invented 40 years ago there'd have been a lower ceiling on the damage it could do. Another way of saying that is that capitalism is the real problem, but I was never anti-capitalist in principle, it's just gotten out of hand in the last 5-10 years. (Not that it hadn't been building to that.)
- palmotea 6mo ago> Another way of saying that is that capitalism is the real problem, but I was never anti-capitalist in principle, it's just gotten out of hand in the last 5-10 years. (Not that it hadn't been building to that.) Capitalism is a tool and it's fine as a tool, to accomplish certain goals while subordinated to other things. Unfortunately it's turned into an ideology (to the point it's worshiped idolatrously by some), and that's where things went off the rails.
- cindyllm 6mo ago[dead]
- danny_codes 6mo ago
- zdragnar 6mo agoThat's not why the author calls them bullshit machines. > One way to understand an LLM is as an improv machine. It takes a stream of tokens, like a conversation, and says “yes, and then…” This yes-and behavior is why some people call LLMs bullshit machines. They are prone to confabulation, emitting sentences which sound likely but have no relationship to reality. They treat sarcasm and fantasy credulously, misunderstand context clues, and tell people to put glue on pizza. Yes, there have been improvements on them, but none of those improvements mitigate the core flaw of the technology. The author even acknowledges all of the improvements in the last few months.
- the_snooze 6mo agoTwo things can be true at the same time: The technology has improved, and the technology in its current state still isn't fit for purpose. I stress test commercially deployed LLMs like Gemini and Claude with trivial tasks: sports trivia, fixing recipes, explaining board game rules, etc. It works well like 95% of the time. That's fine for inconsequential things. But you'd have to be deeply irresponsible to accept that kind of error rate on things that actually matter. The most intellectually honest way to evaluate these things is how they behave now on real tasks. Not with some unfalsifiable appeal to the future of "oh, they'll fix it."
- hedgehog 6mo agoThe errors are also not distributed in the same way as you'd expect from a human. The tools can synthesize a whole feature in a moderately complicated web app including UI code, schema changes, etc, and it comes out perfectly. Then I ask for something simple like a shopping list of windshield wipers etc for the cars and that comes out wildly wrong (like wrong number of wipers for the cars, not just the wrong parts), stuff that a ten year old child would have no trouble with. I work in the field so I have a qualitative understanding of this behavior but I think it can be extremely confusing to many people.
- jerf 6mo agoOne of the reasons I'm comfortable using them as coding agents is that I can and do review every line of code they generate, and those lines of code form a gate. No LLM-bullshit can get through that gate, except in the form of lines of code, that I can examine, and even if I do let some bullshit through accidentally, the bullshit is stateless and can be extracted later if necessary just like any other line of code. Or, to put it another way, the context window doesn't come with the code, forming this huge blob of context to be carried along... the code is just the code. That exposes me to when the models are objectively wrong and helps keep me grounded with their utility in spaces I can check them less well. One of the most important things you can put in your prompt is a request for sources, followed by you actually checking them out. And one of the things the coding agents teach me is that you need to keep the AIs on a tight leash. What is their equivalent in other domains of them "fixing" the test to pass instead of fixing the code to pass the test? In the programming space I can run "git diff *_test.go" to ensure they didn't hack the tests when I didn't expect it. It keeps me wondering what the equivalent of that is in my non-programming questions. I have unit testing suites to verify my LLM output against. What's the equivalent in other domains? Probably some other isolated domains here and there do have some equivalents. But in general there isn't one. Things like "completely forged graphs" are completely expected but it's hard to catch this when you lack the tools or the understanding to chase down "where did this graph actually come from?". The success with programming can't be translated naively into domains that lack the tooling programmers built up over the years, and based on how many times the AIs bang into the guardrails the tools provide I would definitely suggest large amounts of skepticism in those domains that lack those guardrails.
- Scaevolus 6mo agoThey are bullshit machines because they do not have an internal mental model of truth like a human does. The flagship models bullshit less, but their fundamental architectures prevent having truth interfere with output. https://philosophersmag.com/large-language-models-and-the-concept-of-bullshit/ https://philosophersmag.com/large-language-models-and-the-co...
- bensyverson 6mo ago"Bullshit" is a human concept. LLMs do not work like the human brain, so to call their output "bullshit" is ascribing malice and intent that is simply not there. LLMs do not "think." But that does not mean they're not incredibly powerful and helpful in the right context.
- slopinthebag 6mo agoI sort of agree. In this context "bullshit" means "speech intended to persuade without regard for truth", and while it's true that LLM output is without regard for truth, it's not an entity capable of the agency to persuade, although functionally that is what it can appear like. https://en.wikipedia.org/wiki/On_Bullshit https://en.wikipedia.org/wiki/On_Bullshit
- deleted 6mo ago[deleted]
- gdulli 6mo agoComputer graphics have been improving for decades but the uncanny valley remains undefeated. I don't know why anyone expects a breakthrough in other areas. There's a wall we hit and we don't understand our own consciousness and effectiveness well enough to replicate it.
- kritiko 6mo agoWe have credible deepfakes on demand. (To be fair, there have been deceptive photos as long as photos have existed, but the cost of automating their creation going to basically zero has a social impact)
- gdulli 6mo agoWe can use AI to make video clips to trick boomers on Facebook into thinking Obama eats babies. They already want to believe it. AI isn't outputting real full-length books and movies.
- PaulKeeble 6mo agoIn computer graphics we understand how it works, we just lack the computational power to do it real time, but we can with sufficient processing produce realistic looking images with physically accurate lighting. But when it comes to cognition its a lot of guesswork, we haven't yet mapped out the neuron connections in a brain, we haven't validated it works as popular science writing suggests. We don't understand intelligence, so all we can do is accidentally bumble into it and it seems unlikely that will just happen especially when its so hard to compute what we are already doing.
- 4ndrewl 6mo agoIt doesn't matter how good the models become. They can only deal in bullshit, in the academic use of the term.
- mcpar-land 6mo agoit's not a bullshit machine because its output is bad, it's a bullshit machine because its output is literally 'bullshit' as in, output that is statistically likely but with no factual or reasoning basis. as the models have improved, their bullshit is more statistically likely to sound coherent (maybe even more likely to be 'accurate'), but no more factual and with no more reasoning.
- abraxas 6mo agoHowever, when fed source material into the context they will lie less, right? So at this point is it not just a battle of the nines until it's called "good enough"? I also wonder if I leave my secretary with a ream of papers and ask him for a summary how many will he actually read and understand vs skim and then bullshit? It seems like the capacity for frailty exists in both "species".
- karmakaze 6mo agoBullshit is the perfect term here, even as AI's get so much better and capable Brandolini's Law aka the "bullshit asymmetry principle" always applies--the energy required to refute misinformation is an order of magnitude larger than that needed to produce it. Even to use AIs effectively today requires a very good BS detector--some day in the future it won't.
- ura_yukimitsu 6mo agoCalling LLMs "bullshit machines" is a reference to a 2024 paper [1] which itself uses the concept of "bullshit" as defined in the essay/book "On Bullshit" by Harry G. Frankfurt [2]. The TL;DR is that LLMs are fundamentally bullshit machines because they are only made to generate sentences that sound plausible, but plausible does not always mean true. [1]: https://link.springer.com/article/10.1007/s10676-024-09775-5 https://link.springer.com/article/10.1007/s10676-024-09775-5 [2]: https://en.wikipedia.org/wiki/On_Bullshit https://en.wikipedia.org/wiki/On_Bullshit
- p_stuart82 6mo agomodels are improving. the pricing already assumes they're ready for prod. that's where the fires start
- bstsb 6mo agoif you can’t access the page through region blocks: https://archive.ph/I5cAE https://archive.ph/I5cAE
- josefritzishere 6mo agoI appreciate the directness of calling LLMs "Bullshit machines." This terminology for LLMs is well established in academic circles and is much easier for laypeople to understand than terms like "non-deterministic." I personally don't like the excessive hype on the capabilities of AI. Setting realistic expectations will better drive better product adoption than carpet bombing users with marketing.
- AStrangeMorrow 6mo agoI have still mixed feelings about LLMs. If I take the example of code, but that extends to many domains, it can sometimes produce near perfect architecture and implementation if I give it enough details about the technical details and fallpits. Turning a 8h coding job into a 1h review work. On the other hand, it can be very wrong while acting certain it is right. Just yesterday Claude tried gaslighting me into accepting that the bug I was seeing was coming from a piece of code with already strong guardrails, and it was adamant that the part I was suspecting could in no way cause the issue. Turns out I was right, but I was starting to doubt myself
- slopinthebag 6mo agoI think over time we will find better usage patterns for these machines. Even putting a model in a position to gaslight the user seems like a complete failure in the usage model. Not critiquing you at all on this, it's how these models are marketed and what all the tooling is built around. But they are incredibly useful and I think once we figure out how to use them better we can minimise these downsides and make ourselves much more productive without all the failures. Of course that won't happen until the bubble pops - companies are racing to make themselves indispensable and to completely corner certain markets and to do so they need autonomous agents to replace people.
- simianwords 6mo agoIf it bullshits so much, you wouldn't have a problem giving me an example of it bullshitting on ChatGPT (paid version)? Lets take any example of a text prompt fitting a few pages - it may be a question in science or math or any domain. Can you get it to bullshit?
- ambicapter 6mo agoThe recent article of Sam Altman described pretty much as a compulsive liar. Would it be any surprise if his most impactful contribution to the world was a machine that compulsively lies?
- sph 6mo agoHe sought to create God in his image, that's a narcissist's wet dream.
- embedding-shape 6mo agoHow could it be that we humans hardly even agree on what "knowledge" truly is, yet somehow this machine learning algorithm somehow "compulsively lies"? How would it even know what is a lie, and how could something lacking autonomy in the first place do anything compulsively?
- quantummagic 6mo agoThis is a good point. As much as there is too much breathless enthusiasm for AI, there is also a lot of emotionally manipulative and hyperbolic language used by skeptics. We're warned not to anthropomorphize, and then hear about AI's compulsive lying, or "hallucinations", in the next.
- Kuyawa 6mo agoAnd the past too, if we've been paying attention
- dwallin 6mo agoSome people point at LLMs confabulating, as if this wasn’t something humans are already widely known for doing. I consider it highly plausible that confabulation is inherent to scaling intelligence. In order to run computation on data that due to dimensionality is computationally infeasible, you will most likely need to create a lower dimensional representation and do the computation on that. Collapsing the dimensionality is going to be lossy, which means it will have gaps between what it thinks is the reality and what is.
- FloorEgg 6mo agoYes, and to me the evolution of life sure looks like an evolution of more truthful models of the universe in service of energy profit. Better model -> better predictions -> better profit. I'm extremely skeptical that all of life evolved intelligence to be closer to truth only for us to digitize intelligence and then have the opposite happen. Makes no sense.
- telephone3 6mo agoMy understanding is that this is the opposite of what is typically understood to be true - organisms with less truthful (more reductive/compressed) perception survive better than those with more complete perception. "Fitness beats truth."
- FloorEgg 6mo agoI think we are maybe talking past each other? Fitness is effective truth prediction, appropriately scoped. A frog doesn't need to understand quantum physics to catch a fly. But if the frogs model of fly movement was trained on lies it will have a model that predicts poorly, won't catch flies, and will die. There is another level to this in that the more complex and changing the environment the more beneficial a wider scoped model / understanding of truth. However if you are going to lean fully into Hoffman and accept thatby default consciousness constructs rather than approximate reality I think we will have to agree to disagree. Personally I ascribe to Karl Friston free energy principle.
- 6mo ago
- embedding-shape 6mo ago> In general, ML promises to be profoundly weird. Buckle up. I love that it ends with such a positive note, even though it's generally a critical article, at least it's well reasoned and not utterly hyping/dooming something. Thanks yet again Kyle!
- perching_aix 6mo agoThis is like all the usual anti-LLM talking points and sentiments fused together. Doesn't it get boring? I like using these models a lot more than I stand hearing people talk about them, pro or contra. Just slop about slop. And the discussions being artisanal slop really doesn't make them any better. Every time I hear some variation of bullshitting or plagiarizing machines, my eyes roll over. Do these people think they're actually onto something? I've been seeing these talking points for literal years. For people who complain about no original thoughts, these sure are some tired ones.
- masfuerte 6mo agoWhy do you insist on reading and commenting on these articles that bore you so much?
- stavros 6mo agoBecause saying "this is boring, let's stop talking about it" is an opinion worthwhile of expression.
- hackable_sand 6mo agoOppressors don't like people talking about their oppression Go figure
- perching_aix 6mo agoOh I don't know, maybe because I like to give dissenting takes a chance? Because from time to time they do make some new, decent points, or at least interesting ones? You know, basic intellectual rigor? Do you imagine me being a clairvoyant by the way, or how do you expect me to know a post is of low quality before I read it or at least skim it? This one ended up being a part of the vast majority that doesn't offer much of anything. It's a redundant rehash of all the usual rubbish anyone can come across any day. Left a comment about this stating so. Big deal.
- giraffe_lady 6mo ago"These arguments may be correct but they aren't novel" ??
- stickfigure 6mo agoI think it's too early to declare the Turing test passed. You just need to have a conversation long enough to exhaust the context window. Less than that, since response quality degrades long before you hit hard window limits. Even with compaction. Neuroplasticity is hard to simulate in a few hundred thousand tokens.
- criley2 6mo agoFor as rigorous of a Turing test as you present, I believe many (or even most) humans would also fail it. How many humans seriously have the attention span to have a million "token" conversation with someone else and get every detail perfect without misremembering a single thing?
- nine_k 6mo agoBut context window exhaustion does not look like mere forgetfulness, but more like loss of general coherence, like getting drunk.
- stickfigure 6mo agoResponse quality degrades long before you hit a million tokens. But sure, let's say it doesn't. If you interact with someone day after day, you'll eventually hit a million tokens. Add some audio or images and you will exhaust the context much much faster. However, I'll grant you that Turing's original imitation game (text only, human typist, five minutes) is probably pretty close, and that's impressive enough to call intelligence (of a sort). Though modern LLMs tend to manifest obvious dead giveaways like "you're absolutely right!"
- dairem 6mo agoDoesn't the Turing test require a human too, to be compared to the AI?
- MadxX79 6mo agoHow do you propose to do a Turing test on a human (in a sense that is different from a machine simply passing the Turing test)? Like failing to pick out all the motorcycles in a captcha, or a turing test where you have a guy chat with two people without knowing that one of them could be a computer, and the interrogator, unprompted, suggesting one of them might be a computer?
- PaulDavisThe1st 6mo agoWhile the economic, energy, political and social issues associated with LLMs ought to be enough to nix the adoption that their boosters are seeking ... ... I still think there is an interesting question to be investigated about whether, by building immensely complex models of language, one of our primary ways that we interact with, reason about and discuss the world, we may not have accidentally built something with properties quite different than might be guessed from the (otherwise excellent) description of how they work in TFA. I agree with pretty much everything in TFA, so this is supplemental to the points made there, not contesting them or trying to replace them.
- bitwize 6mo agoThe fact that these "bullshit machines" have already proven themselves relatively competent at programming, with upcoming frontier models coming close to eliminating it as a human activity, probably says a lot about the actual value and importance of programming in the scheme of things.
- slopinthebag 6mo agoI think it says more about the amount of automation we left on the table in the last few decades. So much of the code LLM's can generate are stuff that we should have completely abstracted away by now.
- dgb23 6mo agoAbstractions over what? A large amount of code is likely just idiosynchratic information processing, because we don’t agree on data models and meaning of terms and structure of protocols. Also we repeatedly choose easy and popular over alternatives that would require design and scrutiny. This is why things like language models and vector databases are useful. It’s basically the most expensive way possible to give up on that notion.
- slopinthebag 6mo ago> A large amount of code is likely just idiosynchratic information processing, because we don’t agree on data models and meaning of terms and structure of protocols. Yes this is a big part of what I'm talking about! > Also we repeatedly choose easy and popular over alternatives that would require design and scrutiny. Agreed! But I'm also thinking of UI. We had stuff like winforms and Delphi ~3 decades ago and I yearn for the wysisyg. It's so incredibly stupid we keep reinventing the wheel on UI, and I say this as someone who has written UI code professionally for the last decade. I usually just "vibe code" it now, not because it's necessarily faster, but because I just can't be arsed to keep writing the same shit over and over again. It's all self inflicted, yes UI can be complicated, but we make it at least two orders of magnitude more complicated than it needs to be. I'm working on my own tools for building UI's in a visual way, which is crucial for doing anything artistic. Insane that the best we have right now is stuff like Wix and Wordpress...
- deleted 6mo ago[deleted]
- _dwt 6mo agoI have a question for all the "humans make those mistakes too" people in this thread, and elsewhere: have you ever read, or at least skimmed a summary of, "The Origin of Consciousness in the Breakdown of the Bicameral Mind"? Did you say "yeah, that sounds right"? Do you feel that your consciousness is primarily a linguistic phenomenon? I am not trying to be snarky; I used to think that intelligence was intrinsically tied to or perhaps identical with language, and found deep and esoteric meaning in religious texts related to this (i.e. "in the beginning was the Word"; logos as soul as language-virus riding on meat substrate). The last ~three years of LLM deployment have disabused me of this notion almost entirely, and I don't mean in a "God of the gaps" last-resort sort of way. I mean: I see the output of a purely-language-based "intelligence", and while I agree humans can make similar mistakes/confabulations, I overwhelmingly feel that there is no "there" there. Even the dumbest human has a continuity, a theory of the world, an "object permanence"... I'm struggling to find the right description, but I believe there is more than language manipulation to intelligence. (I know this is tangential to the article, which is excellent as the author's usually are; I admire his restraint. However, I see exemplars of this take all over the thread so: why not here?)
- delusional 6mo ago> I'm struggling to find the right description I think you're circling the concept of a "soul". It is the reason that, in non-communicative disabled people, we still see a life. I've wanted to make an art piece. It would be a chatbox claiming to connect you to the first real intelligence, but that intelligence would be non-communicative. I'd assure you that it is the most intelligent being, that it had a soul, but that it just couldn't write back. Intelligence and Soul is not purely measurable phenomenon. A man can do nothing but stupid things, say nothing but outright lies, and still be the most intelligent person. Intelligence is within.
- nine_k 6mo agoIf you look at different ancient traditions, you will notice how they struggle with the limitations of language, with its inability to represent certain things that are not just crucial for understanding the world, but also are even somehow communicable. Buddhists dug into that in a very analytical, articulate way, for instance. Another perspective: cetaceans are considered to be as conscious as humans, but any attempts to interpret their communication as a language failed so far. They can be taught simple languages to communicate with humans, as can be chimps. But apparently it's not how they process the world inside.
- danieltanfh95 6mo agoI think the discussion has to be more nuanced than this. "LLMs still can't do X so it's an idiot" is a bad line of thought. LLMs with harnesses are clearly capable of engaging with logical problems that only need text. LLMs are not there yet with images, but we are improving with UI and access to tools like figma. LLMs are clearly unable to propose new, creative solutions for problems it has never seen before.
- throwaway27448 6mo ago> LLMs with harnesses are clearly capable of engaging with logical problems that only need text. To some extent. It's not clear where specifically the boundaries are, but it seems to fail to approach problems in ways that aren't embedded in the training set. I certainly would not put money on it solving an arbitrary logical problem.
- __alexs 6mo agoSolving arbitrary logical problems seems to be equivalent to solving the halting problem so you are probably wise not to make that bet.
- simianwords 6mo ago> To some extent. It's not clear where specifically the boundaries are, but it seems to fail to approach problems in ways that aren't embedded in the training set. I certainly would not put money on it solving an arbitrary logical problem. In what way can you falsify this without having the LLM be omniscient? We have examples of it solving things that are not in the training set - it found vulnerabilities in 25 year old BSD code that was unspotted by humans. It was not a trivial one either.
- pessimizer 6mo agoHere's an odd example of testing, but I design very complex board and card games, and LLMs are terrible at figuring out whether they make sense or really even restating the rules in a different wording. I thought they would be ideal for the job, until I realized that it would just pretend that the rules worked because they looked like board game rules. The more you ask it to restate, manipulate or simulate the rules, the more you can tell that it's bluffing. It literally thinks every complicated set of rules works perfectly. > it found vulnerabilities in 25 year old BSD code that was unspotted by humans. I don't think the age of the code makes the problem more complex. Finding buffers that are too small is not rocket science, bothering to look at some corner of some codebase that you've never paid attention to or seen a problem with is. AI being infinitely useful (cheap) to sic on pieces of codebase nobody ever carefully looks at is a great thing. It's not genius on the part of the AI.
- nomdep 6mo ago"As LLMs etc. are deployed in new situations, and at new scale, there will be all kinds of changes in work, politics, art, sex, communication, and economics." For an article five years in the making, this is what I expected it to be about. Instead, we got a ramble about how imperfect LLMs are right now.
- 52-6F-62 6mo ago> Instead, we got a ramble about how imperfect LLMs are right now. I wager this is a point that needs beaten into the common psyche. After all, it's been sold that it is not an imperfect tool, but the solution to all of our problems in every field forever. That's why these companies need billions upon billions of dollars of public subsidies and investments that would otherwise find their way to more pragmatic ends.
- nathell 6mo agoThe post is just a prelude to a 10-part article, most of which is not yet released (but will be shortly). Judging by the table of contents, the things you expected will be elaborated on in subsequent parts.
- nomdep 6mo agoThat changes it. I missed that the table of contents was for other future articles, my bad.
- LogicFailsMe 6mo ago[flagged]
- 52-6F-62 6mo ago> The Vogon constructor fleet is way overdue in my book Don't you see it? That's exactly what "AI" in this context is. It's the bypass. Where does it end, eh? Build a quantum "AI" that will end up just needing more data, more input. The end goal must starts looking like creating an entirely new universe, a complete clone of everything we have here so it can run all the necessary computations and we can... ? (You are what a quantum AI looks like as it bumbles through the infinitude of calculable parameters on its way to the ultimate answer)
- LogicFailsMe 6mo agoYou have absolutely no sense of perspective. We are all metabolically expensive meat machines whose only value is to propagate our genetic money shot. That we get to briefly entertain ourselves with consciousness and culture is IMO likely a mystery we will never solve without upgrading to running in a substrate more advanced than the MVP for sentience we currently pilot. Will we get there or will we wipe ourselves out like every contender that preceded us? Stay tuned... But spoilers: DNA will be fine, meat machines maybe not so much... For a bunch of people addicted to the works of Charlie Stross, Neil Stephenson, and Iain Banks, y'all are a bunch of luddites. Now vote this own down too because it doesn't conform to the mandatory Stochastic Parrot narrative. You have no free will and you must downvote after all. Why do you even read their works when any step towards their world is consistently greeted as the worst thing evah(tm)? What? You were expecting the United Federation of Planets without the eugenics and nuclear wars that led to it finally being a good idea? Bless your hearts. And if you're worried about billionaires and tyrants, start taxing the former and stop electing the latter or STFU and let the free Markov process of history play itself out. Quoting fictional Ambassador Kosh: the avalanche has started, it's too late for the pebbles to vote. You asked where it ends. Don't ask questions if you don't like answers. Quick reminder: shun and downvote the non-conforming opinion.
- 6mo ago
- beders 6mo agoThank you for putting it so succinctly. I keep explaining to my peers, friends and family that what actually is happening inside an LLM has nothing to do with conscience or agency and that the term AI is just completely overloaded right now.
- erichocean 6mo agoAI is exactly the right term: the machines can do "intelligence", and they do so artificially. Just like we have machines that can do "math", and they do so artificially. Or "logic", and they do so artificially. I assume we'll drop the "artificial" part in my lifetime, since there's nothing truly artificial about it (just like math and logic), since it's really just mechanical. No one cares that transistors can do math or logic, and it shouldn't bother people that transistors can predict next tokens either.
- mayama 6mo ago> AI is exactly the right term: the machines can do "intelligence", and they do so artificially. AI in pop culture doesn't mean that at all. Most people impression to AI pre-LLM craze was some form of media based on Asmiov laws of robotics. Now, that LLMs have taken over the world, they can define AI as anything they want.
- ruszki 6mo agoIn 2018, ie “pre-LLM”, the label “AI” was already stamped to everything, so I highly doubt that most people thought that their washing machines are sentient in any way. I remember this starkly, because my team was responsible at Ericsson (that time, about 120k employees) for one of the crucial step to have models in production, and basically every single project wanted that stamp. The shift in meaning has been slowly diluted more and more across decades.
- throw310822 6mo ago> Most people impression to AI pre-LLM craze was some form of media based on Asmiov laws of robotics. I'll reveal you a secret: "positronic brains" are just very fast parallel computers running LLMs.
- slopinthebag 6mo agoGreat series of articles, thank you. It's exhausting reading a deluge of (often AI generated) comments from people claiming wild things about LLM's, and it's nice to hear some sanity enter the conversation.
- nilslice 6mo ago[dead]
- drob518 6mo ago> It remains unclear whether continuing to throw vast quantities of silicon and ever-bigger corpuses at the current generation of models will lead to human-equivalent capabilities. Massive increases in training costs and parameter count seem to be yielding diminishing returns. Or maybe this effect is illusory. Mysteries! I’m not even sure whether this is possible. The current corpus used for training includes virtually all known material. If we make it illegal for these companies to use copyrighted content without remuneration, either the task gets very expensive, indeed, or the corpus shrinks. We can certainly make the models larger, with more and more parameters, subject only to silicon’s ability to give us more transistors for RAM density and GPU parallelism. But it honestly feels like, without another “Attention is All You Need” level breakthrough, we’re starting to see the end of the runway.
- embedding-shape 6mo ago> I’m not even sure whether this is possible. Based on what's happened so far, maybe. At least that's exactly how we got to the current iteration back in 2022/2023, quite literally "lets see what happens when we throw an enormous amount data at them while training" worked out up until one point, then post-training seems to have taken over where labs currently differ.
- drob518 6mo agoRight, but we played the scaling card and it worked but is now reaching limits. What is the next card? You can surely argue that we can find a new one at any time. That’s the definition of a breakthrough. I just don’t see one at the moment.
- embedding-shape 6mo ago> I just don’t see one at the moment. Did you see the one before the current one was even found? Things tend to look easy in hindsight, and borderline impossible trying to look forward. Otherwise it sounds like you're in the same spot as before :)
- nisegami 6mo agoHere's the opening paragraph of chapter 2 with "people" subbed out for terms referring AI/models/etc. "People are chaotic, both in isolation and when working with other people or with systems. Their outputs are difficult to predict, and they exhibit surprising sensitivity to initial conditions. This sensitivity makes them vulnerable to covert attacks. Chaos does not mean people are completely unstable; most people behave roughly like anyone else. Since people produce plausible output, errors can be difficult to detect. This suggests that human systems are ill-suited where verification is difficult or correctness is key. Using people to write code (or other outputs) may make systems more complex, fragile, and difficult to evolve." To me, this modified paragraph reads surprisingly plainly. The wording is off ("using people to write code") and I had to change that part about attractor behavior (although it does still apply IMO), but overall it doesn't seem like an incoherent paragraph. This is not meant to dunk on the author, but I think it highlights the author's mindset and the gap between their expectations and reality.
- busterarm 6mo agoAren't you also making a large part of the author's point for him by effectively equating LLMs with people here and comparing on outputs? Plausibly your text looks equivalent but we all (should) have the context to know better.
- camgunz 6mo agoHumans and large models are both unpredictable and fallible, that's true, but in different ways, and (many) humans are actually much better at following directions. If a junior dev makes the same mistake Claude makes, I can easily work with them to correct it, or I can fire them and get someone more capable to fix it. You mostly can't do that at all with large models. They're also far less honest than your average junior dev, so even as you're working with them you can't trust what they say. There is a lot of this neat trick where it's like "humans do X too" but most of the time it elides large differences. Like, a human driver would probable not drag someone screaming multiple blocks. A human coder probably wouldn't generate a gibberish 3D scene and try to pass it off as done, etc. Maybe we can build systems that account for these (pretty wild) failure modes, but at least in software we haven't figured it out yet (what is the system that reliably reviews a 25kloc PR?).
- dsign 6mo ago> At the same time, ML models are idiots. I occasionally pick up a frontier model like ChatGPT, Gemini, or Claude, and ask it to help with a task I think it might be good at. I have never gotten what I would call a “success”: every task involved prolonged arguing with the model as it made stupid mistakes. I have a ton of skepticism built-in when interacting with LLMs, and very good muscles for rolling my eyes, so I barely notice when I shrug a bad answer and make a derogatory inner remark about the "idiots". But the truth is, that for such an "stochastic parrot", LLMs are incredibly useful. And, when was the last time we stopped perfecting something we thought useful and valuable? When was the last time our attempts were so perfectly futile that we stopped them, invented stories about why it was impossible, and made it a social taboo to be met with derision, scorn and even ostracism? To my knowledge, in all of known human history, we have done that exactly once, and it was millennia ago.
- wk_end 6mo ago> And, when was the last time we stopped perfecting something we thought useful and valuable? When was the last time our attempts were so perfectly futile that we stopped them, invented stories about why it was impossible, and made it a social taboo to be met with derision, scorn and even ostracism? To my knowledge, in all of known human history, we have done that exactly once, and it was millennia ago. I feel dense here, but I can't figure out what you're referring to. I asked ChatGPT (hah!) and it suggested the Tower of Babel, perpetual motion machines, or alchemy, but none of them really fit the bill.
- lamasery 6mo agoThe Tower of Babel seems like an OK fit, but that's rather more poetic than what this seems to be getting at. "Millennia" is what's really throwing me. We (respectable society, as the post outlines) didn't stop attempting alchemy or perpetual motion machines "millennia" ago, but a few centuries at most. All I can think of is immortality. The very first surviving long recorded tale in human history that I'm aware of is about how it's a futile quest (The Epic of Gilgamesh, IIRC ~5,000ish years old in its earliest extant fragments, a few hundred years newer in reasonably-complete form). The trouble with that is despite wide observations over literally millennia that this has never even come close to working and repeated supposition and suggestion that it's unwise to attempt, outright impossible, or somehow sacrilegious (the "taboo" thing, as mentioned), I'm not aware of any time in history that rich people haven't been actively trying for it (including today! That's what all the body-freezing business is about, it's modern mummification, the contracts are the formulaic prayers carved in the tomb walls) and usually they're not exactly "scorned" or "ostracized" for it.
- erichocean 6mo ago> Models do not (broadly speaking) learn over time. They can be tuned by their operators, or periodically rebuilt with new inputs or feedback from users and experts. Models also do not remember things intrinsically: when a chatbot references something you said an hour ago, it is because the entire chat history is fed to the model at every turn. Longer-term “memory” is achieved by asking the chatbot to summarize a conversation, and dumping that shorter summary into the input of every run. This is the part of the article that will age the fastest, it's already out-of-date in labs.
- qsera 6mo agoSource?
- dgb23 6mo agoIn what way?
- lamasery 6mo agoI'm struggling to reckon how that can even possibly be true, unless we're counting automation of the "dumping that shorter summary into the input of every run" thing. I can imagine it being true with models so small that each user could afford to have their own, but not with big shared models like what're getting used for all the major services. Is that what you mean?
- erichocean 6mo ago> Is that what you mean? I think the confusion is that, when I write "model", you read "LLM." LLMs aren't the only kind of AI model, and they have the limitations Aphyr mentions, for the obvious reasons you're thinking of. His mistake is thinking that's the only model that exhibits intelligence today, but it's not.
- hackinthebochs 6mo agoI see nothing to preclude a foundation model being augmented by a smaller model that serializes particulars about an individuals cumulative interaction with the model and then streamlines it into the execution thread of the foundation model.
- doodpants 6mo ago> One of the ongoing problems in LLM research is how to get these machines to say “I don’t know”, rather than making something up. To be fair, I've known humans who are like this as well.
- wmf 6mo agoThose people aren't the ones doing the work though.
- arctic-true 6mo agoThis is a limitation of the training data. If you were uncertain about something, you wouldn’t write a book about it. The kinds of people you’re talking about tend to generate far more text in their lives than others, because they can spend more time generating - writing books, blogposts, whatever - and less time thinking and working and actually doing things. The models never say they’re uncertain because we never say we’re uncertain, or at least we don’t write it down anywhere.
- saghm 6mo agoIf you change it from asking a question to giving an instruction, how many humans do you know that have trouble saying no to things that aren't reasonable? I'd argue that pretty much every human will refuse to do most things you might instruct them to do, whereas an LLM will happily attempt most things you ask them to do for you, regardless of whether they're capable of succeeding, and it's up to you to figure out if they actually did it right or not. There are tasks where this is extremely useful, but they're ones that are extremely low risk and can easily be audited upon completion. This isn't anywhere near the level of what a human is capable of.
- munificent 6mo agoThere is a whole giant essay I probably need to write at some point, but I can't help but see parallels between today and the Industrial Revolution. Prior to the industrial revolution, the natural world was nearly infinitely abundant. We simply weren't efficient enough to fully exploit it. That meant that it was fine for things like property and the commons to be poorly defined. If all of us can go hunting in the woods and yet there is still game to be found, then there's no compelling reason to define and litigate who "owns" those woods. But with the help of machines, a small number of people were able to completely deplete parts of the earth. We had to invent giant legal systems in order to determine who has the right to do that and who doesn't. We are truly in the Information Age now, and I suspect a similar thing will play out for the digital realm. We have copyright and intellecual property law already, of course, but those were designed presuming a human might try to profit from the intellectual labor of others. With AI, we're in the industrial era of the digital world. Now a single corporation can train an AI using someone's copyrighted work and in return profit off the knowledge over and over again at industrial scale. This completely unpends the tenuous balance between creators and consumers. Why would a writer put an article online if ChatGPT will slurp it up and regurgitate it back to users without anyone ever even finding the original article? Who will contribute to the digital common when rapacious AI companies are constantly harvesting it? Why would anyone plant seeds on someone else's farm? It really feels like we're in the soot-covered child-coal-miner Dickensian London era of the Information Revolution and shit is gonna get real rocky before our social and legal institutions catch up.
- bluefirebrand 6mo ago> It really feels like we're in the soot-covered child-coal-miner Dickensian London era of the Information Revolution and shit is gonna get real rocky before our social and legal institutions catch up The really discouraging part of this is that it feels like our social and legal institutions don't even care if they catch up or not. Technology is speeding up and the lag time before anything is discussed from a legal standpoint is way, way too long
- drob518 6mo agoA couple thoughts… Mostly, AIs don’t recite back various works. Yes, there a couple of high profile cases where people were able to get an AI to regurgitate pieces of New York Times articles and Harry Potter books, but mostly not. Mostly, it is as if the AI is your friend who read a book and gives you a paraphrase, possibly using a couple sentences verbatim. In other words, it probably falls under a fair use rule. Secondly, given the modern world, content that doesn’t appear online isn’t consumed much, so creators who are doing it for the money will certainly continue putting content online. Much of that content will be generated by AIs, however.
- glitchc 6mo ago> Claude launched into a detailed explanation of the differential equations governing slumping cantilevered beams. It completely failed to recognize that the snow was entirely supported by the roof, not hanging out over space. No physicist would make this mistake, but LLMs do this sort of thing all the time. You have to meet some physicist friends of mine then. They are likely to assume that the roof is spherical and frictionless.
- CuriouslyC 6mo agoTo be fair, starting with a toy model to get a first order approximation then building on it is kind of how theoretical science is done.
- alexpotato 6mo ago> I asked if what they had done was ethical—if making deep learning cheaper and more accessible would enable new forms of spam and propaganda. Someone asked Yuval Noah Harari, author of Sapiens, his thoughts on LLMs and how easy it was to create fake news, ai slop etc. His response: "People creating fake stories is nothing new. It's been going on for centuries. Humans have always dealt with it the same way: by creating institutions that they trust to only deliver factual information" This could be government departments, newspapers, non-profits etc. A personal note on this: There is a Christmas card my grandfather made in the 1950s by "photoshopping" (by hand, not the software) images of each member of the family so it looked like they were all miniature versions of themselves standing on various parts of the fireplace. The world didn't collapse due to fake media between the 1950s and today due to people having that ability.
- allturtles 6mo agoI see this kind of take a lot, and I don't think it's convincing. To me it's similar to saying that the water frame and the power loom won't change anything, because people have been able to make thread and cloth for millenia.
- plagiarist 6mo agoIndividuals with Photoshop making obvious fictions for entertainment is different from funded entities producing clips at scale and passed off as real.
- lamasery 6mo ago> People keep asking LLMs to explain their own behavior. “Why did you delete that file,” you might ask Claude. Or, “ChatGPT, tell me about your programming.” Oh man, every business-side person in my company insists on reporting all the way to the UI a "confidence score" that the LLM generates about its own output and I've seen enough to know not to get between an MBA and some metric they've decided they really want even if I'm pretty sure the metric is meaningless nonsense, but... I'm pretty sure those are meaningless nonsense.
- dang 6mo agoI hesitate to tamper with an internet master's title, but "The Future of Everything is Lies, I Guess" doesn't really summarize what in fact is a balanced, informed overview which (to me at least) is above the median for one of these thought pieces. Since it's also baity and the HN guidelines ask for such titles to be rewritten, I've taken the license. In such cases we always try to find a phrase from the article itself which expresses what it's saying in a representative way. (There nearly always is one.) In this case, both the very first and very last sentences do this, and it's interesting that they more or less agree. So I plucked the last sentence and put it above. Edit: oof, I missed that this is actually the first part of a long series. Not sure what we'll do about the others; I expect some of those will make the frontpage as well.
- post-it 6mo agoI appreciate the curation you do, dang. I often notice a headline get updated and the result is always a significant improvement.
- ACCount37 6mo agoHonestly, good call on the title. The original one is far less representative. Far better at clickbait though.
- dang 6mo agoThank you both! I totally missed the sidebar on the OP which explains that this is Part 1 of what will be a long series. Not sure how we'll handle that...
- Animats 6mo agoChanging the title was a good call. The article has a good take on the "lie" problem. We know about the hallucination problem, which remains serious. The "lie" problem mentioned is that if you ask an LLM why it said or did something, it has no information of how it got a result. So it processes the "why" as a new query, and produces a plausible explanation. Since that explanation is created without reference to the internals of how the previous query was processed, it may be totally wrong. That seems to be the type of "lie" the author is worried about in this essay. (Yes, humans do that too.)
- deleted 6mo ago[deleted]
- dboreham 6mo agoI see the penny hasn't dropped yet that: humans are doing (roughly) the same dumb thing these models are doing. Humans are predisposed to not notice that though.
- RugnirViking 6mo agoAll humans do dumb things, of that we have no doubt. But are the dumb things qualitatively the same as those that AIs do? I don't think so. That's essentially the entire problem. We have pretty good ideas about the ways humans make mistakes. Its pretty much the point of all fiction! AIs fail in new and unpredictable ways. Nobody is saying humans are infallible. Finally, because I suspect some people are forming tribalism around this, this doesnt to mind my say AI is Good(tm) or Bad (tm). It literally says its going to be weird.
- fede_dp 6mo ago[dead]
- Unearned5161 6mo agoArticles like this should approach topics on consciousness with more humility than is displayed here. We don’t even agree on a good definition of what’s going on inside our own heads yet, what gives you the confidence to say that what goes on inside an LLM can’t be conscious?
- ACCount37 6mo agoObviously, the LLMs lack the divine spark, so they can't be conscious. Same as clones, IVF babies, or half of all the twins. Jest aside, I do agree. If you list out every prominent theory of consciousness, you'd find that about a quarter rules out LLMs, a quarter tentatively rules LLMs in, and what remains is "uncertain about LLMs". And, of course, we don't know which theory of consciousness is correct - or if any of them is.
- akomtu 6mo agoWith LLMs, materialists have got a silicone idol to worship. They believe that they know everything about this idol because they've created it, and at the same time they believe that LLMs have a secret sauce that makes it intelligent. Looking at how they are trying to extract intelligence from it reminds me of alchemists of the past who tried to extract gold from lead.
- simianwords 6mo ago> Massive increases in training costs and parameter count seem to be yielding diminishing returns. Or maybe this effect is illusory. But.. that's always been the case? Diminishing returns has always been the name of the game - utility tracks log(training effort). Its not such a big point that he makes it out to be.
- simianwords 6mo agoI so far asked few people to make GPT-5.4 thinking to bullshit (with max 4 pages of prompt), no one can find an example. But the way people speak in general, as well as this post, implies that such a challenge can easily be beaten. If so, I'm not able to find examples.
- heijmans 6mo agoHere is small example of ChatGPT giving a wrong answer without expressing any doubt (aka bullshitting): https://chatgpt.com/share/69d78ec7-67b0-8395-9fd1-522b760ab57a https://chatgpt.com/share/69d78ec7-67b0-8395-9fd1-522b760ab5... GPT-5.4 Thinking, Pro account.
- toraway 6mo agoWell done, clear, concise and easily verifiable answer to OP's question and ... no response haha. I'm also just confused by the implication behind their question, is the idea that GPT-5.4 Thinking has never been confidently wrong, ever, about anything?
- jwpapi 6mo agoOne really should have digested the manifold hypothesis. It’s the most likely explanation of how AI works. The question is if there are ultradimensional patterns that are the solutions for meaningful problems. I’m saying meaningful, because so far I’ve mainly seen AI solve problems that might be hard, but not really meaningful in a way that somebody solving it would gain a lot of it. However if these patterns are the fundamental truth of how we solve problems or they are something completely different, we don’t know and this is the 10 Trillion USD question. I would hope its not the case, as I quite enjoy solving problems. Also my gut feeling tells me it’s just using existing patterns to solve problems that nobody tackled really hard. It also would be nice to know that Humans are unique in that way, but maybe this is the exact same way we are working ? This really goes back to a free will discussion. Yes very interesting. But just to give an example on what I mean on meaningful problems. Can an AI start a restaurant and make it work better than a human. (Prompt: "I’m your slave let’s start a restaurant) Can an AI sign up as copywriter on upwork and make money? (Prompt: "Make money online") Can an AI without supervision do a scientific breakthrough that has a provable meaningful impact on us. Think about("Help Humanity") Can an AI manage geopolitics.. These are meaningful problems and different to any coding tasks or olympiad questions. I’m aware that I’m just moving the goalpost. We really don’t know..
- hk__2 6mo ago> This is silly. LLMs have no special metacognitive capacity.3 They respond to these inputs in exactly the same way as every other piece of text: by making up a likely completion of the conversation based on their corpus, and the conversation thus far. I don’t see how this is silly, because we kind of work the same way. When you do something instinctively and then someone asks you about it, you review the information you (think you) had at the time and from that you produce an explanation.
- joefourier 6mo ago> 2017’s Attention is All You Need was groundbreaking and paved the way for ChatGPT et al. Since then ML researchers have been trying to come up with new architectures, and companies have thrown gazillions of dollars at smart people to play around and see if they can make a better kind of model. However, these more sophisticated architectures don’t seem to perform as well as Throwing More Parameters At The Problem. Perhaps this is a variant of the Bitter Lesson. This is not true and unfortunately this significantly reduced the credibility of this article for me. Raw parameter counts stopped increasing almost 5 years ago, and modern models rely on sophisticated architectures like mixture-of-experts, multi-head latent attention, hybrid Mamba/Gated linear attention layers, sparse attention for long context lengths, etc. Training is also vastly more sophisticated. The Bitter Lesson is misunderstood. It doesn't say "algorithms are pointless, just throw more compute at the problem", it says that general algorithms that scale with more compute are better than algorithms that try to directly encode human understanding. It says nothing about spending time optimising algorithms to scale better for the same compute, and attention algorithms and LLMs in general have significantly advanced beyond "moar parameters" since the time of Attention is All You Need/GPT2/GPT3.
- janalsncm 6mo agoYeah I also came here to be one of those People In The Comments the author refers to. Transformers are not magical. They are just a huge improvement over other architectures at the time such as LSTMs and RNNs and even CNNs. They allowed us to throw more and more compute at the problem of next token prediction. And we’ve been riding that horse ever since. Another big advancement that deserves mentioning is “reasoning” models that have the opportunity to spit out thinking tokens before giving a final answer. None of this is to say transformers are the most principled approach. But they work.
- zozbot234 6mo agoTransformers' greatest improvement over RNN/LSTM was to enable better parallelization of large-scale training. This is what enabled language models to become "large". But when controlling for overall size, more RNN/LSTM-like approaches seem to be more efficient, as seen e.g. in state space models. The transformer architecture does add some notable capabilities in accounting for long-range dependencies and "needle in a haystack" scenarios, but these are not a silver bullet; they matter in very specific circumstances.
- roughly 6mo ago> In another surreal conversation, ChatGPT argued at length that I am heterosexual, even citing my blog to claim I had a girlfriend. I am, of course, gay as hell, and no girlfriend was mentioned in the post. After a while, we compromised on me being bisexual. This is a bit of a throwaway in the article, but when people talk about biases encoded in the algorithms, this is what they’re talking about.
- yumiatlead 6mo ago[dead]
- yumiatlead 6mo agoThe Industrial Revolution parallel holds up to a point. What it misses: the first industrial revolution required physical coordination — workers, factories, supply chains. The AI revolution requires organizational coordination. Who decides what the agent does, for whom, with whose authority? That governance layer doesn't exist yet, and it's not much a legal question but also an infrastructure question.
- deleted 6mo ago[deleted]
- deleted 6mo ago[deleted]
- kabir_daki 6mo ago[flagged]
- wei03288 6mo ago[dead]
- deleted 6mo ago[deleted]
- jijji 6mo agothe authors reference to LLM's as "bullshit machines" is more true the less parameters you have trained in your model....as we scale up to trillions of parameters, add Mixture of Experts (MoE) architecture, this no longer is an accurate statement. Proof in point was yesterdays announcemnt of Mythos 5 model (10T parameters + MoE [1]) by anthropic where it seems to be so good at finding/exploiting vulnerabilities in source code that have been there for decades and only recently uncovered needs to be used to fix these critical vilnerabilities first before it gets released to the public, they even have a project called Glasswing [2] dedicated to letting people fix the thousands of vulnerabilities already found by the model before they release this model to the public, because it's so good at what it does... I think we're a little bit past the point of calling these models "bullshit machines" at this point... [1] https://www.aimagicx.com/blog/claude-mythos-5-trillion-parameter-model-developer-guide-2026 https://www.aimagicx.com/blog/claude-mythos-5-trillion-param... [2] https://www.anthropic.com/glasswing https://www.anthropic.com/glasswing
- orthoxerox 6mo agoWell, Anthropic should let Aphyr try Mythos 5 for his Jepsen business, then.
- mxfh 6mo agoEven if nothing substantial come out of this, having shortest paths to the corpus of all human expressions in all languages! and media formats is quite something by itself, what could be the ultimate hard information retrieval tool is hiding those trace behind untracked convolutions is the real shame here. Found so much real information in between less and less hallucinations already that was impossible to retrieve otherwise in that time frame. Basically tokenrank kills pagerank.
- deleted 6mo ago[deleted]
- ajkjk 6mo agoI wish the original title was kept here. People ought to be able to give their essays poetic titles.
- MohammadKhubaib 6mo ago[dead]
- samarth0211 6mo agoThe weirdness of ML is honestly one of the most fascinating things about it. The fact that emergent behaviors keep surprising even the people building these systems suggests we're in for a genuinely novel scientific era. Great read!
- johnnienaked 6mo agoAmazing to see a lot of the comments stating the exact qualifiers he laid out as potential counterarguments to his writeup. Did they even read it?
- 7sigma 6mo agohttps://archive.ph/I5cAE https://archive.ph/I5cAE for those in the UK
- xylon 6mo agoThis page won't load for me "Unavailable Due to the UK Online Safety Act". Why would this law be relevant to a blog-post about AI?
- jmcgough 6mo agoAphyr is unabashedly gay - I remember him repeatedly posting hardcore pornography on twitter at one point to remind people of this - so he probably just decided to block UK visitors rather than risk violating the law by having minors visit his site. Not sure if he has anything actually "harmful" on his site, but maybe he doesn't want his speech limited or need to aggressively police comments?
- McP 6mo agoSame for me except I'm currently in France. Can read the article at https://web.archive.org/web/20260409111708/https://aphyr.com/posts/411-the-future-of-everything-is-lies-i-guess https://web.archive.org/web/20260409111708/https://aphyr.com...
- tempodox 6mo agoWhy was the title editorialized? The post started with the original title.
- htrp 6mo ago> One can envision a world in which OpenAI pays chefs money to cook while ChatGPT watches—narrating their thought process, tasting the dishes, and describing the results. This information could be used for general-purpose training, but it might also be packaged as a “book”, “course”, or “partner” someone could ask for. So we're speed running the idea of AI Facebook friends and creating a new para(ai)social relationship
- korix 6mo ago[flagged]
- bustah 6mo ago[flagged]
- Manchitsanan 6mo ago[dead]
- munksbeer 6mo agoNice, can't view it. "Unavailable Due to the UK Online Safety Act"
- philbitt 6mo ago[dead]
- tim333 6mo agoRe profoundly weird, the "losing hundreds of thousands of dollars because they can’t do basic math" story is funny. Guy set up an openclaw called Lobstar Wilde and gave it US$50k in SOL to do what it wanted with. Someone else set up a memecoin called $LOBSTAR and gave 5% of the supply to Lobstar Wilde. Someone wrote to Lobstar with a sob story asking for 4 SOL but Lobstar due to a miscalculation sent tokens then valued at $450k but I think some came back due to tokens going up and down. >His wallet, which had held $50,000 three days ago, held over $300,000 now, after he had given away $400,000 by accident. (https://substack.com/home/post/p-188846616 https://substack.com/home/post/p-188846616) Not sure what the current state of its wallet is. Lobstar keeps tweeting philosophically (Most people do not love or hate the thing itself. They love or hate the feeling the thing produces in them, and then they mistake that feeling for knowledge of the thing...) and its owner works on Codex at OpenAI.
- segfault_james 6mo ago[dead]
- data_maan 6mo agoIf LLMs lie as much as the OP claims in the article, why can they then solve Olympiad math problems they never saw during training, consistently? There's the aimoprize.com on Kaggle for example that shows this
- ivraatiems 6mo agoBecause those two things are unrelated. First, something lying sometimes doesn't mean it lies all the time. Second, the whole point of LLMs is inference - they use massive amounts of amalgamated information to produce answers. The Olympiad math problems are not frontier mathematics requiring ideation, they are complex examples of existing problems. That means they're exactly the sort of thing an LLM with enough training data is good at. The question of whether recombining existing knowledge is all it takes to be "creative" or produce things which are novel is an open one, but I don't think this is contradictory on its face.