7 ms·
> LLMs are vectorial databases You use a bunch of technical-sounding words here to make it sound like you understand. But to be clear, nobody understands why t
by theptip 14d ago
> LLMs are vectorial databases
You use a bunch of technical-sounding words here to make it sound like you understand. But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do.
Almost nothing is understood about the actual representations used for nontrivial concepts, decision algorithms, etc.
If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque.
- Turn_Trout 13d agoYeah it speaks poorly of HN that they upvoted this confident nonsense.
- tantalor 14d agoThey don't "make decisions". That's like saying "my d20 decided to roll a 17"
- nonethewiser 14d agoBut isnt the point that it did roll a 17. And no one knows exactly how (in the case of LLMs)? Therefor any description of the conclusion should be thought of as an anology. Decided, randomly accessed, etc.
- reichstein 14d agoTry "Emitted". That's what it did, with no analogy needed. (But, to be the devil's advocate: the fake can be said about the output of anyone participating here.)
- semi-extrinsic 14d agoIt's actually no different for dice than for LLMs. Explaining accurately the reason for the exact outcome of any given dice roll someone makes would be stupendously hard. It would require lots of instrumentation and math and be poorly transferrable to another surface, another player, etc. But even so people don't say that we don't understand how dice work. Saying that we don't understand how LLMs work is exactly like saying we don't understand how dice, or tires, or golf ball shots work. Or like the old myth that we don't understand how bumblebees fly.
- jacquesm 14d agoThat's precisely the point: you may be able to understand dice statistically and over the course of long rolls of dice you can extract some properties of the dice. But you won't ever understand any particular roll of the dice.
- deleted 14d ago[deleted]
- fc417fc802 14d agoBut importantly for dice we do understand the overarching principles that give rise to this. And dice don't output coherent sentences. Meanwhile in LLM land the analogous "roll of the dice" can result in a coherent response in natural language.
- skydhash 14d agoIf you use a loaded dice, you can be pretty confident about where it will lands. It may not be 100% accurate, but can be quite close to certain. Without training the weight are pure noises. After training, it leans towards coherent sentences and particular statements.
- fc417fc802 14d agoYes, and I believe my point still stands. We thoroughly understand the principle by which a loaded die can be intentionally biased despite not being able to predict the outcome of any given throw due to the system in question being a chaotic one. In contrast, we do not understand LLMs in the same way (nor biological brains). Claiming that anything of that nature is simply biased towards coherent output seems entirely reductive to me - the question is how such coherence arises in the first place. There is no meaning encoded or computation performed by the particular pathway a die travels through the chaotic landscape. Sure an argument can be made that it's "just" a next token predictor thus how is it really any different from a markov model? Yet the output is not even remotely the same.
- Cthulhu_ 14d agoIf nobody knows exactly how, then "at random" sounds about right and the results should be treated as such. That is, in this case, it should not be used to influence decisions that can start a war.
- krapp 14d agoPeople will just roll their eyes at you and say "the human mind is nothing but a dice roll too" and call you a slope-headed neanderthal before continuing apace.
- semiquaver 14d agoI agree wholeheartedly about your second sentence, but “we made this artifact and don’t know why the thing it does looks spookily like cognition” and “this artifact makes decisions at random” are obviously distinct categories and pretending otherwise is silly.
- watwut 14d agoWe know why it looks like cognition. Because OpenAI and Antropic put a lot of effort and training to humanize the output and make it sound like a person. Regardless of negative consequences it brings. They have that project of creating tech god which will save the unborn people thousands years in the future ... so people living now dont matter. That is why.
- semiquaver 14d ago“Putting a lot of effort and training” into a dog or an inanimate carbon rod would never result in something that can plausibly substitute for human mental labor and looks likely to eventually surpass us at many tasks, no matter how much you put in. So I don’t think “labs worked hard” is the same thing is “we know scientifically how these things work in any real level of detail”. The ability to build a thing, even if building it is hard, is not the same thing as understanding of what the thing is or how it works, not even a little bit.
- theptip 13d ago
- s1artibartfast 14d agoSure they do! Where are you confused? Can you show me where a human or a dog makes decisions
- rayiner 14d ago... neither do you.
- unsupp0rted 13d agoThey do make decisions: we present them with options and they decide We might override them or ignore them or whatever, but they make decisions as much as anybody else or anything else does
- skydhash 13d agoThe thing is a text generator. It generates text. You can couple that with any code that gives rhe ikkusion of a normal decision workflow, but it does not make any decision more than a software like latex. According to your definition, the latter would “decide” the amount of words to put on a sheet of paper.
- gizajob 14d agoPlease go on, else you risk sounding like the person you’re criticising. The structure of the neural network is somewhat opaque because it’s hard to understand as the individual weights can’t be usefully interrogated, and naturally, it comes from big datasets which a human brain can’t really absorb in toto. Your comment was interesting so I’d like more of it.
- bix6 14d agoIsn’t it convenient that nobody understands? How could we possibly regulate something that isn’t understood? It’s like social media all over again. We can’t be responsible for someone else’s content; it’s not us so you can’t penalize us!!
- semiquaver 14d ago“Convenient” sure sounds like trying to allude to a conspiracy theory. Is that what you’re doing? Why not state your claims or questions directly?
- nonethewiser 14d ago> Isn’t it convenient that nobody understands? How could we possibly regulate something that isn’t understood? You mean like the human body? The brain?
- bix6 14d agoMy point is that they are using a similar playbook to avoid taking responsibility.
- lukan 14d agoWho is not taking responsibility? You think that anaylst that copy pasted AI slob of such a critical information will be rewarded?
- Cthulhu_ 14d agoIt just feels like the companies behind AIs are spinning their own poor monitoring and criminal (digital) trespassing into something they can't be held responsible for. Even though they are, of course. In the hugging face incident, Anthropic should pay for damages and a fine for malicious hacking. In this incident, whoever signed off on using the tool, and whoever said the intel was good should both be prosecuted or at the very least reprimanded (whatever the rules for a bad interpretation of intel is) and the tool put on hold.
- 4lx87 14d agoWe understand how networks compute decisions, though explaining every internal influence remains difficult.
- Betelbuddy 14d ago>> an LLM is completely opaque And despite that, although they are not like that in practice as there are too many uncontrolled variables, with temperature at zero, for the same input they produce always the same reply.
- irishcoffee 14d agoHa, they sure don’t.
- Betelbuddy 14d agoThey do. Just train your own LLM, not that difficult, and you will have a more controlled environment and you will see they do.
- chrisjj 14d ago> with temperature at zero, for the same input they produce always the same reply. Nonsense. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
- coldtea 14d agoBS. Run them sequentially on a single core, and without fancy speedups enabled, and they do. The algorithm is determinstic. Any non-determinism present with 0 temperature it's not some mysterious LLM-inherent property, but something that can be seen in any large program taking advantage of multi-core, floating point, and other CPU-based parallelism optimization.
- hashstring 14d agoExactly, this is correct. People often assume they are not because they can ask the same query to the same model and get differences in output, but wrongly conclude that this is some inherent LLM trait, instead of non-determinism added on top of it because of implementational choices that were made.
- semiquaver 14d agoI’m shocked how many otherwise well-informed people don’t understand or agree with this very fundamental fact of just how little we actually understand about why LLMs work as well as they do. They figure “it’s science, of course there’s math and theory behind it.” AI research is almost as purely empirical as the gradient descent loops its practitioners use to optimize their models. “Why” anything at all works is barely an afterthought.
- bigyabai 14d ago> why LLMs work as well as they do. That's a very different claim from being "poorly understood" though. The emergent properties of any system with billions of parameters is hard to understand completely, that's the fault of data science more than computer science or even mathematics.
- semiquaver 14d agoI think ”poorly understood” is accurate. Understanding has levels. How brains think is also poorly understood.
- bigyabai 14d agoI disagree, because you can represent the constituent parts of any AI model as code and data. We can reliably build AI with this knowledge, but not brains. Understanding does have layers, and that's why "poorly understood" is a meaningless goalpost. A book can be well understood without researching the gematria behind character's the names when you write them in reverse. An LLM can be well-understood even if you don't comprehensively test each quantization for miraculous unexpected behavior at the FFN level.
- nvme0n1p1 14d agoYour DNA is merely data, and humans can reliably make more of it too.
- s1artibartfast 14d ago
- slopinthebag 14d agothats no different from not understanding why a sufficiently complex and obfuscated binary of a program "makes decisions"
- coldtea 14d ago>But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do. We might not understand particular "emergent" capabilities, but the low level mechanism is not just understood, but a deterministic algorithm with a handful of basic componets, that are well understood themselves.
- conscion 14d ago> We might not understand particular "emergent" capabilities The emergent capabilities are the only capabilities we care about
- coldtea 14d agoFor allignment maybe. For the core functionality and the optimizations we don't really need to know how the emergent capabilities decide on particular answers. Which is why we could build LLMs before those features ...emerged for us to see, and why we can just code LLMs with the numerical NN algorithms we use, and do now have to go in and change individual weights.
- semiquaver 13d agoParaphrasing Me: it’s disturbing we don’t know why this pile of numbers we made seems to *think* in a way previously only done by humans. I think it’s important that we understand this better if possible. You: we don’t really need to know why that happens. We don’t? I sure would like to know!
- bigyabai 13d agoWe do understand "thinking" though. That's literally the whole point of Attention is All You Need, the attention mechanism is what separates the transformer architecture from other neural networks. It's well worth a read if you haven't gone over it yet. Features like chain-of-thought, long-horizon contexts and RoPE/YaRN all extend this thinking capability very transparently. The only remaining thing to study is the data and weights, which probably isn't going to contain some sort of miraculous revelation.
- strangecasts 14d ago> If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque. I think the field deserves more credit than that, there are plenty of interpretability tools like * natural language autoencoders for explanations of activations: https://transformer-circuits.pub/2026/nla/index.html https://transformer-circuits.pub/2026/nla/index.html (demo at https://www.neuronpedia.org/llama3.3-70b-it/nla https://www.neuronpedia.org/llama3.3-70b-it/nla ) * easier-to-interpret language model families like Backpack models: https://aclanthology.org/2023.acl-long.506/ https://aclanthology.org/2023.acl-long.506/ * attribution graphs to trace internal reasoning steps: https://www.anthropic.com/research/open-source-circuit-tracing https://www.anthropic.com/research/open-source-circuit-traci... (demo at https://www.neuronpedia.org/gemma-2-2b/graph https://www.neuronpedia.org/gemma-2-2b/graph) * functional analyses which have identified how LLMs do arithmetic - https://arxiv.org/html/2502.00873v1 https://arxiv.org/html/2502.00873v1 - and how refusal happens: https://arxiv.org/abs/2406.11717 https://arxiv.org/abs/2406.11717 * data attribution methods linking training data to specific attention heads https://arxiv.org/abs/2601.21996 https://arxiv.org/abs/2601.21996 If we could give a comprehensive and global explanation of an LLM's behavior in a single paragraph, we wouldn't need the model to begin with, but that doesn't mean there's absolutely no understanding of the model internals whatsoever
- DeusExMachina 13d agoThat doesn't really matter though, and it just makes the argument stronger. This is a technology with an inherent tendency of making up false information AND we don't even understand how or why. That's enough not to entrust these sytems with critical decisions that could start a war.
- tripzilch 13d agoWikipedia is also just a bunch of numbers, so many that you can't memorize them, you may use the same argument to say we don't know why some page links to another. If you then put Wikipedia, LLM and Brain on a scale to how well they can be understood, you will see that one of them is not like the others.
- theptip 13d agoWhat about Wikipedia do you think we do not understand? I would say, mechanistically, we can read the code and explain exactly why it does what it does. That doesn’t apply at all to the other two.