5 ms·
I have to be careful of confirmation bias when I read stuff like this because I have the intuition that we are uncovering a single intelligence with each of the
by stillpointlab 1y ago
I have to be careful of confirmation bias when I read stuff like this because I have the intuition that we are uncovering a single intelligence with each of the different LLMs. I even feel, when switching between the big three (OpenAI, Google, Anthropic) that there is a lot of similarity in how they speak and think - but I am aware of my bias so I try not to let it cloud my judgement.
On the topic of compression, I am reminded of an anecdote about Heidegger. Apparently he had a bias towards German and Greek, claiming that these languages were the only suitable forms for philosophy. His claim was based on the "puns" in language, or homonyms. He had some intuition that deep truths about reality were hidden in these over-loaded words, and that the particular puns in German and Greek were essential to understand the most fundamental philosophical ideas. This feels similar to the idea of shared embeddings being a critical aspect of LLM emergent intelligence.
This "superposition" of meaning in representation space again aligns with my intuitions. I'm glad there are people seriously studying this.
- b112 1y agoLLMs don't think, nor are they intelligent or exhibiting intelligence. Language does have constraints, yet it evolves via its users to encompass new meanings. Thus those constraints are artificial, unless you artificially enforce static language use. And of course, for an LLM to use those new concepts, it needs to be retokenized by being trained on new data. For example, if we trained LLMs only on books, encyclopedias, newpapers, and personal letters from 1850, it would have zero capacity to speak comprehensibly or even seem cogent on much of the modern world. And it would forever remain in that disconnected positon. LLMs do not think, understand anything, nor learn. Should you wish to call tokenization, learning, then you'd better call a clock "learning" from the gears and cogs that enable its function. LLMs do not think, learn, or exhibit intelligence. (I feel this is not said enough). We will never, ever get AGI from an LLM. Ever. I am sympathetic to the wonder of LLMs. To seeing them as such. But I see some art as wonderous too. Some machinery is beautiful in execution and to use. But that doesn't change truths.
- kevin42 1y agoYou made a lot of bold assertions there. It's as if you have a complete and definitive theory of human intelligence to compare it against. Which if true, would be incredible, because there isn't a scientifically accepted theory, nor is there consensus from a philosophical standpoint. I can't say that you are wrong, you might be right, especially about AGI. And I think it's unlikely that LLMs are the direct path to AGI. But, just looking at how human brains work, it seems unlikely that we would be intelligent either if we used your same reductionist logic. An individual neuron doesn't "think" or "understand" anything. It's a biological cell that simply fires an electrochemical signal when its input threshold is met. It has no understanding of language or context. By your logic, since the fundamental components are just simple signal processors, the brain cannot possibly learn or be intelligent. Yet, from the complex interaction of ~86 billion of these simple biochemical machines, the emergent properties of thought, understanding, and consciousness arise. Dismissing an LLM's capabilities because its underlying operations are basically just math operating on tokenized data is like dismissing human consciousness because it's "just electrochemistry" in a network of cells. Both arguments mistake the low-level mechanism for the high-level emergent phenomenon.
- b112 1y agoWhen an LLM can re-tokenize on the fly, due to newly learned data, let me know. It won't prove intelligence, but at least it won't be static like a book.
- hatthew 1y agohttps://x.com/sukjun_hwang/status/1943703574908723674 https://x.com/sukjun_hwang/status/1943703574908723674 > Tokenization has been the final barrier to truly end-to-end language models. > We developed the H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data
- b112 1y ago