6 ms·
Sliding past the mistakes pointed out, shifting goalposts and trying to recover I see. Let’s say I’m a complete moron and don’t know what a sparse embedding is
by randomImmigrant 29d ago
Sliding past the mistakes pointed out, shifting goalposts and trying to recover I see.
Let’s say I’m a complete moron and don’t know what a sparse embedding is.
Pretty please, can you define it for me and then tell me, in detail, where in whatever region of the brain you think this is going on… how is it going on?
Explain how “memories must be stored as embeddings with single multi-neuron assemblies (cortical columns?) storing multiple embeddings as a kind of contents-addressable memory”
You have moved past the cortical column. But still seem to be insisting it’s a bunch of neurons, somewhere… or has that also conveniently changed? Whatever your current position is, please go ahead and explain what components of what cells or otherwise are involved in this process you’re describing.
- HarHarVeryFunny 28d agoAn embedding space is a (typically) high dimensional space that has enough dimensions such that examples of some type of entity (e.g. faces, words, or thoughts) can be represented as points in that space, positioned such that they are nearby to other entities with which they have things in common. An entity embedding doesn't need to use all the dimensions of the space it is positioned in - some dimensions may be unused (sometimes represented as a coorrdinate of 0 in that dimension). These are called "sparse" embeddings. For example, an LLM's tokens are represented as embeddings in what is typcially an approximately ~1000 dimensional space, but start out as sparse embeddings just representing a short letter sequence (but then go on to be transformed/augmented with additional information and so become less sparse). As an example, let's say an embedding space has 10 dimensions, then a couple of sparse embedding examples could be: [0 0 1 0 0 1 1 0 0 0] [1 1 0 0 0 0 0 1 0 0] These two embeddings have no overlap (where both are non-zero), and the more dimensions you have the more likely it is that two random sparse embedding will have little in common. Embeddings are used in many types of artificial neural networks, not just LLMs, for example face recognition networks, where they are trained such that similar faces (multiple photos of the same person) are close together in the embedding space, and post-training you can then "look up" any arbitrary photo (in the training set or not) by embedding it and seeing what is nearby in the embedding space, which will be similar looking faces. Presumably real neural networks used embeddings in a similar way, since, for example, it obviously requires many neurons to represent the many differences between different faces, and there is going to be overlap between the neurons used to represent multiple faces (this is not a computer with one storage location for face #1, and a different location for face #2). A neural network, real or artificial, uses groups of neurons (e.g. a cortical column) to represent an embedding space, with each neuron corresponding to a dimension. A single group of neurons (column) can store multiple embeddings (e.g. faces) represented as different activity patterns (which neurons are firing), and if these are sparse embeddings then the firing patterns corresponding to different memories stored in the same column will have little in common. Now, I don't know how you believe associative recall is implemented in the brain - how does someone's voice, or half obscured face, recall their entire face, so feel free to imagine it as implemented however you will, but I'd suggest that in an assembly such as a cortical column that when a set of synaptic inputs are triggered the assembly as a whole will learn to reactivate the entire pattern when only part of the original set of synaptic inputs are triggered, and this is the basis of associative recall. There are papers that suggest exactly how this may work given the cortical column microcircuit. So, with all that said, the suggestion I was making for why (or at least one reason why) memory degrades with age, with memories blending together, is that with a finite quantity of "storage" (cortical columns) you will eventually be storing so many memories (absent a deliberate forgetting mechanism) that there will inevitably be overlap between the sparse embedddings, and this associative recall will therefore not cleanly recall individual memories but rather recall blended memories according to what they have in common. Obviously some types of memory are at least initially stored in the hippocampus, so no reason to focus on cortical columns, but I expect the use of embeddings is universal.
- randomImmigrant 27d agoOk great, thanks. Now you’ve brought up facial recognition, and that’s actually a great example to show where the analogy breaks, and the idea of a few sparse cells encoding specific faces has been conclusively disproved: https://authors.library.caltech.edu/records/znzhp-4j547 https://authors.library.caltech.edu/records/znzhp-4j547 Faces live in a ~50-dimensional continuous space (25 shape axes, 25 appearance axes). They measured about 205 neurons across 2 macaques (human studies have substantiated much of this, some from the same lab), and the key thing is: every neuron participates in every face. The paper shows faces are embedded, but as points in a dense linear space where neurons are axes, not as sparse activity patterns where neurons are on/off slots. The mapping between the neuronal activity and the facial structures is invertible. Record these same cells, and their firing pattern can be used to reconstruct the face. Or, if you generate a novel face, you can predict the firing rates of these neurons for it. As far as I understand, this doesn’t work for sparse embeddings. Some cells carry the shape coordinates and others carry the appearance coordinates, in a heirarchy. There’s an embedding space, yes. But that space isn’t defined by a network of “on” and “off” neurons. The embedding space is instead constructed by the activity of neurons, and the differences in activity distinguish the faces, using the same set of neurons. And distance in the ensemble activity of these neurons tracks the distance in face space. If faces use sparse embeddings, you wouldn’t expect similar faces to evoke similar activity would you? Yet that is exactly what this paper shows, and the same has been shown in the human brain for faces. There are places where it’s sparse activity of a subset of neurons that maps to specific memories. What you’re describing is what you’d see if you look at how the dentate gyrus (part of the hippocampus) handles your memories in the same location. But even there, the sheer number of cells makes this combinatorially such a vastly overdetermined system for a lifetime that there’s no capacity limit of the kind you’re describing. Even 1% of these cells lighting up for a specific memory leaves you with so many possible combinations that you’d have to live for a few million years to be in the right scale to at least being to talk about capacity issues. The brain just isn’t capacity limited by the number of neurons the way your intuition is pointing you. If you say this has nothing to do with the Von Neumann bottleneck or computational functionalism, fine, but how do you square that with the statement below, which you made further down responding to another post? > but it's hard to imagine that all of the classical chemistry, let alone quantum, details are important. It's necessarily built out of chemistry, but selection is happening at the level of behavior - presumably depending only on a much higher level set of abstract capabilities (ability to learn, etc), not the exact details of chemistry. The success of LLMs, a crude prediction mechanism built atop a crude ANN, does tend to support the idea that low level details don't matter. Timing will matter if we want to go beyond LLMs to AI that can learn time-based things and not just sequence order, but how much else will matter remains to be seen! It’s really odd to see these two paragraphs, because the second actually tells you why your first is wrong. Simply put, the biochemistry is timed. I urge you to study how temperature compensation of circadian rhythms is achieved. That anticipatory function goes all the way down to the molecular level. It might go down to the quantum level too. In birds, magnetoception depends on a protein called cryptochrome IV, which uses a singlet born, entangled radical pair of electrons to sense the very weak magnetic field of earth. Now cryptochrome 4 is bird specific and mammals don’t have it. Other cryptochromes are critical clock molecules. And the whole shebang of these evolved initially to be sensitive to blue light and repair DNA. Try as you might, you can’t separate out the deep linkages from the molecular to the behavioral in biology. Trying is perfectly fine for stuff like language models. But if you’re going to build models with internal time, best of luck if you ignore the molecular and the energetic considerations. Time emerges from the ground up, in biology, as in physics. Doubt we’ll get a free ride with computers.