3 ms·
I think this post suffers from anthropomorphizing LLMs. Humans have values which inform our behavior, these values permeate all our communication and choices (t
by snordgren 3y ago
I think this post suffers from anthropomorphizing LLMs. Humans have values which inform our behavior, these values permeate all our communication and choices (to varying extents).
LLMs are trained to minimize loss on a dataset, and the previous tokens are used to inform which token is most likely to come next. The LLM will reflect the values of its training data, which may be wildly inconsistent.
We can use a system prompt like OpenAI to align LLM output with human values, by simply asking for it. By framing a query with an ethical system prompt we instruct the LLM to pull its output from a different part of its latent space.
This "constitutional AI" is just doing the same, but encoding the system prompt into the training data. As long as the training data concerning ethical behavior is more ethical than the rest of the training data, it is fair to assume that this method will be quite effective at aligning the model.
This is not perpetual motion, because the required information was within the model the whole time. It just wasn't the most likely output with regard to the original dataset.
- DalasNoin 3y agoWhen using system prompts, you still need to finetune the language model to follow the content of the system prompt. I think that you can't quite replace constitutional AI with a constitution-systemprompt even if you do this. In practice, LMs still need chain of thought to determine whether an output is consistent with the constitution.
- noduerme 3y ago> the required information was within the model the whole time. This, being a restatement of the concept of "latent space", though... makes me realize even more how much I hate that concept. Yes, every color or every pixel in a 2D square exists in the enormous multidimensional latent space of a model. But when that space is essentially "all images never seen before, all books never written," then consigning all that to potential model output if you rub the genie the right way is the same as wasting your three wishes on nothing. For your and my latent creative space is just as multidimensional and broad, but has the benefit of intersecting with reality. Intersecting with reality narrows all our potential output down significantly and also increases the chances that we'll fill a blank square with more signal than noise, so therefore, we know that limits are good and beneficial to structured thought... and yet still, there's this crazy obsession with the unbounded latent space. Is it a signal that we've become unmoored from reality, that we'd build a tower to heaven out of actual babble?
- haswell 3y ago> Is it a signal that we've become unmoored from reality, that we'd build a tower to heaven out of actual babble? I think there’s a case to be made that we’ve been unmoored from reality for quite some time now. Recent developments seem closer to finishing that unmooring than starting it.
- twic 3y agoAnd once we are finally unmoored from reality, we can set sail for whatever ultraterrestrial destination our cybernetic helmsmen have in mind.
- gnramires 3y agoThe latent space is useful because it reduces the dimensionality, in some cases to one-hot encoding. Yes, the "latent space" (it's not latent if it's in the input :) ) of a pixel map is of very little information (but not zero[1]). But neural nets can build very informative latent dimensions, which can correspond to abstractions: the network can abstract things like "Is this object a car?", "how red is this object?", "how large is it?", and so on, as latent parameters. Abstraction is one of the fundamental properties of (efficient) thought, that allow dealing with such large amounts of information efficiently by reducing its complexity, reducing it to essential data, that's essentially what latent parameters can be. Indeed I believe several experiments have shown emergence of such latent parameters (which confirms expectations). Also several papers do "latent space interpolation" for generative modeling, which allows mixing of abstract notions (such as gender, size, age, and so on) right in latent space to produce interpolated results (such as images or even text, etc.) that are not pixel-space averages but "conceptual averages" (conceptual interpolation) of some results. This possibility is also evidence of the validity of interpreting the latent space as abstracted data[2]. [1] For example, you can use the pixel space directly when comparing very similar images, say affected by additive noise, using Euclidean distance in pixel space is a pretty good distortion metric for human vision (and you can do classification in some cases as well of say digits using euclidean KNN in pixel space but depending on very large datasets). [2] I think it's important to note that abstraction isn't necessarily quite 'dimensionality reduction', or removal of redundant data. I think abstraction is more like computational simplification, and in some cases you might need even more intermediary simplified data than you began with (consider the memory usage of some efficient algorithms that far exceeds the dataset size). So the abstract classes and dimensions of the data could be even more than the data size itself, but they tend to signal computationally independent characteristics, like say 'is there a car in the image?' and 'is there a person in the image?' -- you could have thousands of such classes each represented in a single dimension in binary encoding.
- EGreg 3y agoWhy not just tell the AI to also create the constitution? And AI all the way down. We should just be able to bootstrap the AI and it learns exactly what we want with zero input.
- powerapple 3y ago"Humans have values which inform our behavior, these values permeate all our communication and choices (to varying extents)." Imaging if you want to "share" your value with others, how do you do it? The prediction from LLM, the choice of words, is influenced by our language, and our value.