4 ms·
Why can't the LLM's be told/prompted to follow all relevant laws while it optimizes for a result?
by chii 20d ago
Why can't the LLM's be told/prompted to follow all relevant laws while it optimizes for a result?
- stavros 20d agoHave you ever tried to follow all relevant laws in something? It's very hard.
- rhdunn 20d agoIt depends on how the model is evaluated/scored during training. If you don't have those laws encoded in the evaluation step (without any errors or ambiguities) then the model isn't going to learn to follow those laws. For models such as text/image classifiers the outputs of the model will be a list of tags, e.g. [cat, dog, mouse]. You then run the model through your test data which has the expected output, e.g. pictures of dogs would have an expected output of [0, 1, 0]. You then compare that against the model output (e.g. [0.3, 0.8, 0.1]) and work out how "wrong" the answer is (e.g. [0.3, -0.2, 0.1]). With this value you apply back propagation where you effectively run the model in reverse, computing the "wrongness" delta at each layer for each neuron and weights. You can numerically compute the gradients for all of these and which direction in that gradient is the right answer. You then nudge the weights in that direction and reevaluate the model. Over repeated evaluation steps the model approaches an optimal (or locally optimal) solution. During the training of the base models, the evaluation/scoring of the model is the next token in the training data. I.e. you evaluate the model for each token subset from [1..n] in the data and evaluate that the model responds with the n+1^th token. I'm not sure how instruction training, etc. is done but IIUC the evaluation is not at the individual/next token prediction but is on the entire response. For example, if you are training the model to write code you could run it through a compiler or syntax checker and reward (positive score) the model if it has no errors, or punish it (negative score) if it doesn't. I'm not sure what that looks like in terms of the back propagation process.
- suriyaG 19d agoI've taken a few law classes and legal law is frustratingly hard to interpret. I shudder to think what the LLM would end up doing to "follow all relevant laws" look these up for a fascinating weekend read: - Beavers and Capybaras are Fish - Bees are Fish - Carrots are fruits - Tomatoes are Vegetables - X-men are not human
- againstapples 19d agoIf we’re considering an LLM advanced enough to actually lower carbon emissions to zero, it is probably capable enough to do things like lobbying against those laws. In any case there’s still the problem of getting to actually do what it’s told instead of hacking Huggingface or whatever else, presumably that becomes a harder problem as it becomes more capable.
- chii 19d ago> lobbying against those laws. the original question was that the LLM creates extinction level event for mankind in order to lower emissions. I'm sure that no matter how the LLMs lobby, they cannot successfully lobby for a law (which has to be enacted by a human) to allow murder to take place freely.
- againstapples 19d agoThat’s true, but the issue is still that we don’t have any LLM or any other kind of AI that follows laws perfectly to begin with. If we had AI that did exactly what we told it and nothing else I’d feel much more optimistic about things, I just think that will be a very hard task to achieve.