3 ms·
We can have perfect insight into the bias of the LLM, but we can't have perfect insight into the bias of a human. An LLM's judgements can be verified as consis
by courseofaction 2y ago
We can have perfect insight into the bias of the LLM, but we can't have perfect insight into the bias of a human.
An LLM's judgements can be verified as consistent with past judgements, or other criteria. A judge's cannot.
The bias can be quantified in advance of making actual rulings.
- czl 2y ago> We can have perfect insight into the bias of the LLM, You will analyze billions of weights the LLM is made from? What will you look for? You will analyze the trillions of text training inputs that created the LLM weights? What will your analysis do? How to get this "perfect insight" that you claim?
- mewpmewp2 2y agoYou would do statistical analysis of the output. E.g. whether it is preferring a certain gender if gender is variable in the prompt, but everything else is the same etc.
- deleted 2y ago[deleted]
- dragonwriter 2y ago> You would do statistical analysis of the output. Assuming reproducible settings are used, that will give you perfect insight into the behavior in the tested circumstances, but unless it lets you reconstruct the entire network, it won't tell you the impact that untested changes even if they are recombinations of tested elements behave. It is not perfect insight into the biases.
- czl 2y agoWhat if the LLM is biased against words that start with e or biased towards sentences having five words? Using your technique how would you discover such biases? If highly paid lawyers know about some such innate LLM quirks might they exploit them to get a biased decision?