9 ms·
One thing that would be fun to know is, e.g., "If the LLM answers question X correctly, then what's a minimal-sized set of things we could remove from the train
by keithwinstein 3y ago
One thing that would be fun to know is, e.g., "If the LLM answers question X correctly, then what's a minimal-sized set of things we could remove from the training set and cause it to get that question wrong?" I think with current methods this would be pretty expensive to find out, but, in principle I'm guessing it would be pretty illuminating.
- hansvm 3y agoCausal modeling lets you determine where particular facts are stored without too much computational cost and also how you can edit those facts in the model. That might allow you to decouple the problem into first finding the desired set of gradients (converging at the target modification, zero for the weights you don't care about), and just do a linear solve (since probabilistic weightings on how likely an input is in the training set will linearly impact each of the gradients in basically every neural architecture of note, including most LLM modifications) to find an approximately minimal (IIRC, a true minimal subset is NP-hard or something) subset of the inputs which when combined would give you the inverse of the target gradient. Remove that much of the weighting for each of the inputs.