4 ms·
I'd just as easily believe that someone updated the model, it overfit to the input, and so minor changes gave incorrect answers.
by staticassertion 5y ago
I'd just as easily believe that someone updated the model, it overfit to the input, and so minor changes gave incorrect answers.
- mannykannot 5y agoIt is, of course, possible that, just by chance, someone happened to update the model in exactly the way that would change the model's response, to one specific question, from a generic evasion to a specific correct answer, yet overfitting so that very similar questions still get the generic evasion. It is even possible that, by chance, this happened within a day after Garry submitting the particular phrasing of the question for which the updated model does give a correct answer. That it is possible is not enough to satisfy my curiosity, however - but then, I don't rank it as being at least as likely as any other scenario. On the other hand, if your scenario involves someone updating the model (or, more likely, InstructGPT) in response to Garry's question, with the intent of having GPT3 return the correct answer to that question, I am not seeing how that would be materially different, in any relevant sense, from hard-coding the answer.