3 ms·
Are these authors aware of the contents of the training set? My understanding is that they are not. If not, how can they know that the model is not being tested
by carbocation 4y ago
Are these authors aware of the contents of the training set? My understanding is that they are not. If not, how can they know that the model is not being tested on the training set?
In the paper they say that they came up with a "MELD" algorithm to try to detect testing on the training set, but in my view it has the wrong properties to answer this question (from the paper, it has “high precision but unknown recall”).
I don't at all doubt that a language model could perform exceedingly well at this task, but I think that the way to make this paper into a valuable scientific work would be to present the model with questions that had not yet been written as of the end of its training time.
- yawnxyz 4y agoI wonder this every time I work with a physician in Australia (here working on a clinical trial), as compared to a US physician. Seems like physician training here and in Canada is much laxer / easier, but that's as much awareness I have of the physicians' training set.