3 ms·
I think the point of the study was precisely to see how well the original model without adaptations (my understanding is that they only use prompting, which doe
by c7b 3y ago
I think the point of the study was precisely to see how well the original model without adaptations (my understanding is that they only use prompting, which does not affect weights) can perform. I think that question is arguably even more interesting than 'How well can we train a model to perform on med school questions?'. I'm not saying this is generalization, because the training set surely included a lot of medical literature, but if the base model without fine-tuning can perform well in one important domain, that's a very interesting data point (especially if we find the same to hold true for multiple domains).
- ramoz 3y agoCaveat this with the fact that they did rag.
- practice9 3y agoYeah the Medprompt name is misleading
- owl_brawl 3y agoThe "RAG" part (over the training set) is by far the smallest contribution to the performance gains reported (see the ablation study in section 5.2). I don't think the model is actually learning in-context from the selected samples, but rather is continuing to be better conditioned to sample from the right part of the pre-training distribution here, which does a slightly better job when the samples are on topic (vs. random)