4 ms·
Illusory generalizability of clinical prediction models
- lamename 3y agoI'm glad this is getting attention. If you spend time reading this literature, especially in the mental health space, there's a lot of sloppy ML published. Unbelievably low sample sizes, little to no mention of careful cross-validation or other measures taken to prevent data leakage, etc. In one case, a supervised ML model built for diagnosis using binary classification, with no training examples in the negative class, touting great true positive rates in the abstract. I suspect part of the problem is people coming from other disciplines (ab)using machine learning who haven't been taught how to use it properly. However, even when the authors include a biostatistician/mathematician in the research, that doesn't mean the paper will be of high quality.