5 ms·
Also, their classifier just uses manual features rather than doing any sort of meaningful analysis. As in, the input to their model is a tuple of ~20 features c
by nsaftarli 3y ago
Also, their classifier just uses manual features rather than doing any sort of meaningful analysis. As in, the input to their model is a tuple of ~20 features consisting of things like, whether the article contains the word "but", or the character "?".
All of this is easily fixable by using better prompts. It's 99% effective for a small set of articles from one journal (sampling not specified), assuming that the synthetic samples are created with the minimum level of effort using GPT-3.5.
It's not that their methods are wrong, it's just that they're absolutely useless for anything beyond an extremely narrow range of data. In the ML field this is called overfitting.