4 ms·
In summary they took 24 songs (13 hits and 11 flops) and evaluated these against some neurological parameters from 33 volunteers, measuring average immersion, p
by barbegal 3y ago
In summary they took 24 songs (13 hits and 11 flops) and evaluated these against some neurological parameters from 33 volunteers, measuring average immersion, peak immersion and "Retreat" (low immersion). They then synthesised 10000 observations which were labelled either hit or flop and had a similar distribution of the three parameters as the original 24 songs.
A machine learning algorithm was trained on 5000 of these observations and tested on the other 5000 observations and the 24 songs.
It got 97% of the synthetic observations correctly labelled and 23 out of the 24 songs.
Unfortunately, it can clearly be seen that the generation of the synthetic data based on all 24 songs means overfitting to the data can easily take place (despite what the authors think their 10-fold cross validation proves)
Without proper separation of training and test data this is a garbage study and tells us virtually nothing.