4 ms·
This is awesome! I wonder if anyone's done something similar with beers? Anyway a few "next thing to try" suggestions from a machine learning perspective: The
by mjw 13y ago
This is awesome! I wonder if anyone's done something similar with beers?
Anyway a few "next thing to try" suggestions from a machine learning perspective:
The model selection process used here is by its own admission quite ad-hoc, based on a gut feel about diminishing returns. There are various more principled methods you can use to find the sweet spot between over- and under-fitting with these kind of models, a lot of them based on held-out validation data.
One way to do this would be leave-one-out cross validation (LOO-CV): hold out one whisky, fit the model, and see how 'surprised' the model is by the held-out whisky, repeat for the next whisky and average over all the folds. Because the dataset is tiny this should be quite feasible.
To measure 'surprisal' you could e.g. look at the distance from the held-out data point to the nearest cluster, although something better motivated would be if you switched to a probabilistic model and used likelihood of the held-out data. Probably the simplest next thing you could try in that direction would be a Gaussian mixture model (GMM) trained using EM. K-means is actually a degenerate limiting case of this.
A probabilistic model would also allow you to use Bayesian model selection criteria, which can get quite interesting (and might lead you eventually to things like Dirichlet process mixture models).
It would make it easier to compare the model's explanatory power with other unsupervised probabilistic models. For example some kind of latent factor model like Factor analysis or pPCA would be quite interesting to investigate too, whether taken alone or in combination with clustering as a dimensionality reduction step as tlarkworthy is suggesting.
Also concur that doing multiple runs with different randomised initialization is generally a good idea for k-means or EM, since they can get stuck in poor local minima. Perhaps more common practise to pick the best of multiple runs than to average them though.
- _deh 13y agoSomething on beer: http://bit.ly/JLNgZA http://bit.ly/JLNgZA (wisc.edu)