3 ms·
> - How significant are the estimates of lifestyle factors? Do you have p-values? If you bootstrap resample, how much do the rankings change at the extremes? H
by aab0 10y ago
> - How significant are the estimates of lifestyle factors? Do you have p-values? If you bootstrap resample, how much do the rankings change at the extremes?
His VW script does do bootstrapping ('--bootstrap 16') but he doesn't report it anywhere I see. https://github.com/JohnLangford/vowpal_wabbit/wiki/using-vw-varinfo https://github.com/JohnLangford/vowpal_wabbit/wiki/using-vw-... seems to not report any sort of p-value or confidence interval which might be derived from the bootstrapping. (The 'relevance' is 'the relative distance of each variable from the best constant prediction', not sure what that means.)
So if you want to know, it looks like you'll have to run it yourself and visualize the output. I would guess that the uncertainties are huge and none of them reach even p<0.05 - it simply should not be possible to get mean loss of like 0.2 and reliable estimates of hundreds of variables like 'melon' out of less than 4 months of data when the random measurement error of the scale itself is on the order of half a pound (I have an Omron body fat scale, and even taking 2-3 measurements daily, there's a lot of error) unless his VW regression is grossly overfitting.
- benkuhn 10y agoI mean, the uncertainties in the middle are less interesting than the uncertainties at the extremes. I'm sure you can't say anything useful about melon, but it would be interesting to know if (e.g.) "sleep" was consistently that big. He did regularize somewhat (`--l2 1.85201e-08`) but unclear whether that's enough regularization. Basically I'd love to see some actual diagnostics I guess :) PS I believe the "relevance" is just the coefficient divided by the biggest coefficient (the `RelScore` column in the printout in the README file).
- aab0 10y agoWell, he also disables the holdout set (`--holdout_off ariel.train`), so between that and not reporting the uncertainties and his absolutely tiny dataset with high measurement error and the low prior probability that any of these effects could be that big... Nah. It's just overfitting noise. I tried to run the makefile, but apparently the version of Vowpal Rabbit that ships on my Ubuntu (7.3) is so outdated that it doesn't support the bootstrap option.
- ariel-faigon 10y agoThanks so much for all the excellent comments. There was definitely an over-fit with 4-passes. No more. I've updated the Makefile to run only one pass, changed the options so it runs with older-version vw, Fixed misspellings of 'gioza', removed 'mayo' which found itself on the wrong side because it appeared only twice and always alongside the bun and regenerated the chart. All the main conclusions remain intact. In the end, I urge everyone to use their own data, that was the main purpose of sharing this code. My data-set is small, awfully noisy and insufficient. There are no p-values and no rigorous statistics, so please don't read too much into the minute details. It is the discovery journey into the top factors that is the important part, in my view. The ML was just one aid in this discovery process. The proof for me was my actual, and sustainable, weight loss that came after (very slowly) realizing the top factors that eventually worked for me. Thanks again.
- aab0 10y agoI don't think it matters whether you run 4 passes or 1 pass, it's still going to overfit. You can run an online linear regression in a single pass too, but that doesn't magick away the uncertainties. The results are still going to be garbage, and any effects you get are due to your health-consciousness and not any specific dietary choices you make (how could it be, when the data is so weak and noisy that each item can easily flip signs?).
- ariel-faigon 10y agoThanks so much. Your comments are really helpful. I realized early on that the data is hopelessly noisy, due to the small daily changes and the scales resolution so rather than trying to build a perfect model to gauge the variable importance of each and every kind of food, I focused on the few days when weight change was more significant hoping I could detect some signal in those, and extrapolate and further explore from that. That's why I sorted the data-set by abs(delta) and that's what consistently pointed me towards sleep/fasting as the #1 factor. I do agree that the full list/model is garbage in the sense that probably 80% or so of it is woefully inaccurate/flipped, noisy, overfitted etc. The main point was to lead me in the right direction by looking at the big picture and what stood out. And what stood out were 2 things 1) sleep (fasting duration), and 2) fat vs carbs. I think everything else should be ignored. I think we're in total agreement on this point. Does this sound more sensible to you?