4 ms·
Cool project! Something you might want to keep in mind for future iterations or similar projects - You ended up with 127 usable scans, and then used all 127 sc
by Yen 11y ago
Cool project! Something you might want to keep in mind for future iterations or similar projects -
You ended up with 127 usable scans, and then used all 127 scans as your test data set while developing. You run the risk, here, of 'overfitting' your data set. That is to say, you figure out a set of parameters and techniques that gets a good match rate (98%) on your test data set, but fails when you try it on new data.
There's a useful technique in machine learning and statistics, called Cross Validation. There's different specific techniques, but basically, you should split up your dataset into training (80%ish) and validation (20%ish) sets, randomly sampled. You develop with just the training set, and when you have a good match rate, you test if it also is good on your validation set. This helps you detect if your technique actually generalizes to new input well, or if it just happens to match what you trained on.
- omn1 11y agoThanks for the hint. This is an important note. I tried to be as generic as possible but still I need to test it with more scans.