4 ms·
If the dataset mostly contains edge cases, model performance on the dataset is going to be poor, but it's not an issue I think. But how could the real-world ac
by dest 5y ago
If the dataset mostly contains edge cases, model performance on the dataset is going to be poor, but it's not an issue I think.
But how could the real-world accuracy be computed? Is a separate dataset needed for that purpose?
- isusmelj 5y agoWhen learning ML at university one assumes that the data you have well represents the environment. We do the famous train/validation/test split and train our model. However, in practice we see that it is very hard to collect a good dataset. There is a great twitter thread from Abubakar(CEO Gradio) about this topic: https://twitter.com/abidlabs/status/1423067498862219267 https://twitter.com/abidlabs/status/1423067498862219267