4 ms·
I'm curious, how do you prevent overfitting, where the model will simply learn the exact formats of the training data? Then it will not generalize to a format
by 2bitencryption 6y ago
I'm curious, how do you prevent overfitting, where the model will simply learn the exact formats of the training data? Then it will not generalize to a format it has never seen before?
Unless the training data is an extremely diverse set of invoices, maybe randomly generated?
- malux85 6y agoLook at the loss and val_loss numbers in the sample image, it's way overfitting.
- ottolin 6y agoThat's exactly the first thing pop up to my mind also.
- visarga 6y agoIn general I think invoice extraction models will only generalise up to about 90% F1, after which they plateau off. What you can do at this point is to include examples of the specific formats your clients need to process and overfit on those formats, making the system 95% or 97% accurate. But for new layouts it's only going to be around 90%. The problem with invoices is that some fields have extreme variability - the address, the company names and the product descriptions. So a synthetic invoice generation approach might not work when you want to process in a new industry or language.