3 ms·
> To evaluate Form Recognizer, I split the data randomly into 26 training documents and 25 test documents. Training on just 26 documents seems woefully inadequ
by hermitdev 5y ago
> To evaluate Form Recognizer, I split the data randomly into 26 training documents and 25 test documents.
Training on just 26 documents seems woefully inadequate. I'm not a data scientist and have only cursory exposure to ML, but I'm not surprised to see terrible results with such a small training set.
- ctk_brian 5y agoAgreed, but ground truth labeling is a lot of work! The thing is, Form Recognizer has a hard limit of 500 total pages (not documents) in the training set. I'm skeptical it's possible to achieve good performance with an unsupervised model with only 500 pages, unless those documents are very similar. In which case, why would you need a service like Form Recognizer at all? From a product perspective, it just makes no sense to me.