3 ms·
A big factor in producing a good analysis is the feedback modality -- chat transcripts are different from emails, which are different from web forms or operator
by Radim 8y ago
A big factor in producing a good analysis is the feedback modality -- chat transcripts are different from emails, which are different from web forms or operator notes.
We've had several "customer feedback / intent / support case analysis" projects in the past. Some for large customers with millions of individual records (Autodesk), where there's the additional challenge of "What should the categories be in the first place? What's in the data?" (discovery).
What we learned is a model trained on one type of feedback will not necessarily perform well on others, because the relevant signals manifest differently across modalities: feedback length / writing style / typos, lexical richness / repetition / boilerplate, OCR noise / how long is the long tail… Your model may learn to pick up on cues that are orthogonal to the sentiment or categorization problem.
This is especially true for black box models (deep learning) where introspection is limited: Did the model learn to rely on syntax? Specific words or character ngrams? Exclamation marks? Something else? Does an Indian-looking name imply sentiment negativity?
Slapping a generic ML technique (Stanford NLP, Naive Bayes, bi-LSTM, whatever) onto a bunch of tokens is a reasonable first step, that's the low-hanging fruit. The tricky part is defining the problem space and the QA process correctly, and managing the devil that comes with the details.
- prabhatjha 8y agoI totally agree with this. We learned it pretty quickly that classification does not generalize across domains so we narrowed the problem space by focusing one domain at a time followed by predefined and fixed set of categories so that we can measure effectiveness of our solution as we experimented with different algorithms and deployment pipeline.