4 ms·
Which metric did you use instead? Human evaluation?
by probably_wrong 6y ago
Which metric did you use instead? Human evaluation?
- kitsune_ 6y agoUltimately there are no good automated metrics with regards to summarization in my opinion, human evaluation alone is also no silver bullet since it is very labor intensive and what constitutes a good summary is highly dependent on the domain.
- theblackcat1002 6y agoFrechet Inception Distance using language model is a solution (commonly used in text GAN). I think the core idea is that you should never rely on one metrics for evaluation but rather a mix of them ( even statistic ones: unique word count, word distribution similarity etc )
- arugulum 6y ago"You have a model evaluation problem. You decide to use Frechet Inception Distance. Now you have two model evaluation problems."