2 ms·
I think in addition to all the benchmarks used right now for LLM evaluation (HumanEval and the like). It would be interesting to have a 'hallucination benchmark
by TastyLamps 3y ago
I think in addition to all the benchmarks used right now for LLM evaluation (HumanEval and the like). It would be interesting to have a 'hallucination benchmark' with a summarization based hallucination dataset.