3 ms·
The HumanEval benchmark scores are confusing to me. Why does Haiku (the lowest cost model) have a higher HumanEval score than Sonnet (the middle cost model)? I
by memothon 3y ago
The HumanEval benchmark scores are confusing to me.
Why does Haiku (the lowest cost model) have a higher HumanEval score than Sonnet (the middle cost model)? I'd expect that would be flipped. It gives me the impression that there was leakage of the eval into the training data.