4 ms·
Deep Dive into G-Eval: How LLMs Evaluate Themselves
- eeasss 11mo agoAre there any llms in particular that work best with g-evals?
- kirchoni 11mo agoInteresting overview, though I still wonder how stable G-Eval really is across different model families. Auto-CoT helps with consistency, but I’ve seen drift even between API versions of the same model.
- zlatkov 11mo agoThat's true. Even small API or model version updates can shift evaluation behavior. G-Eval helps reduce that variance, but it doesn’t eliminate it completely. I think long-term stability will probably require some combination of fixed reference models and calibration datasets.
- sirlapogkahn 11mo agoWe’ve tried geval but it hasn’t been super useful in practice. If we run the same input on the same model and same geval 10 times we get significantly different results, so you can’t really arrive at any conclusions based on the results.