3 ms·
How are you measuring “rightness”? Does it have to exactly match the predefined JSON output? Or is there some qualitative “sameness” check going on? If you’re
by yeldarb 3y ago
How are you measuring “rightness”? Does it have to exactly match the predefined JSON output? Or is there some qualitative “sameness” check going on?
If you’re not already, you could probably give GPT-4 the correct JSON and the predicted JSON and ask it whether the answer is qualitatively correct and assign a letter grade much like a teacher would grade something qualitative like an essay.
The “parse the graph” one seems almost correct. It’s labeled as “Failing” which is correct if failing means “not perfect” but I’d probably give it a B- rather than an F.
- zerojames 3y agoHow we validate each response varies depending on the test, but generally we are looking for exact equality. This is useful because we can measure whether GPT-4 with Vision does or does not, with full accuracy, solve a problem. This is a category of common questions related to LMM performance. The model being "almost correct" is a valuable condition for us to document. I'll note the idea of having an "Almost" state that shows in yellow if the response for a given day is close to, but not exactly, right. There are use cases where this is valuable information. For instance, an OCR response being 90% correct may be tolerable if you can run error checking to reconstruct the response (i.e. the case for ISBNs or serial numbers constructed in a certain way).
- yeldarb 3y agoIn the simplest case, { “a”: 1, “b”: 2 } seems equally as valid as { “b”: 2, “a”: 1 } (also other things like different key names and white space seem like they shouldn’t effect the “rightness”.