3 ms·
How we validate each response varies depending on the test, but generally we are looking for exact equality. This is useful because we can measure whether GPT-4
by zerojames 3y ago
How we validate each response varies depending on the test, but generally we are looking for exact equality. This is useful because we can measure whether GPT-4 with Vision does or does not, with full accuracy, solve a problem. This is a category of common questions related to LMM performance.
The model being "almost correct" is a valuable condition for us to document. I'll note the idea of having an "Almost" state that shows in yellow if the response for a given day is close to, but not exactly, right. There are use cases where this is valuable information. For instance, an OCR response being 90% correct may be tolerable if you can run error checking to reconstruct the response (i.e. the case for ISBNs or serial numbers constructed in a certain way).
- yeldarb 3y agoIn the simplest case, { “a”: 1, “b”: 2 } seems equally as valid as { “b”: 2, “a”: 1 } (also other things like different key names and white space seem like they shouldn’t effect the “rightness”.