3 ms·
Show HN: GPT-4 with Vision Checkup
When new Large Multimodal Models (LMMs) are released, there is excitement as we explore new capabilities. What can a model do? What can't a model do? What strange behaviors does the model exhibit?
With that said, such analyses are frozen in time.
At a hackathon toward the end of last year, the Roboflow team made a tool that runs the same set of tests with the GPT-4 with Vision API every day. This allows people to see how the model performs over time as updates are made.
The last seven days of results are displayed on a web page; the rest of the data is archived in GitHub.
We started the site with common vision tasks like OCR, object detection, and object counting. We welcome anyone to submit a PR to add new tasks, too!
- yeldarb 3y agoHow are you measuring “rightness”? Does it have to exactly match the predefined JSON output? Or is there some qualitative “sameness” check going on? If you’re not already, you could probably give GPT-4 the correct JSON and the predicted JSON and ask it whether the answer is qualitatively correct and assign a letter grade much like a teacher would grade something qualitative like an essay. The “parse the graph” one seems almost correct. It’s labeled as “Failing” which is correct if failing means “not perfect” but I’d probably give it a B- rather than an F.
- zerojames 3y agoHow we validate each response varies depending on the test, but generally we are looking for exact equality. This is useful because we can measure whether GPT-4 with Vision does or does not, with full accuracy, solve a problem. This is a category of common questions related to LMM performance. The model being "almost correct" is a valuable condition for us to document. I'll note the idea of having an "Almost" state that shows in yellow if the response for a given day is close to, but not exactly, right. There are use cases where this is valuable information. For instance, an OCR response being 90% correct may be tolerable if you can run error checking to reconstruct the response (i.e. the case for ISBNs or serial numbers constructed in a certain way).
- yeldarb 3y agoIn the simplest case, { “a”: 1, “b”: 2 } seems equally as valid as { “b”: 2, “a”: 1 } (also other things like different key names and white space seem like they shouldn’t effect the “rightness”.