5 ms·
First Impressions with Google Gemini
- yeldarb 3y agoWeird that the object detection prompt refused to answer. What other variations of that have we tried? Do you think that's an intentional task that was RLHF'd out or a quirk of some kind? Thinking out loud some things I'd try: Find the x/y position of the dog. What is the center point of the dog? Give the pixel coordinates (xywh) of the dog in this photo. Simulate an object detection model trained to find dogs run on this photo; give your output in JSON format.
- zerojames 3y agoI have experienced intermittent performance with the web interface. I am really curious about the object detection performance so I’ll be rerunning that test regularly.
- zerojames 3y agoAfter some prompting, I got a response. It identified the dog in the image, but the dog was in the center point. I gave Gemini the Home Alone image and it couldn't identify the Christmas tree, returning invalid coordinates. The post has been updated accordingly.
- deleted 3y ago[deleted]
- s0laster 3y agoJust guessing but it might be some kind of safety such that Gemini can't be used to solve captchas.
- YetAnotherNick 3y agoYou gave very different prompts to gemini and GPT 4("how much money" vs "how many coins", different coordinates systems etc).
- zerojames 3y agoThank you! We took screenshots of our tests, which you can download here: https://media.roboflow.com/lmms-tests.zip https://media.roboflow.com/lmms-tests.zip (Gemini ones are not in there yet.)
- uxp8u61q 3y agoThat's not really answering anything.
- YetAnotherNick 3y agoThe prompt for the coordinates is still different in GPT 4. Coins one look to be the same but the question is little bit ambiguous as GPT might be answering the total value of coins and did 1+1+1+2. For me, the tax one is working in GPT 4.
- mvdtnz 3y ago> Coins one look to be the same but the question is little bit ambiguous as GPT might be answering the total value of coin Well that wouldn't be a very intelligent "AI" would it? The prompt was completely unambiguous.
- mewpmewp2 3y agoTrying myself with GPT 4 Vision (directly using the API), I get response: "The image shows a total of four coins." For the exact same image with a prompt "how many coins do I have?" I keep getting the same thing every time. If I ask "what is the value shown on the image?" I get the following: "The image shows four coins, and from what I can discern, three of the coins have the number "1" on them, which likely indicates that they are one-unit coins in their respective currency. The fourth coin, which is a different color and appears to be bi-metallic, has the number "2" on it, suggesting it is a two-unit coin. The total value shown by the coins, therefore, appears to be 5 units of their currency. However, without knowing the specific currency, I cannot determine the exact monetary value in any other terms." So to me it's obvious gpt4-vision model can handle this question. Edit: I also get for regular ChatGPT "You have a total of four coins in the image." Now I tried to ask GPT4 Vision to find the Dog, and it said: "The dog is located approximately at 0.29, 0.36, 0.70, 0.90." - Just eyeing it, it may seem decently accurate, but not 100%. My difference is that I took a screenshot from your pictures rather than providing the original though, so that could also probably affect answers. Prompt for the dog was: "Find the dog. Return its location in the format x_min, y_min, x_max, y_max. Respond with 0-1 as percentages."
- volandovengo 3y agoThanks for the thorough analysis. Wish articles like this were not written optimized for SEO. They’re much harder to read!
- zerojames 3y agoReadability is important to me. What can I do better? I always appreciate feedback.
- serjester 3y agoObviously the core content wasn't but the article reads like unfiltered AI output. It doesn't really 'flow' for a lack of a better word.
- blindstitch 3y agoIt is an interesting article and some of the results are unexpected, but the layout is way too long and could be greatly condensed. At first glance and speaking from how I like to lay things out in a paper: - You do not really need to show the prompt interface. You can show it once and thereafter use a bulleted list format, or simply show the input image if it responded correctly. - Your figures should be about half of their size, they don't need to fit the width of the body. - For comparative results against other models you can use a table with colored cells, with the test name on the row and model names on the columns. - For the dog, show a side-by-side figure with the raw image on the left and the box on the right, and include the coordinates it gave you in the body. - In your conclusion show the full matrix table of comparative results and summarize the relative strengths of the model against the others. In terms of the writing and methods your conclusion says little and your tests do not go into significant depth. For example with the tire image you could show that it succeeds when cropped but as the photo gets wider it begins to fail to correctly identify the text in the image's center. For example see the methodology and presentation this article used: https://dynomight.net/ducks/ https://dynomight.net/ducks/ Also, the OCR test is too simple, even a 20-year-old OCR algorithm would probably recognize that. Experimenting with progressive degradation of the image could show its strengths, and analysis could show its accuracy at each level of degradation.
- 3y ago
- starwin1159 3y agogood work!
- mips_r4300i 3y agoInteresting why it decided to draw a bounding box around the dog's head, maybe because its training images are mostly dog portraits
- mewpmewp2 3y agoI would like to see a few more tests to make sure it's definitely not a coincidence that it was specifically a box in the center.
- Racing0461 3y agoInteresting results since Gemini scored "better" than GPT4 but Gemini is more like a 3.5 equilivalent model. The ultra one being released later is like gpt4.
- abi 3y agoThese results say that Gemini performs better than GPT4-V, which is clearly not true from experience. Clearly, we need better evals.
- hqmhqm 3y agoI am trying to use the HTTP API. I asked this simple question "How many fingers does a baby have?" and all I get is the empty reponse: [{'candidates': [{'content': {'role': 'model'}, 'finishReason': 'OTHER'}]}, {'candidates': [{'content': {'role': 'model'}, 'finishReason': 'OTHER'}], 'usageMetadata': {'promptTokenCount': 3, 'totalTokenCount': 3}}] If I use their example "Give me a recipe for banana bread.", it returns a recipe. Why won't it answer a simple question about number of fingers?