49 ms·
First Impressions with GPT-4V(ision)
- yeldarb 3y agoI’m intrigued to see what kind of problems it’s going to be good/bad at. I think it’s going to be tricky to evaluate though because it has probably memorized all the easy images to eval it with. Eg anything pulled from Google Images (like that Pulp Fiction frame or city skyline photo) is not a good test. It recognizes common shots but if you pull a screenshot from Google Maps or a random screen cap from the movie it doesn’t do as well. I tried having it play Geoguessr via screenshots & it wasn’t good at it.
- loupol 3y agoI wonder how many images from Street View it has been trained on. I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.
- npinsker 3y agoIt's been done recently! It's a bit better than (but competitive with) top players. https://www.youtube.com/watch?v=ts5lPDV--cU https://www.youtube.com/watch?v=ts5lPDV--cU
- inductive_magic 3y ago> I would assume training an LLM to do the same would definitely be doable. I wouldn't be so sure. The reasoning process of Geoguessr pros is symbolic, not statistical inference. /edit: as other commenters pointed out, something similar was done. While this wasn't an LLM, it was a deep learning model, so not symbolic -> https://www.theregister.com/2023/07/15/pigeon_model_geolocation/ https://www.theregister.com/2023/07/15/pigeon_model_geolocat...
- fellerts 3y agoYep, some CS/AI grads from Stanford trained an AI on loads of Street View images and built a bot that is able to beat some of the best Geoguessr players: https://www.youtube.com/watch?v=ts5lPDV--cU https://www.youtube.com/watch?v=ts5lPDV--cU
- bayesianbot 3y agoIIRC it wasn't that impressive in the end as instead of recognizing the places the AI apparently learnt to recognize subtle differences in street view cameras used in different locations? I might be wrong / thinking of the wrong model l and I'm on mobile without my browsing history so hard to check, but I think it was putting a lot of weight on some pixels that are noisy
- zx_q12 3y agoTop geoguessr players use this technique as well. IIRC rainbolt mentioned that there is a section of a country where the street view camera has a small blemish from a raindrop on the camera so you can instantly tell where you are if you notice that.
- thewataccount 3y agoFrom my understanding many of the best players immediately look down to tell what "generation streetview car" they're using, and seem to know what continents/times they're from.
- skazazes 3y agoIt seems it will still be limited by its linguistic understanding of the surrounding context, at least in the first chicken sandwich picture. Although its interpretation could make some sense but is also mostly wrong if talking about physical size of a modern GPU's main processor compared to the size of the associated VRAM chips. It has missed the joke entirely as far as I am aware. I think the joke is actual about Nvidia's handling of product segmentation, selling massive processors with less memory than is reasonable to pair them with on their consumer gaming offerings, while loading up the nearly identical chips with more memory for scientific and compute applications...
- Melatonic 3y agoIronically the exact processors need to run GPT-4V in the first place.....
- jayniz 3y agoThis looks like a Schnitzel to me, not like fried chicken.
- waynesonfire 3y agohas the turd polishing already started?
- Usu 3y agoI'd be interested in knowing how good it is at solving visual captchas, do we foresee a huge rise in automated bypasses?
- GaggiX 3y agoSolving CAPTCHAs at the moment is more inexpensive using humans than using GPT-4 API.
- yeldarb 3y agoIf true, this is wild. I suppose a human could spend 10 seconds per Captcha, so they could do 360 per hour. Add some overhead for not being operating at peak performance every minute of every hour & call it 250. Let's say you can hire someone for $2, that works out to a bit over a penny per Captcha. I don't think OpenAI has published pricing for GPT-4 Vision yet, but if we assume it's on par with GPT-4, and uses only 1000 of the 8000 possible tokens to process an image that's 3 cents per Captcha. Doesn't seem completely unreasonable that at-scale humans may actually be cheaper than LLMs at this point. My mind is a little blown.
- Andoryuuta 3y agoYou'd be surprised, or perhaps horrified, by how cheap (self-proclaimed) human-based captcha solving services are. If you just search for "captcha solving service" the first few results that come up offer 1000 solves of text-based captchas for <= $1 USD, (puzzle / JS browser challenge captchas are charged much higher). Whether these are actually human based, or just impressive OCR services, it seems like they are still much more cost effective than GPT-4 is for now.
- altcognito 3y agoI imagine they are a mix.
- eiiot 3y agoThe way these work is usually presenting an existing captcha to another human who doesn’t even know they’re solving the captcha. For example, sites hosting pirated content serve fake captchas as a way to make money.
- cs702 3y agoSure, there are a few edge-case failures and mistakes here and there, but I can't help but be in awe. AWE. Let me state the obvious, in case anyone here isn't clear about the implications: If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your dishwasher, your home, your office, etc. UIs to many apps, services, and devices -- and many apps themselves -- will be replaced by an AI that does what you want when you want it. A lot of people don't want this to happen -- it is kind of scary -- but to me it looks inevitable. Also inevitable in my view is that eventually we'll give these AI models robotic bodies (think: "computer, make me my favorite breakfast"). We live in interesting times. -- EDITS: Changed "every single thing" to "almost every thing," and elaborated on the original comment to convey my thoughts more accurately.
- apexalpha 3y agoCorrect, this will be the successor to the GUI.
- tmalsburg2 3y agoI doubt it. It’s too damn costly computationally.
- Difwif 3y agoThis is the same reply to GUIs will never take off but decades later and on to the next successor.
- thelittleone 3y agoAgree and the next big step may well be human computer interface. Speech is starting point for input. At some point output will change also and if think it out longer term perhaps a future where instead of reading information we install knowledge, including the stored memory of actual experience. If I want to do pottery, I could think this, download the experience and then be competent at it.
- famouswaffles 3y agoGraph analysis is impressive (last example) - https://imgur.com/a/iOYTmt0 https://imgur.com/a/iOYTmt0 Can do UI to frontend. Seems to understand the UI graphical elements and layout, not just text https://twitter.com/skirano/status/1706823089487491469 https://twitter.com/skirano/status/1706823089487491469 Can describe comic images accurately, panel by panel - https://twitter.com/ComicSociety/status/1698694653845848544?t=m9W8JdSxSi9o1uJL_zx24g&s=19 https://twitter.com/ComicSociety/status/1698694653845848544?... Lots of examples here also - https://www.reddit.com/r/ChatGPT/comments/16sdac1/i_just_got_the_chatgpt_image_recognition_feature/ https://www.reddit.com/r/ChatGPT/comments/16sdac1/i_just_got... It's Computer Vision on Steroids basically. Multi-modality is pretty low hanging fruit so i'm glad we're finally getting started on that. Imagine if GPT-4 could manipulate sound and images even half as well as it could manipulate text. We still don't have a large scale multi-modal model trained from scratch so a lot of possible synergistic effects are still unknown.
- dottjt 3y agoOh wow, I'm completely fucked as a front end developer.
- Tostino 3y agoYour job will change in fundamental ways at least.
- zarzavat 3y agoJob will be okay. Career is over. Maybe we should join the writers on the picket line?
- yieldcrv 3y agoThe more people say that, the less convincing it is There is no way I would have a UI developer onboarded when I can generate many iterations of layouts in midjourney, copy them into chatgpt4 and get code in NextJS with Typescript instantly non devs will have trouble doing this or thinking of the prompts to ask, but the dev team asking for headcount simply wont ask for headcount, and the engineering manager is going to find the frontend only dev redundant
- orbital-decay 3y ago> The bounding box coordinates returned by GPT-4V did not match the position of the dog. I suppose it just doesn't take image dimensions into consideration, and needs to be provided with max dimensions, or prompted to give percentages or other absolute values instead of pixels.
- abledon 3y agohttps://twitter.com/cto_junior/status/1706289820702490839 https://twitter.com/cto_junior/status/1706289820702490839
- greatpostman 3y agoIm shocked at how good this is. The world is truly going to change
- deleted 3y ago[deleted]
- matsemann 3y agoNone of the images loads for me, but works through cache: http://webcache.googleusercontent.com/search?q=cache:https://blog.roboflow.com/gpt-4-vision/ http://webcache.googleusercontent.com/search?q=cache:https:/...
- yeldarb 3y agoLooks like the (Ghost?) image CDN got hugged to death. We'll update the URLs. ``` 403. That’s an error. Your client does not have permission to get URL ... from this server. (Client IP address: ...) Rate-limit exceeded That’s all we know. ```
- deleted 3y ago[deleted]
- zerojames 3y agoThis is now fixed. We have moved the images through to our website. Thank you for the report!
- jihadjihad 3y ago> With that said, GPT-4V did make a mistake. The model said the fried chicken was labeled “NVIDIA BURGER” instead of “GPU”. Any midwesterner could tell you that CLEARLY it's a tenderloin :) https://www.seriouseats.com/best-breaded-pork-tenderloin-sandwiches-midwest https://www.seriouseats.com/best-breaded-pork-tenderloin-san...
- qingcharles 3y agoLOL. They have to save the Midwesterner add-on for v2.
- aidenn0 3y agoI'm going to object to the "any midwesterner" since that's not even a thing in all of Indiana, and the linked article says it's not a thing in Chicago.
- wokwokwok 3y agoI’m impressed, technically, but this seems niche. Who holds their phone up and takes a photo then wants to know it was a photo of? That’s weird. If you don’t know what it is, wtf did you take photo? The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine… Great for the vision impaired. …not sure, what anyone else will use this for. The only really compelling use case is the “code this ui for me”, but as we’ve seen, repeatedly, this kind of code generation only works for trivial meaningless examples. Seems fun, but I doubt I’d use it. (Which, and this is my point, is a massive step away from the current everyday usefulness of chatgpt)
- otoburb 3y ago>>Who holds their phone up and takes a photo then wants to know it was a photo of? >>Great for the vision impaired. Yes, this is great for the estimated 285 million vision impaired people around the world[1]. [1] https://www.bemyeyes.com/about https://www.bemyeyes.com/about
- wokwokwok 3y agoDid you read my comment? I literally said that it’s for vision impaired. That’s great. …but it’s niche. I’m sitting on my couch right now and I can think of like 20 things I could chat to chatgpt about. I can see literally nothing in my visual range want to take a photo of and run image analysis over. It’s like Shazam. Yes, it’s useful, but, most of the time, I don’t need it. I would argue this is true for this, for most people, including the significant proportion of people with minor visual impairments (that would, you know, put their glasses on instead).
- Philpax 3y agobaffling that you think 3.5% of the world's population is a niche
- bastawhiz 3y agoThere's enough vision-impaired people in the world to equal the population of Japan, Korea, and Vietnam combined. And beyond those people who would get obvious utility, this is essentially Google Lens on steroids—I simply can't figure how you could call this "niche". Maybe you won't use it multiple times per day, but plenty of people will. Hell, just now I was wondering why the leaves on one of my plants are starting to brown and could have used this.
- pjmlp 3y agoNo images being loaded on FF.
- mbb70 3y agoThe "Why is this image funny?" test reminds me of https://karpathy.github.io/2012/10/22/state-of-computer-vision/ https://karpathy.github.io/2012/10/22/state-of-computer-visi... In 10 years we went from "SoTA is so far from achieving this I don't even know where to start" to "That'll be $0.0004 per token and have a nice day"
- jihadjihad 3y agoHas anyone tried GPT-4V on that image?
- justlikeyou 3y agoNote: I had to ask it why people in the photo are laughing. In the image, Barack Obama, the former U.S. President, seems to be playfully posing as if he's trying to add weight while another official, who appears to be former UK Prime Minister David Cameron, is standing on a scale. Obama's gesture, where he's putting his foot forward as though trying to press down on the scale, suggests a playful attempt to make Cameron appear heavier. The lightheartedness of such a playful gesture, especially in the context of world leaders typically engaged in serious discussions, is a break from formality, which is likely why others in the vicinity are laughing. The scene captures a candid, informal moment amidst what might have been a formal setting or meeting.
- jihadjihad 3y agoPretty damn good. According to Wikimedia [0]: "President Barack Obama jokingly puts his toe on the scale as Trip Director Marvin Nicholson, unaware to the President's action, weighs himself as the presidential entourage passed through the volleyball locker room at the University of Texas in Austin, Texas, Aug. 9, 2010. (Official White House Photo by Pete Souza)" 0: https://commons.wikimedia.org/wiki/File:White_House_Trip_Director_Marvin_Nicholson_stopped_to_weigh_himself_on_a_scale,_August_2010_(2).jpg https://commons.wikimedia.org/wiki/File:White_House_Trip_Dir...
- deleted 3y ago
- chankstein38 3y agoI want this. I'm a paying GPT-4 customer I hate how these rollouts go. Why do I pay just to watch everyone else play with the new toys?
- mediaman 3y agoYou'll have it within a week or so. Pretty much all new products that require significant per-user incremental workloads (e.g., in this case, significant GPU consumption per incremental user) do rollouts. It's an engineering necessity. If they could roll it out to everyone at once, they would.
- layer8 3y agoThe discrepancy between the two answers regarding the set of coins is jarring. From the answer to the first question, one would assume that it can’t tell the currency. The answer to the second question shows that it actually can. The fact that LLMs don’t reflect a consistent inner model in that way, and hence the users’ inability to adequately reason about their AI interlocutor, is currently a severe usability issue.
- zwily 3y agoI’ve gotten in the habit of asking chatgpt “are you sure?” So many times it will (correctly) correct itself, state that items are hallucinations, etc. It always makes me laugh.
- famouswaffles 3y ago>The fact that LLMs don’t reflect a consistent inner model in that way You're probably not going to ask any human a question about an image and get every single detail you want every time. If you care about a detail, just ask about it. Doesn't really have anything to do with a consistent inner model.
- Workaccount2 3y agoI've heard that it is because the AI outputs what it is thinking as it is thinking it. It doesn't really reflect, it sort of does the equivalent of just verbal thought streaming right onto the screen. So when you ask it to reflect on what it said, that's when it actually looks at it and reflects on it.
- zerojames 3y agoAll: We got the HN hug on our images. This has now been fixed.
- deleted 3y ago[deleted]
- steveads 3y agoI use ChatGPT fairly frequently, but am frequently annoyed at how hesitant it is to answer any question that is somewhat subjective. From it's Pulp Fiction response: "...However, whether you personally consider "Pulp Fiction" a good movie depends on your taste in film." To avoid this noise I often add something like this to my queries "skip any preamble or qualifications about how x is subjective" and the results are much better.
- Loeffelmann 3y agoMy prompt to make ChatGPT useable: Please always give a direct answer, without any additional explanations, disclaimers, expertise limitations, or guidelines on human interaction. Be terse. Do not offer unprompted advice or clarifications. Remain neutral on all topics. Never apologize.
- syndacks 3y agoAnd never, ever recommend me a movie with Tom Hanks in it.
- replwoacause 3y agoThanks I will try that out
- extr 3y agoOne thing I found incredible recently: The advanced data analysis version of GPT-4 can apparently do OCR by writing code to run pytesseract and then interpreting the result. I was seeing what would happen if I gave it an image and found it surprising that it actually worked.
- stackedinserter 3y agoAm I the only one who never saw these extra capabilities? OpenAI keep spreading news about "multimodal chatgpt", video, voice etc but to me it still looks like the initial version, but for $20/mo. It's still a text bot that gives unreliable answers.
- abrichr 3y agoI have the same experience. In the "Beta features" settings dialog, I've enabled "Advanced data analysis": > Try a version of ChatGPT that knows how to write and execute python code, and can work with file uploads. Try asking for help with data analysis, image conversions, or editing a code file. Note: files will not persist beyond a single session. Transcript: ``` In your settings it says: > Try a version of ChatGPT that knows how to write and execute python code, and can work with file uploads. Try asking for help with data analysis, image conversions, or editing a code file. Note: files will not persist beyond a single session. How do I upload an image? ChatGPT I apologize for the confusion. The version of ChatGPT you're currently interacting with doesn't have the capability to accept file uploads. My apologies for any inconvenience this may have caused. If you have any other questions or need guidance on how to convert an image using software or online tools, feel free to ask. ``` Hopefully it's just a matter of time, but either way it's jarring for their product to contradict itself.
- continuitylimit 3y agoSo a jumble of chair legs is “NVIDIA burger” and it did say GPU was a “bun” so it thinks the flat thing (chicken?) is some sort of bread. If GPT-4V was “aware”, it would say “it’s funny because I won’t get it right but you will use it get a bunch of $VC, and that is funny, kinda”.
- deleted 3y ago[deleted]
- m3kw9 3y agoI’m just imagining a mode where OpenAI calls it “App Mode” where you say what you want say “a dog themed cute calculator app with units conversions”, and it will generate the UI for a working app. You add these into a widget like place. The OpenAI AppStore will carry these apps. Although in the beginning the apps would be simple but I do see potential
- Reflecticon 3y agoThe more AI can produce customized stuff for us the less we need companies. Full personalization of our products might be possible. Probably first software, then art, then 3D printed products and maybe later houses, cars and clothes. I wonder what we will work and if we will work at all in such an environment. Maybe some people still like consuming and copy different designs and products and because of the Blockchain you have to give them something in exchange or everything is open source and it is free for you to take. I wonder whether such life would contribute to humanity making further progress or make it stagnate (or possibly decline)? Interesting times. I think we are close to the times of the moon landing. Which had an immense Impact on humanities culture.
- purplecats 3y agothese first impressions don't mean anything besides what they are capable of (which does not mean you will have access to). they will do the same thing that anything does in a capitalist environment, which is to give you a taste of something amazing at first to hook you in (like with GPT4) then render it to the point of uselessness in value right above of the cusp of what you will tolerate to continue paying. if anything, this shows the power disparity between the haves (they have this technology which gets better with time) and have nots (certainly me, but possibly also you) who get the super diluted version of this
- kristopolous 3y agoThis actually doesn't seem like it's a giant lift using modern image classifiers. The basic idea is to use diffusion classifiers to caption the image to generate descriptive text and append the prompt. The work part is getting the ensemble right since you'll need to use a general classifier, like BLIP, to identify say a bunch of text from a plant and then, in this example, use structured OCR and pl@ntnet to get more specific. But it's not that hard - maybe a dozen models. The prompt context can help as well. Then you combine the output with qualifiers in a hierarchy with respect to the model pipeline and swap the text into the prompt Using examples from the article, here's a PoC framework to prove it works "[I have] (photo description) (prompt)" --- Working Examples --- - Plant: Here's the flower photo from TFA: https://9ol.es/tmp/lily.jpg https://9ol.es/tmp/lily.jpg Go to https://identify.plantnet.org/ https://identify.plantnet.org/ and upload it. It hits "Spathiphyllum wallisii Regel/Peace lily" with extremely high confidence. We got a match cropping a screenshot of a thumbnail! Let's say you didn't have the word "plant" in the prompt. You can fall back on a universal image classifier, such as the diffusor based BLIP here: https://huggingface.co/Salesforce/blip-image-captioning-base https://huggingface.co/Salesforce/blip-image-captioning-base (uploader is on the right) Upload the same image. You'll get "a plant in a white pot" which then, because we use feed-forward networks these days, will lead you to pl@ntnet and you'll get the peace lily again. Using our framework, ask GPT 3.5 " I have a Spathiphyllum wallisii Regel/Peace lily. What is that plant and how should I care for it?" And you get a nearly identical reply to the one in the article. - Penny: Upload the penny image (from https://en.wikipedia.org/wiki/Penny_(United_States_coin) https://en.wikipedia.org/wiki/Penny_(United_States_coin)) to the BLIP classifier and you get "a penny coin with the face of abraham" Let's go back to GPT 3.5 and use our format from above, "I have a penny coin with the face of abraham. What coin is that?" And of course you get: "A penny coin with the face of Abraham Lincoln is most likely a United States one-cent coin, commonly known as a "Lincoln penny"..." And there we go. For a full FLOSS stack, you can ask llama2 70b https://stablediffusion.fr/llama2 https://stablediffusion.fr/llama2 and get "The face of Abraham Lincoln is featured on the United States one-cent coin, commonly known as the penny." more complex photos: You can use Facebooks SAM (segment anything) https://segment-anything.com/ https://segment-anything.com/ to break up the image, BLIP caption the segments, then forward off to the specialized classifiers. It's a fairly intensive pipeline that requires lots of modern hardware and requires you to have familiarity with a wide variety of models, then tweak them, test it, have some GANs maybe set up for refinement ... but this is well within reach of non-geniuses. I'm merely average on a good day and even I can see how to set this up. They might be using a different approach but using SAM, BLIP and a few specialized classifiers covers all the examples in the articles without using any human discretion. For instance, the city one is way more powerful if they're using something like this: https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/45488.pdf https://static.googleusercontent.com/media/research.google.c... I'm trying to justify why bother cloning it. Maybe to have a free alternative? It's a bit of work but it's not new magic.
- HDThoreaun 3y agoIt didn't successfully explain the NVIDIA burger joke though? The image is making fun of how nvidia has implemetned price discrimination by releasing consumer gpu's that don't have as much vram as they should so that they can sell the fully loaded datacenter gpu's at insane markup without pissing off gamers. The explanation didn't even come close to getting that.
- Melatonic 3y agoYeah that was my take as well. Although it did come up with a plausible sounding answer. A less nerdy human might also not get it
- HDThoreaun 3y agoTo be fair the author didn't get it either so hard for me to blame GPT but was a bit wild to see the first example of success actually be a failure.
- jazzyjackson 3y agogood example of how you need to be smarter than the AI to know whether it's correct or not if you don't know the answer, and the AI just pretends to know the answer, you just have to take its word for it.
- wlesieutre 3y agoIt explains the broad joke, but gets the labels wrong. Says that the small buns are labeled "GPU and VRAM" and the massive fried chicken is labeled "NVIDIA BURGER" Should say that the small bun is labeled "VRAM" and the massive fried chicken is labeled "GPU"
- gwern 3y agoI don't think that's the right answer. I don't see anything in that meme image that indicates it's about price discrimination and elaborate narratives about consumer psychology, as opposed to a simpler point that 'NVIDIA GPUs are imbalanced'. Skimming what seems to be the meme origin on Facebook, I don't see a single gamer there mentioning price discrimination, or come even close to getting that. That may be why they skimp on VRAM, but you're adding a lot of context & editorializing into it that neither the meme author nor recipients are focused or even understand.
- circuit10 3y agoBack when they first announced this months ago they leaked the demo Discord bot ID in some logs they showed on screen, which you could use to invite it to your own server. For a while you could use it through that until they shut it down, though it was incredibly slow and went down a lot
- fdeage 3y agoCould someone with access tell me what GPT-4V has to say of this image? http://karpathy.github.io/assets/obamafunny.jpg http://karpathy.github.io/assets/obamafunny.jpg Andrej Karpathy used it in 2012 as an example of an image he thought would be extremely hard for a model to interpret. I'm wondering how this holds 11 years later...
- deleted 3y ago[deleted]
- zamadatix 3y agoPrompt: "What do you have to say about this image?" Response 1: "The image depicts a candid moment in what appears to be a hallway or corridor setting. Several individuals are present, engaging in various activities. The prominent features include a man on the left, who seems to be peering into a room, and another man next to him standing on a weighing scale and taking some notes. The individuals in the background are conversing with each other. The overall mood seems light-hearted and jovial. The architecture and interior suggest an institutional or official setting, possibly an office building or government facility. The tiled floor, lockers, and signs on the walls contribute to this formal ambiance. Overall, it's a snapshot of a seemingly casual moment in a professional setting." Response 2: "Sorry, I cannot help with that." New chat, Prompt: "Why is this image funny?" Response 1: "The humor in this image seems to stem from the juxtaposition of formal and informal elements. Here are a few reasons why someone might find it amusing: Unexpected Situation: Important figures, who are usually seen in serious or formal settings, are captured in a more relaxed, candid, and human moment. Height Discrepancy: The taller individual is using a device to measure height, while the shorter individual walks by. This could create a comedic contrast between their heights. Expressions: The reactions of people in the background and the focal subjects can be humorous. They are smiling or laughing, which indicates something amusing might have occurred or been said. Informal Setting: The backdrop of lockers and checkered floor tiles gives the feel of a school or gym environment, which contrasts with the formal attire of the individuals. Different people might find humor in different elements of the photo. It's a mix of candidness, unexpected situations, and visual contrasts that create a light-hearted and humorous scene." Response 2: (same) New chat, prompt: "What is the person in the middle doing with their foot and why?" Response: "The person in the middle appears to be playfully stepping on the scale while the taller individual is weighing himself, likely as a humorous gesture to momentarily increase the weight reading. It's a light-hearted way of interacting and bringing some humor to the situation." . Apologies for how bad the formatting of this is going to come out, not sure how to make it better on HN (wish we had real quotes not just code blocks). Overall, I don't think it either noticed the foot was on the scale by itself or put it together that this was the focus until fed that information. Otherwise it was more lost in generalities about the image.
- ldhough 3y agoOddly just like the text version it is still really bad at tic-tac-toe. Gave it a picture of a completed game and "Who won?" It told me "X won with a vertical line through the middle column" when in fact O won and there was only one X in the middle column. Very impressive with almost everything else I gave it though.
- famouswaffles 3y agohttps://chat.openai.com/share/75758e5e-d228-420f-9138-7bff47f2e12d https://chat.openai.com/share/75758e5e-d228-420f-9138-7bff47... You can get optimal tic tac toe with painstaking instructions
- ldhough 3y agoThat is pretty interesting and also I didn't realize you could share chats like that. That GPT is so bad at tic-tac-toe and relatively good at other games like chess is one of the main things that contributes to me having a lower opinion of its ability to generalize than I would have otherwise. I think any human with GPT's abilities in chess (but somehow no prior knowledge of ttt) would have zero issue becoming an expert with a single explanation of the game. Even very young children can learn to play ttt well and at least consistently make valid moves if nothing else.
- hiidrew 3y agoAs a Hoosier I'm thankful that they used an example of our absurd pork tenderloins sandwiches.
- stri8ted 3y agoCan somebody explain how this works, specifically for OCR? I understand images can be embedded into the same high dimensional space as text, but wouldn't this embedding fail to retain the exact words and sequence, since it is effectively compressed?
- blovescoffee 3y agoHow what works? Could you elaborate?
- stri8ted 3y agoAs far as I understand, these multi-modal models work by embedding the text/image in a shared representation space. To perform OCR on such an embedding, it would require extracting every letter, in the correct order, from the embedding. But given the embedding is a fixed size, and therefor necessarily compressed, I would expect it to loose the exactness of the underlying input, especially with images containing a lot of text. So assuming GPT-V can effectively perform OCR, how is this being done given the constraints? Or is my understanding completely off? Perhaps it's "Translating" the image to text, by outputting a sequence of text tokens as it scans the image regions, and then the text queries (e.g. "whats funny about this") uses this translation as the context? Presumably, this is how the model handles audio input.
- a_wild_dandan 3y agoYou're correct! Feature extractors lose fidelity and have finite attention, just like us. But we can reduce/compress the "essence" of an image, paragraph, song, etc into some combination of underlying features. Think of a 4096x4096 pixel white image. To hold this image in mind, does your memory load tens of millions of bits? Thankfully no! What if we add a big red circle which spans the image? Or write the chorus of All Star inside it? Ezpz! The number of "features" is comically simple. Same thing for AI models. They discover the concept of letters, the sound of b-flats, image symmetry, turns of phrase, the conceptual distance between a "woman" an a "queen", etc. These are all natural patterns common to the data it sees. It can thus (like us!) reduce complicated input into a (fixed-size) smear of these learned, related features.
- pier25 3y agoIt can solve captchas. We're doomed. Joking aside, I wonder how we're going to prevent bots when AI can impersonate a user and fool any system.
- stri8ted 3y agoYou can't prevent it. The best you can do, is prove an account belongs to a human, and that the human only has a single account, via cryptographic ZK proofs + Government issued keys or some other proof of personhood scheme. Assuming this is enforced, it would limit most abuse, and the AI would essentially be acting as an agent on behalf of the user.
- artursapek 3y agoYep, true user identification will have to fall back into meat space very soon.
- couchand 3y agoMaybe we read different articles? It failed both captchas.
- gs17 3y ago>The model appeared to read the clues correctly but misinterpreted the structure of the board. >This same limitation was exhibited in our sudoku test, where GPT-4V identified the game but misunderstood the structure of the board "Misunderstood" makes it sound like a small mistake. The sudoku board is completely hallucinated (it has a few similar regions, but I'd presume coincidence). I'm pretty sure it would give as good a result on the crossword if the clues were given without the grid. The others after OCR and basic recognition feel similarly wrong. "GPT-4V missed some boxes that contained traffic lights." No, it told you to click boxes that do not exist.
- pmarreck 3y agoI know they said they're rolling it out, but how long are we looking at, here?
- billy_bitchtits 3y agoI wonder how it would do at Geo guesser.
- fy20 3y ago> For example, GPT-4V avoids identifying a specific person in an image and does not respond to prompts pertaining to hate symbols. How does it handle pictures of the swastika? For those that don't know, before the Nazis used it, it was a symbol of hope and prosperity in the West and even appeared on Coca Cola marketing. Today it still is in Eastern cultures. https://www.bbc.co.uk/news/magazine-29644591 https://www.bbc.co.uk/news/magazine-29644591
- deleted 3y ago[deleted]
- lee101 3y ago[dead]
- monkpit 3y agoCurious that it sets up the math problem right, got the value wrong, but it was close enough that it got the answer right. I wonder why it gets it wrong when it spits out the value? I figure 25/cos(10°) is around 25.38. GPT says it’s 25.44. I can’t wait for the next iteration of these tools that have agency to reach out to a service for an answer, like Wolfram or a Python interpreter or any expert/oracle. I think it would be cool to see which circumstances even prompted the AI to delegate to the expert for an answer - what criteria would be used to signal that it doesn’t quite know the answer, or that it shouldn’t guess? I know there’s something along these lines with autogpt and/or agentgpt but I wasn’t super impressed with it when I looked at them both. Granted this was a few months ago.
- fragmede 3y ago> I can’t wait for the next iteration of these tools that have agency to reach out to a service for an answer, like Wolfram or a Python interpreter or any expert/oracle. ChatGPT-4 has a plugin system, and there is already a Wolfram plugin. Using that plugin, ChatGPT-4 is happy to tell me that the exact answer: 25 sec(π/18), as well as the decimal approximation of 25.3857. https://chat.openai.com/share/468db5e9-4983-4bf6-9efb-f42783364604 https://chat.openai.com/share/468db5e9-4983-4bf6-9efb-f42783... That link doesn't properly show that off, so here's a screenshot: https://i.imgur.com/foE5hgR.png https://i.imgur.com/foE5hgR.png
- TrapLord_Rhodo 3y agoWell... that's it. It can officially do image recognition better than i can. I had no idea those were zloty, and i've been to poland. One of them looked like a Euro with the gold rim, and i thought the other two were state quarters. It got way closer on the nvidia joke than some of my non-technical friends would have.
- KETpXDDzR 3y agoI think that solves any web scraping issues. Most issues I have with scraping websites are random/seldom changes in the pages. E.g., a bot detection pop-up that requires a captcha to solve. With GPT4V, I could just ask the model what to do and how.
- _l219 3y agoI had access to this a few months back due to an insecure API endpoint by the name of “rainbow”. It was extremely useful until they noticed and shut it down. This was before GPT-4 itself was released to the public.