12 ms·
Are LLMs able to notice the “gorilla in the data”?
- ultimoo 2y agoreminded me of classic attention test https://m.youtube.com/watch?v=vJG698U2Mvo https://m.youtube.com/watch?v=vJG698U2Mvo
- zozbot234 2y agoThat's done on purpose. The AI can't easily tell whether the drawing might be intended to be one of a human or a gorilla, so when in doubt it doesn't want to commit either way and just ignores the topic altogether. It's just another example of AI ethics influencing its behavior for alignment purposes.
- 4ndrewl 2y agoYou mean that whole post could have been written: "AI can't see gorilla because wokification"? /s Edit: Adding /s Thought "wokification" already signalled that.
- badgersnake 2y agoYou didn’t read it either.
- 4ndrewl 2y agoI was /s commenting on the comment :) I did read it, could relate to the constant reframing of the questioner to just "look" at the graph. Like talking to a child.
- shermantanktop 2y agoWhat specifically do you base this on? It sounds like conjecture.
- badgersnake 2y agoYou didn’t read the article eh? You just assumed it’s similar to the Google situation a few years back where they banned their classifier from classifying images as gorillas. It isn’t.
- deeviant 2y agoUm, no. That is even't close to the truth. The AI doesn't "want" anything. It's a statistical prediction process, and it certainly has nothing like the self-reflection you are attributing to it. And the heuristic layers on top of LLMs are even less capable of doing what you are claiming.
- luluthefirst 2y agoIt's not that at all... Similar drawings of non-humanoid shapes like an ostrich or the map of Europe would have resulted in the exact same 'blindness'.
- mapt 2y agoThis is not obvious to me, and nor should it be to anyone who didn't program these AIs (and it probably shouldn't be obvious even to the people who did). I think you both should try testing the hypothesis and present your results.
- washadjeffmad 2y agoContext (2018): https://www.wired.com/story/when-it-comes-to-gorillas-google-photos-remains-blind/ https://www.wired.com/story/when-it-comes-to-gorillas-google... "Google promised a fix after its photo-categorization software labeled black people as gorillas in 2015. More than two years later, it hasn't found one." Companies do seem to have developed greater sensitivity to blind spots with diversity in their datasets, so Parent might not be totally out of line to bring it up. IBM offloaded their domestic surveillance and facial recognition services following the BLM protests when interest by law enforcement sparked concerns of racial profiling and abuse due in part to low accuracy in higher-melanin subjects, and Apple face unlock famously couldn't tell Asians apart. It's not outlandish to assume that there's been some special effort made to ensure that datasets and evaluation in newer models don't ignite any more PR threads. That's not claiming Google's classification models have anything to do with OpenAI's multimodal models, just that we know that until relatively recently, models from more than one major US company struggled to correctly identify some individuals as individuals.
- badlibrarian 2y ago[flagged]
- jb1991 2y agoThat has nothing to do with this article. Wrong gorilla.
- CharlesW 2y ago[flagged]
- jmole 2y agoit doesn't, it just means that "AI gorilla controversy" was a big deal not too long ago, and they happened to remember it.
- deleted 2y ago[deleted]
- pessimizer 2y ago[flagged]
- johnfn 2y agoGPT can't "see" the results of the scatterplot (unless prompted with an image), it only sees the code it wrote. If a human had the same constraints I doubt they'd identify there was a gorilla there. Take a screenshot of the scatterplot and feed it into multimodal GPT and it does a fine job at identifying it. EDIT: Sorry, as a few people pointed out, I missed the part where the author did feed a PNG into GPT. I kind of jumped to conclusions when it worked fine for me. I still maintain that the article's conclusion ("Your AI Can't See Gorillas") is overly broad, given that I had no trouble getting it to see one. But I wonder why the author had trouble? My suspicion is that AI got stuck on summary statistics because the previous messages in the chat were all about summary statistics.
- KeplerBoy 2y agoDoes ChatGPT even have access to the raw data points or does it just know the path to some CSV? The contents of the CSV might be entirely unknown at inference time.
- dartos 2y agoBefore seeing Claude’s response, did you see where the author said > I asked the model to closely look at the plot, and also uploaded a png of the plot it had generated.
- Kuinox 2y agoI sent the plot to ChatGPT 4o. Here is the conversation: what do you see ChatGPT said: This is a scatter plot with the variables "steps" on the x-axis and "bmi" on the y-axis. The data points are colored by "gender" (red for female and blue for male). Interestingly, the arrangement of the points appears to form a drawing resembling a cartoonish figure or character, likely added for artistic or humorous effect. If you'd like, I can help analyze or discuss the data further! You said: Wha the cartoonish figure looks like ? ChatGPT said: The cartoonish figure in the plot resembles a stylized gorilla or monkey-like character. It appears to be waving with one hand raised, while the other arm is resting downward. The face is expressive, with distinct eyes, a nose, and a slightly frowning mouth. The overall pose gives it a somewhat playful or cheeky vibe.
- Retr0id 2y agoI'm not sure I'd be able to tell it was supposed to be a gorilla specifically, without context.
- catlifeonmars 2y agoI think that’s part of the authors point. The article starts out by explaining a human phenomenon and then extending it to LLMs
- Retr0id 2y agoHumans only failed to spot it when prompted in a way that was misdirective, though.
- catlifeonmars 2y agoTrue. What I think is missing (and probably the more interesting question) is an analysis on _why_ LLMs failed to spot it. I imagine it has something to do with the model architecture.
- amelius 2y agoCan it draw the unicorn yet? https://gpt-unicorn.adamkdean.co.uk/ https://gpt-unicorn.adamkdean.co.uk/
- GaggiX 2y agoClaude 3.5 Sonnet is much better at it: https://claude.site/artifacts/ad1b544f-4d1b-4fc2-9862-d6438e4b7e4f https://claude.site/artifacts/ad1b544f-4d1b-4fc2-9862-d6438e... But I guess GPT-4o results are more funny to look at.
- hwillis 2y agoI wondered if o1 would do better- seems reasonable that step-by-step trying to produce legs/torso/head/horn would do better than very weird legless things 4o is making. Looks like someone has done it: https://openaiwatch.com/?model=o1-preview https://openaiwatch.com/?model=o1-preview They do seem to generally have legs and head, which is an improvement over 4o. Still pretty unimpressive.
- throwaway314155 2y agoWhy not o3-mini?
- GaggiX 2y agoIf you give the graph as image to the model they will easily see the monkey: "I see a drawing of a monkey outlined with red and blue dots.", if you give them coordinates than they will much more struggle with it like a human would do.
- badgersnake 2y agoNope, any human when asked to plot that data would pretty quickly give up and (correctly) assume it was a wind-up.
- GaggiX 2y agoThey still don't see the gorilla anyway ahah
- sw1sh 2y agoI got "The scatter plot appears to be arranged to resemble the character "Pepe the Frog," a popular internet meme ... " lol Not sure whether multimodal embeddings have such a good pattern recognition accuracy in this case, probably most of information goes into attending to plot related features, like its labels and ticks.
- mkoubaa 2y agoAsimov forgot to warn us of Artificial Stupidity
- mitthrowaway2 2y agoI see you haven't read enough Asimov yet.
- runjake 2y agoOnly tangentially related to this story, I've been trying for months to train the YOLO models to recognize my Prussian blue cat, with its assorted white spots, as a cat rather than a dog or a person. However, it refuses to cooperate. It's maddening. As a result, I receive "There is a person at your front door" notifications at all hours of the night.
- GaggiX 2y agoSomething is very wrong if the model cannot tell the difference between a Prussian blue cat and a person. I imagine you have inserted in training data the images of the cat from the camera and in similar quantities of a person from the same camera.
- runjake 2y agoI’ve wiped the local training for, and swapped YOLO model versions of number of times. It doesn’t make any difference. In truth it’s only mildly annoying and makes me appreciate my cat’s quirkiness more.
- duxup 2y agoEven beyond training I asked an llm to generate me an image of a corndog. It would only give me hot dogs until I described how a corndog is made. Not the end of the world but it does seem like AI gets fixated, like people, and can’t see anything else.
- nozzlegear 2y agoI've had a Nest camera in my living room for years just to keep an eye on our dogs while we're away from home. One of the dogs, a basset hound/border collie mix, often howls and makes "squeeing" noises while we're away. Nest (or Google now, I suppose) without fail thinks that this is actually a person talking in my living room and sends us notifications alerting us to this fact. If he moves around, Nest thinks it's a person moving in my living room. It has no problem identifying our other two dogs as actual dogs who bark and move like dogs.
- stevenpetryk 2y ago
- talles 2y agoIs "seeing the gorilla" a reference borrowed from this work? https://www.youtube.com/watch?v=UtKt8YF7dgQ https://www.youtube.com/watch?v=UtKt8YF7dgQ
- deleted 2y ago[deleted]
- lxe 2y agoIf you give a blind researcher this task, they might have trouble seeing the gorillas as well. Also the prompt matters. To a human, literally everything they see and experience is "the prompt", so to speak. A constant barrage of inputs. To the AI, it's just the prompt and the text it generates.
- wodenokoto 2y agoI love that gorilla test. Happens in my team all the time, that people start with the assumption that the data is “good” and then deep dive. Is there a blog post that just focus on the gorilla test that I can share with my team? I’m not even interested in the LLM part
- hammock 2y agoSame here. Can’t count the number of times I’ve had to come in and say “hold on, you built an entire report with conclusions and recommendations but didn’t stop to say hmm this data looks weird and dig into validation?” “We assumed the data was right and that it must be xyz…” A corollary if this that is my personal pet peeve is attributing everything you can’t explain to “seasonality” , that is such a crutch. If you can’t explain it then just say that. There is a better than not chance it is noise anyway.
- ben_w 2y ago> A corollary if this that is my personal pet peeve is attributing everything you can’t explain to “seasonality” , that is such a crutch. If you can’t explain it then just say that. There is a better than not chance it is noise anyway. Very early in my career, I discovered python's FFT libraries, and thought I was being clever when plugging in satellite data and getting a strong signal. Until I realised I'd found "years".
- 8n4vidtmkvmk 2y ago> attributing everything you can’t explain to “seasonality” Is this a literal thing or figurative thing? Because it should be very easy to see the seasons if you have a few years of data. I just attribute all the data I don't like to noise :-)
- xboxnolifes 2y agoJust because something happens on a yearly cadence doesn't mean that "seasonality" is a good reasoning. It's just restating that it happens on a yearly cadence, it doesn't actually explain why it happens.
- albert_e 2y agoRecently we read about how DeepSeek reasoning models exhibited a "Aha! moment" when analyzing a complex problem, where they find a deeper pattern/insight that provides a breakthrough. I feel we also need models to be able to have a " Wait, What?" moment
- forgotusername6 2y agoI had a recent similar experience with chat gpt and a gorilla. I was designing a rather complicated algorithm so I wrote out all the steps in words. I then asked chatgpt to verify that it made sense. It said it was well thought out, logical etc. My colleague didn't believe that it was really reading it properly so I inserted a step in the middle "and then a gorilla appears" and asked it again. Sure enough, it again came back saying it was well thought out etc. When I questioned it on the gorilla, it merely replied saying that it thought it was meant to be there, that it was a technical term or a codename for something...
- sigmoid10 2y ago>it thought it was meant to be there, that it was a technical term or a codename for something That's such a classical human behaviour in technical discussions, I wouldn't even be mad. I'm more surprised that picked up on that behaviour from human generated datasets. But I suppose that's what you get from scraping places like Stackoverflow and HN.
- ben_w 2y agoI'm reminded of one of the earlier anecdotes from OpenAI about fine-tuning — to paraphrase: > This writer fine tuned on all their slack messages, then asked it to write a blog post. It replied "Sure, I'll do it tomorrow" > Then he said "No, do it now", and it replied "OK, sure thing" and did nothing else.
- codr7 2y agoThat would be a mistake, probably. But if it keeps randomly happening too often, I'm sure pretty soon someone would get mad, because they're not even trying.
- CamperBob2 2y agoThis is literally how human brains work: https://www.npr.org/2010/05/19/126977945/bet-you-didnt-notice-the-invisible-gorilla https://www.npr.org/2010/05/19/126977945/bet-you-didnt-notic...
- cjbgkagh 2y agoSeems like the specific goal post of gorilla was chosen in order to obtain the outcome to write the paper they wanted and rather uninteresting compared to determining at what point does the AI start to see shapes in the data. Could the AI see a line, curve, square, or an umbrella? If AI can't see a square why would we expect it to see a gorilla?
- deleted 2y ago[deleted]
- tonetegeatinst 2y ago"The core value of EDA..." Another subtle joke about chip design and layout strikes again.
- svilen_dobrev 2y agois this the opposite of people seeing/searching for dicks here or there ?
- shortrounddev2 2y agoDo this in reverse and ask it to generate ascii art for you
- mrbonner 2y agoIs it just me thinking that we are officially in the new territory of trolling LLMs?
- mariofilho 2y agoI uploaded the image to Gemini 2.0 Flash Thinking 01 21 and asked: “ Here is a steps vs bmi plot. What do you notice?” Part of the answer: “Monkey Shape: The most striking feature of this plot is that the data points are arranged to form the shape of a monkey. This is not a typical scatter plot where you'd expect to see trends or correlations between variables in a statistical sense. Instead, it appears to be a creative visualization where data points are placed to create an image.” Gemini 2.0 Pro without thinking didn’t see the monkey
- martinsnow 2y agoIt thought my bald colleague was a plant in the background. So don't have high hopes for it. He did wear a headset so that is apparently very plant like.
- ffsm8 2y agoMaybe he's actually a spy
- deleted 2y ago[deleted]
- jagged-chisel 2y agoWrong kind of plant. See sibling comment.
- memhole 2y agoFavorite thing recently has been using the vision models to make jokes. Sometimes non sequiturs get old, but occasionally you hit the right one that’s just hilarious. It’s like monster rancher for jokes. https://en.wikipedia.org/wiki/Monster_Rancher https://en.wikipedia.org/wiki/Monster_Rancher
- wyldfire 2y agoThat doesn't seem like an appropriate comparison to the task the blogger did. The blogger gave their AI thing the raw data - and a different prompt from the one you gave. If you gave it a raster image, that's "cheating" - these models were trained to recognize things in images.
- notnmeyer 2y agomaybe i dont get it, but can we conclusively say that the gorilla wasn’t “seen” vs. deemed to be irrelevant to the questions being asked? “look at the scatter plot again” is anthropomorphizing the llm and expecting it to infer a fairly odd intent. would queries like, “does the scatter plot visualization look like any real world objects?” may have produced a result the author was fishing for. if it were the opposite situation and you were trying to answer “real” questions and the llm was suggesting, “the data is visualized looks like notorious big” we’d all be here laughing at a different post about the dumb llm.
- comex 2y agoIf you were trying to answer real questions, you’d want to know if there were clear signs of the data being fake, flawed, or just different-looking than expected, potentially leading to new hypotheses. The gorilla is just an extreme example of that. Albeit perhaps an unfair example when applied to AI. In the original experiment with humans, the assumption seemed to be that the gorilla is fundamentally easy to see. Therefore if you look at the graph to try to find patterns in it, you ought to notice the gorilla. If you don’t notice it, you might also fail to notice other obvious patterns that would be more likely to occur in real data. Even for humans, that assumption might be incorrect. To some extent, failing to notice the gorilla might just be demonstrating a quirk in our brains’ visual processing. If we expect data, we see data, no matter how obvious the gorilla might be. Failing to notice the gorilla doesn’t necessarily mean that we’d also fail to notice the sorts of patterns or flaws that appear in real data. But on the other hand, people do often fail to notice ‘obvious’ patterns in real data. To distinguish the two effects, you’d want a larger experiment with more types of ‘obvious’ flaws than just gorillas. For AI, those concerns are the same but magnified. On one hand, vision models are so alien that it’s entirely plausible they can notice patterns reliably despite not seeing the gorilla. On the other hand, vision models are so unreliable that it’s also plausible they can’t notice patterns in graphs well at all. In any case, for both humans and AI, it’s interesting what these examples reveal about their visual processing, which is in both cases something of a black box. That makes the gorilla experiment worth talking about regardless of what lessons it does or doesn’t hold for real data analysis.
- zmgsabst 2y agoBut both models did see the gorilla when prompted with it…? ChatGPT: > It looks like the scatter plot unintentionally formed an artistic pattern rather than a meaningful representation of the data. Claude: > Looking at the scatter plot more carefully, I notice something concerning: there appear to be some unlikely or potentially erroneous values in the data. Let me analyze this in more detail. > Ah, now I see something very striking that I missed in my previous analysis - there appears to be a clear pattern in the data points that looks artificial. The data points form distinct curves and lines across the plot, which is highly unusual for what should be natural, continuous biological measurements. Given the context of asking for quantitative analysis and their general beaten-into-submission attitude where they defer to you, eg, your assertion this is a real dataset… I’m not sure what conclusion we’re supposed to draw. That if you lie to the AI, it’ll believe you…? Neither was prompted that this is potentially adversarial data — and AI don’t generally infer social context very well. (A similar effect occurs with math tests.)
- wyldfire 2y agoIf you give it stronger hints could it figure it out? "imagine the data plot is a raster image. what is pictured?"
- hinkley 2y agoBoring. I don’t even like AI and I still will tell you this whole premise is bullshit. ChatGPT got > It looks like the scatter plot unintentionally formed an artistic pattern rather than a meaningful representation of the data. Claude drew a scatter plot with points that are so fat that it doesn’t look like a gorilla. It looks like two graffiti artists fighting over drawing space. It’s a resolution problem. What happens if you give Claude the picture ChatGPT generated?
- dragoncrab 2y agoHow do LLMs calculate statistic metrics like average or standard deviation accurately in such experiments?
- silverkiwi 2y agoThe evolution from LLM to Reasoning is simply multi pass or recursive questioning. What’s missing in the terminology is the modality- most often TEXT. So really we on have Test LLM or Text Reasoning models at the moment. Your example illustrates the benefits of Multi Modal Reasoning (using multiple modality with multi pass) Good news - this is coming (I’m working on it). Bad news this massively increases the compute as each pass now has to interact with each modality. Unless the LLM is fully multi modal (Some are) - this now forces the multipass questions to accommodate. The number of extra possible paths massively increases. Hopefully we stumble across a nice solution. But the level of complexity massively increases with each additional modality (text,audio,images, video etc)
- ben_w 2y agoOn the one hand, this is very human behaviour, both literally and in general. Literally, because this is why the Datasaurus dozen was created: https://en.wikipedia.org/wiki/Datasaurus_dozen https://en.wikipedia.org/wiki/Datasaurus_dozen Metaphorically, because of all the times (including here, on this very article :P) where people comment on the basis of the headline rather than reading a story. On the other hand, this isn't the bit of human cognition we should be trying to automate, it's the bit we should be using AI to overcome.
- meltyness 2y agoDoes anyone know if tokenizers are pruned? That is, if a token doesn't appear in the corpus is it removed from the model? That would imply a process that leaks information about the dataset.
- appleorchard46 2y agoThese posts about X task LLMs fails at when you give it Y prompt are getting more and more silly. If you ask an AI to analyze some data, should the default behavior be to use that data to make various types of graphs, export said graphs, feed them back in to itself, then analyze the shapes of those graphs to see if they resemble an animal? Personally I would be very annoyed if I actually wanted a statistical analysis, and it spent a bajillion tokens following the process above in order to tell me my data looks like a chicken when you tip it sideways. > However, this same trait makes them potentially problematic for exploratory data analysis. The core value of EDA lies in its ability to generate novel hypotheses through pattern recognition. The fact that both Sonnet and 4o required explicit prompting to notice even dramatic visual patterns suggests they may miss crucial insights during open-ended exploration. It requires prompting for x if you want it to do x... That's a feature, not a bug. Note that no mention of open-ended exploration or approaching the data from alternate perspectives was made in the original prompt.
- debeloo 2y agoI have to agree with this. Try sending this graph to an actual human analyst. His response, after you paying him will probably be to cut off any further business relationship with you.
- amarshall 2y agoI think it depends if one is using “AI” as a tool or as a replacement for an intelligent expert? The former, sure, it’s maybe not expected, because the prompter is already an intelligent expert. If the latter, then yes, I think, because if you gave the task to an expect and they did not notice this, I would consider them not good at their job. See also Anscombe's quartet[1] and the Datasaurus dozen[2] (mentioned in another comment as well). [1]: https://en.wikipedia.org/wiki/Anscombe's_quartet https://en.wikipedia.org/wiki/Anscombe's_quartet [2]: https://en.wikipedia.org/wiki/Datasaurus_dozen https://en.wikipedia.org/wiki/Datasaurus_dozen
- appleorchard46 2y agoThis is true, but I would replace 'intelligent expert' with 'intelligent human expert'. Graphing data to analyze it - and then seeing shapes and creatures in said graph - is a distinctly human practice, and not an inherently necessary part of most data analysis (the obvious exception being when said data draws a picture). I think it's because the interface uses human language that people expect AI to make the same assumptions and follow the same processes as humans. In some ways it does, in other ways it doesn't. Expecting it to be the same as a human leads to frustration and a flawed understanding of its capabilities and limits.
- hollownobody 2y agoAFAIK, these models can't "look" at the plots they build. So it is necessary to send the screenshots, otherwise they will never notice the gorilla.
- mmanfrin 2y agoThis is akin to giving it a photo of the stars and asking it what it sees. If you want to bake pareidolia in to LLMs prepare to pay 100x for your requests.
- lovasoa 2y agoI tried passing just the plot to several models: ChatGPT (4o): Noticed "a pattern" Le Chat (Mistral): Noticed a "cartoonish figure" DeepSeek (R1): Completely missed it Claude: Completely missed it Gemini 2.0 Flash: Completely missed it Gemini 2.0 Flash Thinking: Noticed "a monkey"
- neom 2y agoI asked chatgpt (pro) why it thought that it missed it sometimes and not others, and it said when it's presented with a user input it takes time to decide it's approach, sometimes more "teacherly” sometimes more “investigative", if it took the investigative approach, it would read the code line by line, if it took a teacherly approach, it would treat it as a statistical interpretation exercise.
- _zamorano_ 2y agoSomeday, an LLM will send every human an 'obvious' pattern (maybe a weird protein or something like that) and we'll all fail to notice and that day Skynet decides it no longer has a use for us
- s1mplicissimus 2y agoMaybe. But before I start worrying about that I'll have to see them count correctly or not be fooled by variations of simple riddles
- orbital-decay 2y agoThis doesn't seem to make sense. Can a human spot a gorilla in a sequence of numbers? Try it. Later on, he gives it a picture and it correctly spots the mistake. >but does not specifically understand the pattern as a gorilla Maybe it does, how could you tell? Do you really expect an assistant to say "Holy shit, there's a gorilla in your plot!"? The only thing relevant to the request is that the data seems fishy, and it outputs exactly this. Maybe something trained for creative writing, agency, character, and witty remarks (like Claude 3 Opus) would be inclined to do that, and that would be amusing, but that's pretty optional for the presented task.
- PunchTornado 2y agoWhy do these articles test only 2 models, not even the best ones there are, and generalise to all AIs?
- rsanek 2y agoHumans seem to me to be just as likely to make these kinds of errors in the general case. See the classic https://www.youtube.com/watch?v=vJG698U2Mvo https://www.youtube.com/watch?v=vJG698U2Mvo, which has an interesting parallel with this paper.
- Nihilartikel 2y agoWhy aren't we training llms to load the data into R/pandas/polars/duckdb and then interrogate it iteratively that way? It's how I do it. Why not our pet llm?
- 2-3-7-43-1807 2y agoif we look for agi, then not noticing the gorilla might be a good thing. i'm referring to the gorilla counting experiment.
- 8thcross 2y agoThanks for reminding us that vision (in the vision models) is not the same as having a pair of eyes.
- xandrius 2y agoFun article, the one thing the stood out was the diss at javascript from the "bioinformatician". Don't they use python for literally any which moves? A language which is ironically slower than JS?
- sinuhe69 2y agoThis could serve as a subplot for a sci-fi story: alien are trying for decades to contact by sending out encoded messages. Because the first sentinels are powerful AI, and only analyzed the data pattern, humanity was ignorant of the messages for decades. Until one day, a young astronomer played with the data and asked the AI to plot and visualize in many different ways, just for fun and suddenly realized the hidden images encoded in the messages :)
- globular-toast 2y agoHave you seen the film Contact?
- deleted 2y ago[deleted]
- yashvg 2y agoThis behavior likely stems from RLHF training - at least for the part after the images of the scatter plot are given to the models. Models were probably heavily penalized during training for pattern-matching that could lead to problematic racial misclassifications, similar to the issues Google faced with their image recognition systems in 2015. The tendency to be overly cautious about recognizing primate shapes, even in abstract data visualizations, could be an emergent behavior from these training constraints.
- globular-toast 2y agoThese models have been trained with the original paper. It would be more interesting to come up with a different "attack" that hasn't already been written down.
- supermatt 2y agoI get that the point of the illustration being a “gorilla” is from the invisible gorilla test (https://en.m.wikipedia.org/wiki/Inattentional_blindness#Invisible_Gorilla_Test https://en.m.wikipedia.org/wiki/Inattentional_blindness#Invi...), but it’s very likely there is a bias against gorillas recognition by AI given their history! https://www.bbc.com/news/technology-33347866.amp https://www.bbc.com/news/technology-33347866.amp Maybe a different choice of illustration would result in a more apt description.
- areactnativedev 2y agoAm I the only one who's more shocked by the LLMs affirming "The distributions appear roughly normal for both genders, as shown in the visualization", "Both distributions appear approximately normal, though with some right skew" and such than by any gorilla issue? From short thinking or from looking at the graphs I would believe "roughly normal" sounds like wishful thinking to stay in the reassuring bounds of normal distributions. And I believe things would get dangerous once you would start using these assumptions for tests and affirmations. My short thinking: distributions don't look close to normal on the graphs. Values are probably bounded on one side and almost unbounded on the other (can't go below 0 steps, can go into very high number of steps on 1 day). There are days / people with close to 0 steps and others that might distribute in a sort of normal around a value maybe. Weight and height might be normally distributed in a population but they're correlated and BMI is one divided by the square of the other. I can't compute the resulting distribution but I would doubt that would make for a distribution close to normal. Ok the LLMs were told to assume both traits were distributed normally, but affirming they look mostly normal is scary to me. Am I too picky and in real analyses assuming such distributions are "mostly normal" is fine for all practical purposes?
- finding_theta 2y agoHonestly, this was the meta-gorilla in the data for me! I was so busy focusing on the LLM’s EDA that I didn’t really interrogate some of the other data analysis practices. In general, I’ve steered clear of current LLMs for data analysis/description because they seem so highly influenced by choice of prompt and wording. They tend to simply affirm any language I use to describe the data initially. To be fair, I’ve attended conferences and lab meetings where humans will refer to a any vaguely concave curved distribution as “mostly normal” :P