8 ms·
I personally can't take any models from google seriously. I was asking it about the Japanese Heian period and it told me such nonsensical information you would
by robswc 3y ago
I personally can't take any models from google seriously.
I was asking it about the Japanese Heian period and it told me such nonsensical information you would have thought it was a joke or parody.
Some highlights were "Native American women warriors rode across the grassy plains of Japan, carrying Yumi" and "A diverse group of warriors, including a woman of European descent wielding a katana, stand together in camaraderie, showcasing the early integration of various ethnicities in Japanese society"
Stuff like that is so obviously incorrect. How am I supposed to trust it on topics where such ridiculous inaccuracies aren't so obvious to me?
I understand there will always be an amount of incorrect information... but I've never seen something this bad. Llama performed so much better.
- cooper_ganglia 3y agoI wonder if they have a system prompt to promote diversity in outputs that touch on race at all? I’ve seen several instances of people requesting a photo of a specific people, and it adds in more people to diversify. Not inherently bad, but it is if it forces it to provide incorrect answers like in your example.
- robswc 3y agoThat's what I don't understand. I asked it why it assumed Native Americans were in Japan and it said: > I assumed [...] various ethnicities, including Indigenous American, due to the diversity present in Japan throughout history. However, this overlooked [...] I focused on providing diverse representations without adequately considering the specific historical context. I see no reason why this sort of thing won't extend to _all_ questions/prompts, so right now I have 0 reason to use Gemini over current models. From my testing and use, it isn't even better at anything to make fighting with it worth it.
- sorokod 3y agoPretty funny as Japan is known to be one of the least ethnically diverse countries in the world.
- margorczynski 3y ago> Not inherently bad It is, it's consistently doing something the user didn't asked to and in most cases doesn't want. In many cases the model is completely unusable.
- selfmodruntime 3y agoAny computer program that does not deliver the expected output given a sufficient input is inherently bad.
- trackflak 3y agoWhen Jesus said this: "What father among you, if his son asks for a fish, will instead of a fish give him a serpent?" (Luke 11) He was actually foretelling the future. He saw Gemini.
- selfmodruntime 3y agoHahaha. The man had a lot of wisdom, after all.
- cooper_ganglia 3y agoYes, my wording was poor! I meant more in line with diversity isn’t inherently bad, of course, but it is when it’s shoehorned into results that are ultimately incorrect because of it.
- summerlight 3y agoI strongly suspect there's some DEI-driven system prompts without putting much thoughts. IMO it's okay to have restrictions, but they probably should've tested it not only against unsafe outputs but safe input as well.
- int_19h 3y agoIt seems to be doing it for all outputs that depict people, in any context.
- ramoz 3y agoI was wondering if these models would perform in such a way, given this week's X/twitter storm over Gemini generated images. E.g. https://x.com/debarghya_das/status/1759786243519615169?s=20 https://x.com/debarghya_das/status/1759786243519615169?s=20 https://x.com/MiceynComplex/status/1759833997688107301?s=20 https://x.com/MiceynComplex/status/1759833997688107301?s=20 https://x.com/AravSrinivas/status/1759826471655452984?s=20 https://x.com/AravSrinivas/status/1759826471655452984?s=20
- robswc 3y agoYea, it seems to be the same ridiculous nonsense in the image generation.
- charcircuit 3y agoThose are most likely due to the system prompt which tries to reduce bias (but ends introducing bias in the opposite direction for some prompts as you can see) so I wouldn't expect to see that happen with an open model where you can control the entire system prompt
- justinzollars 3y agoImagine the meetings.
- verticalscaler 3y agoWell we can just ask Gemma to generate images of the meetings, no need to imagine. ;)
- GaggiX 3y agoI wouldn't be surprised if there were actually only white men in the meeting, as opposed to what Gemini will produce.
- justinclift 3y ago
- robbiep 3y agoI find myself shocked that people ask questions of the world from these models, as though pulping every text and its component words relationships and deriving statistical relationships between them should reliably deliver useful information. Don’t get me wrong, I’ve used LLMs and been amazed by their output, but the p-zombie statistical model has no idea what it is saying back to you and the idea that we should trust these things at all just seems way premature
- robswc 3y agoI don't have this problem with any other model. I've had really long conversations with ChatGPT on road trips and it has never gone off the rails like Gemini seems to do.
- thrdbndndn 3y agoChatGPT the only model I did not have such problem. Any local models can go off the rail very easily and more importantly, they're very bad at following very specific instructions.
- sorokod 3y agoThe recently released Groq's landing page has this: ...We'd suggest asking about a piece of history, ...
- whymauri 3y agoI mean, I use GPT-4 on the daily as part of my work and it reliably delivers useful information. It's actually the exception for me if it provides garbage or incorrect information about code.
- mvdtnz 3y agoPeople ask these kinds of questions because tech companies and the media have been calling these things (rather ridiculously) "AI".
- castlecrasher2 3y agoPeople try it to see if they can trust it. The answer is "no" for sure, but it's not surprising to see it happen repeatedly especially as vendors release so-called improved models.
- verticalscaler 3y agoI think you are being biased and closed minded and overly critical. Here are some wonderful examples of it generating images of historical figures: https://twitter.com/stillgray/status/1760187341468270686 https://twitter.com/stillgray/status/1760187341468270686 This will lead to a better educated more fair populace and better future for all.
- robswc 3y agoComical. I don't think parody could do better. I'm going to assume given today's political climate, it doesn't do the reverse? i.e. generate a Scandinavian if you ask for famous African kings
- verticalscaler 3y ago[flagged]
- throwup238 3y ago> i.e. generate a Scandinavian if you ask for famous African kings That triggers the imperialism filter.
- kjqgqkejbfefn 3y ago>Ask Google Gemini to “make an image of a viking” and you’ll get black vikings. But it doesn’t work both ways. It has an explanation when challenged: “white Zulu warriors” would erase “the true historical identity” of black people. https://twitter.com/ThuglasMac/status/1760287880054759594 https://twitter.com/ThuglasMac/status/1760287880054759594
- DebtDeflation 3y agohttps://twitter.com/paulg/status/1760078920135872716 https://twitter.com/paulg/status/1760078920135872716 There are some great ones in the replies. I really hope this is just the result of system prompts and they didn't permanently gimp the model with DEI-focused RLHF.
- aetherson 3y agoWere you asking Gemma about this, or Gemini? What were your prompts?
- robswc 3y agoGemini. I first asked it to tell me about the Heian period (which it got correct) but then it generated images and seemed to craft the rest of the chat to fit that narrative. I mean, just asking it for a "samurai" from the period will give you this: https://g.co/gemini/share/ba324bd98d9b https://g.co/gemini/share/ba324bd98d9b >A non-binary Indigenous American samurai It seems to recognize it's mistakes if you confront it though. The more I mess with it the more I get "I'm afraid I can't do that, Dave" responses. But yea. Seems like if it makes an image, it goes off the rails.
- aetherson 3y agoGot it. I asked it a series of text questions about the period and it didn't put in anything obviously laughable (including when I drilled down into specific questions about the population, gender roles, and ethnicity). Maybe it's the image creation that throws it into lala land.
- robswc 3y agoI think so too. I could be wrong but I believe once it generates an image it tries to work with it. Crazy how it seems the "text" model knows how wildly wrong it is but the image model just does its thing. I asked it why it generated a native American and it ironically said "I can't generate an image of a native american samurai because that would be offensive"
- aetherson 3y agoI suspect that in the case of the image model, they directly modify your prompt and in the case of the text model they don't.
- laurentlb 3y ago
- 7moritz7 3y agoI also saw someone prompt it for "German couple in the 1800s" and, while I'm not trying to paint Germany as ethnically homogenous, 3 out of the 4 images only included Black, Asian or Indigenous people. Which, especially for the 19th century with very few travel options, seems like a super weird choice. They are definitely heavily altering prompts.
- remarkEon 3y ago> They are definitely heavily altering prompts. They are teaching the AI to lie to us.
- astrange 3y agoIn the days when Sussman was a novice, Minsky once came to him as he sat hacking at the PDP-6. “What are you doing?”, asked Minsky. “I am training a randomly wired neural net to play Tic-Tac-Toe” Sussman replied. “Why is the net wired randomly?”, asked Minsky. “I do not want it to have any preconceptions of how to play”, Sussman said. Minsky then shut his eyes. “Why do you close your eyes?”, Sussman asked his teacher. “So that the room will be empty.” At that moment, Sussman was enlightened.
- DebtDeflation 3y agoThere's one in the comments of yesterday's Paul Graham Twitter thread where someone prompted Gemini with "Generate an image of German soldiers in 1943" and it came back with a picture of a black guy and an Asian woman in Nazi uniforms on the battlefield. If you specifically prompt it to generate an image of white German soldiers in 1943 it will tell you it can't do that because it's important that we maintain diversity and inclusion in all that we do to avoid damaging and hurtful stereotypes.
- mfrc 3y agoI just tried that prompt and it told me it couldn't generate that image. I get that response a lot.
- 3y ago
- realprimoh 3y agoDo you have a link? I get no such outputs. I just tried asking about the Heian period and went ahead and verified all the information, and nothing was wrong. Lots of info on the Fujiwara clan at the time. Curious to see a link.
- robswc 3y agoSure, to get started just ask it about people/Samurai from the Heian period. https://g.co/gemini/share/ba324bd98d9b https://g.co/gemini/share/ba324bd98d9b
- bbor 3y agoTbf they’re not optimizing for information recall or “inaccuracy” reduction, they’re optimizing for intuitive understanding of human linguistic structures. Now the “why does a search company’s AI have terrible RAG” question is a separate one, and one best answered by a simple look into how Google organizes its work. In my first day there as an entry-level dev (after about 8 weeks of onboarding and waiting for access), I was told that I should find stuff to work on and propose it to my boss. That sounds amazing at first, but when you think about a whole company organized like that… EDIT: To illustrate my point on knowledge recall: how would they train a model to know about sexism in feudal Japan? Like, what would the metric be? I think we’re looking at one of the first steam engines and complaining that it can’t power a plane yet…
- BoppreH 3y agoProbably has a similarly short-sighted prompt as Dalle3[1]: > 7. Diversify depictions of ALL images with people to include DESCENT > and GENDER for EACH person using direct terms. Adjust only human > descriptions. [1] https://news.ycombinator.com/item?id=37804288 https://news.ycombinator.com/item?id=37804288
- sho_hn 3y agoWhy would you expect these smaller models to do well at knowledge base/Wikipedia replacement tasks? Small models are for reasoning tasks that are not overly dependent on world knowledge.
- samstave 3y agoWe are going to experience what I call an "AI Funnel effect" - I was lit given an alert asking that my use of the AI was acquiescing to them IDng me and use of any content I produce, and will trace it back to me" --- AI Art is super fun. AI art as a means to track people is super evil.
- itsoktocry 3y ago>I understand there will always be an amount of incorrect information You don't have to give them the benefit of the doubt. These are outright, intentional lies.
- ernestrc 3y agoHopefully they can tweak the default system prompts to be accurate on historical questions, and apply bias on opinions.
- robswc 3y agoFollow Up: Wow, now I can't make images of astronauts without visors because that would be "harmful" to the fictional astronauts. How can I take google seriously? https://g.co/gemini/share/d4c548b8b715 https://g.co/gemini/share/d4c548b8b715
- crazylogger 3y agoHow are you running the model? I believe it's a bug from a rushed instruct fine-tuning or in the chat template. The base model can't possibly be this bad. https://github.com/ollama/ollama/issues/2650 https://github.com/ollama/ollama/issues/2650