5 ms·
Woe is Google. How did it get this bad?
by throwaway2562 3y ago
Woe is Google. How did it get this bad?
- wly_cdgr 3y agoIt's always been this bad. It just hasn't always been this visibly bad.
- deleted 3y ago[deleted]
- mrweasel 3y agoOne podcast (Coder Radio) suggested that it might be due to corporate culture and people being afraid of reporting issues upward. While I do think it's culture, I think it stems from Googles advertising based business model. I believe they have attempted to make their LLM safe for advertisers and prefer to err on the side of being safe, even if the risk of being non-brand-safe is minimum.
- miohtama 3y agoThe Economist did a piece on how Google's culture affected Gemini. HN link here: https://news.ycombinator.com/item?id=39550359 https://news.ycombinator.com/item?id=39550359
- a_wild_dandan 3y agoYannic gave his anecdotal experience about Google's culture: https://youtu.be/Fr6Teh_ox-8 https://youtu.be/Fr6Teh_ox-8 It was eye-opening for me. I had no idea things were so bad.
- throwaway2562 3y agoInternal culture hijack successful. Great video - thanks!
- balder1991 3y agoProbably not exactly afraid of reporting things, but no incentive to do so. Corporate culture has a lot of this “not my problem” mentality.
- xvector 3y agoYou are not gonna report dumb DEI initiatives if the DEI crowd will get you fired for it. And this is how you end up with LLMs that refuse to teach C++ due to "safety" or trying to convince us that Abraham Lincoln was Black and Nazis had Hispanics. Fundamentally these megacorporations need to drop these useless non-engineering functions (not that all non-engineering is useless, but these functions are) if they really want to get to AGI.
- IncreasePosts 3y agoProbably simply making the mistake of caring a lot about true positives and false negatives and caring relatively less about false positives.
- danielmarkbruce 3y agoIt's just as likely bad engineering as political bias. There is an assumption that they should be able to catch up to OpenAI and it's likely unreasonable. They haven't been working on building a real, direct "LLM product" until recently. OpenAI have been in the weeds on it for quite a while.
- kmeisthax 3y agoThey used an LLM, that's how. On the surface level, Google tends to release stuff right away to get feedback, which means you get to see all the bullshit right away. OpenAI carefully manages access to their models, which increases hype, even if that isn't what they intended. Going deeper, a lot of Google's core[0] search business relies on having a healthy information ecosystem. Their search algorithms - e.g. PageRank, TrustRank, etc - use scarcity as a proxy for signals of quality. That's been chipped away at by linkspam and blogspam schemes. Furthermore, social media and even Google's own Knowledge Graph feature have created incentives to pull information out of Google. This decay has happened over decades, and Google fights back against it over time, but it keeps being a problem for them. Now, if I wanted a weapon to Fucking Kill Google[1] with, an LLM would be my go-to. While there are ways to defeat Google's antispam measures, they all leave pretty obvious statistical evidence that can be detected and compensated for. LLMs generate garbage text that is nearly indistinguishable from humans, at extremely low cost, which can be used to Sybil-attack the Google search algorithm basically forever. Ok, but what does that matter for the quality of Google's LLM? Well, the quality of that garbage text depends greatly on both the quality and quantity of the training data fed into it. OpenAI specifically stopped crawling the public Internet for text around the release of GPT-3 for fear of feeding new models the output of prior models. In other words, they have a huge cache of freely obtained "low-background metal[2]" that Google is having to scrounge around for. Furthermore, we have to keep in mind that none of these models are pure representations of the training set. If they were, they wouldn't answer questions or follow directions very well. There's a second, parallel training set that OpenAI had to build to turn GPT-3 into ChatGPT, which isn't crawled and harvested text from the Internet, but instead a list of dos and don'ts that are fine-tuned on after the initial model training is complete. This includes both basic instruction-following, refusing unsafe requests, and political alignment[3]. Google also has to build that second training set itself. Except it's almost certainly less well-developed than OpenAI's. In fact, this is the intent behind OpenAI's really long preview periods. The people using the model in preview are specifically being spied on to find out new corner cases for their models. My guess is that every stupid thing Gemini says or does[4] is something Google never even considered and thus didn't put a training set example in for. [0] to consumers, i.e. not counting adtech [1] https://www.theregister.com/2005/09/05/chair_chucking/ https://www.theregister.com/2005/09/05/chair_chucking/ [2] Steel that has been produced before the first detonation of nuclear weapons. Due to the way in which steel is made, it absorbs trace radioactive isotopes from the oxygen in the air, effectively 'freezing' in the background radiation of the time at which the steel was made. [3] i.e. making the bot not immediately start spitting out racist bullshit like Tay did [4] e.g. assuming that memory safety and child safety are the same thing, drawing ethnically diverse Nazi soldiers
- vkou 3y ago> How did it get this bad? Serious answer? It's an LLM. They don't actually understand anything, they chain words together. C and C++ are often used next to words like 'unsafe' and 'dangerous'. Not to mention that 'concept' isn't far from 'conceive' or 'conception' - something that a lot of people think is unsafe or dangerous for children to, uh, do. There's a trillion weird edge cases that need to be dealt with to avoid pie-on-face moments like these.
- smsm42 3y agoit is true that it doesn't understand anything, but it is trained to do things and to associate things. These chains are not random - they are formed by training. And whoever trained it was so obsessed with "safety" and not letting anything "unsafe" to leak through - probably at explicit command of their higher ups - that they trained the model to have this bias. It's not only their blame - in current American culture, it's always better to be insane safety-obsessed coward, than take any kind of risk of offending anyone. The former gets you a mild derision, maybe, the latter gets you people that would hound you till the end of time, forever, and would think that destroying you is their sacred duty. A lot of such people, that have much more free time and energy to obsess about destroying you than you do. And this behavior is considered normal and socially accepted. So no wonder we get what we train for - not only with LLMs but with our culture.
- deleted 3y ago[deleted]