9 ms·
People are just as bad as my LLMs
- vivzkestrel 2y agoshould have started naming them from person 4579 and see if it still exhibits the bias
- th0ma5 2y agoNo they are not randomly wrong or right without perspective unless they have some kind of brain injury. So that's against the title but the rest of their point is interesting!
- tehsauce 2y agoThere has been some good research published on this topic of how RLHF, ie aligning to human preferences easily introduces mode collapse and bias into models. For example, with a prompt like: "Choose a random number", the base pretrained model can give relatively random answers, but after fine tuning to produce responses humans like, they become very biased towards responding with numbers like "7" or "42".
- aidos 2y agoWhy is that? Whenever I’m giving examples I almost always use 7, something ending in a 7 or something in the 70s
- Ethee 2y agoVeritasium actually made a video on this concept about a year ago: https://www.youtube.com/watch?v=d6iQrh2TK98 https://www.youtube.com/watch?v=d6iQrh2TK98
- d4mi3n 2y agoMy guess is that we bias towards numbers with cultural or personal significance. 7 is lucky in western cultures and is religiously significant (see https://en.wikipedia.org/wiki/7#Culture https://en.wikipedia.org/wiki/7#Culture). 42 is culturally significant in science fiction, though that's a lot more recent. There are probably other examples, but I imagine the mean converges on numbers with multiple cultural touchpoints.
- Jensson 2y agoI have never heard of 7 being a lucky number in western culture and your link doesn't support that. 3 is a lucky number, 13 is an unlucky number, 7 is nothing to me. So I don't think its that, 7 is still a very common "random number" here even though there is no special cultural significance to it.
- kbenson 2y agoIt's definitely used in slot machines as a lucky number. Which came first I'm not sure (but I suspect from a sibling comment in the same thread it's based on perceived commonality and primeness historically and became "lucky" in the past because of that).
- II2II 2y agoWhile I have never heard of someone referring to 7 as a lucky number, 7 is the most common sum of two rolled dice. So I can see how people would regard it as a lucky number. Along the same lines, I assume that someone who mentions 42 as a random number has at least some interest in science fiction.
- dontseethefnord 2y agoYou must be living under a rock if you’ve never heard of 7 as a lucky number.
- II2II 2y agoI prefer to calculate with numbers, and don't pay much attention to superstitions around them. I don't gamble, nor much pay attention to conversations about gambling, so I pretty much ignore any mention of lucky numbers when such topics arise (aside from knowing that some people have lucky numbers). If you refer being isolated from a particular aspect of life living under a rock, so be it. Though I will point out that I like wide open space. I'm more of an astronomer than a geologist!
- 2y ago
- da_chicken 2y agoThe theory I've heard is that the more prime a number is, the more random it feels. 13 feels more awkward and weird, and it doesn't come up naturally as often as 2 or 3 do in everyday life. It's rare, so it must be more random! I'll give you the most random number I can think of! People tend to avoid extremes, too. If you ask for a number between 1 and 10, people tend to pick something in the middle. Somehow, the ordinal values of the range seem less likely. Additionally, people tend to avoid numbers that are in other ranges. Ask for a number from 1 to 100, and it just feels wrong to pick a number between 1 and 10. They asked for a number between 1 and 100. Not this much smaller range. You don't want to give them a number they can't use. There must be a reason they said 100. I wonder if the human RNG would improve if we started asking for numbers between 21 and 114.
- foota 2y agoOkay, this is a nitpick, but I don't think ordinal can be used in that way. "Somehow, the ordinal values of the range seem less likely". I'd probably go with extremes of the range? Or endpoints?
- smohare 2y ago[dead]
- da_chicken 2y agoNope I just mixed up a rephrase. I originally said "ordinal extremes" and meant to say "extreme values". I replaced the wrong word.
- thfuran 2y agoPeople also tend to botch random sequences by trying to avoid repetition or patterns.
- d0liver 2y agoI like prime numbers. Non-primes always feel like they're about to fall apart on me.
- p1necone 2y ago1 and 10 are on the boundary, that's not random so those are out. 5 is exactly halfway, that's not random enough either, that's out. 2, 4, 6, 8 are even and even numbers are round and friendly and comfortable, those are out too. 9 feels too close to the boundary, it's out. That leaves 3 and 7, and 7 is more than 3 so it's got more room for randomness in it right? Therefore 7 is the most random number between 1 and 10.
- LoganDark 2y agoThat's all well and good, but 4 is actually the most random number, because it was chosen by fair dice roll.
- HappMacDonald 2y agoAlso because humans are biased towards viewing prime numbers as more counterintuitive and thus more unpredictable.
- wruza 2y agoLast time I hallway tested it, people couldn’t tell what prime numbers are, and to my surprise even the ones with tech/math-y background forgot it. My results were something 1.5/10 (ages 30+-5) and I didn’t go to cabinets where I knew there are zero chances.
- HappMacDonald 2y agoBut there's a difference between "knowing what the formal definition is" and "having a feeling that a number is somehow unique due to it's indivisibility".
- moffkalast 2y agoIt's very funny that people hold the autoregressive nature of LLMs against them, while being far more hardline autoregressive themselves. It's just not consciously obvious.
- absolutelastone 2y agoI think people tend to just not understand what autoregressive methods are capable of doing generally (i.e., basically anything an alternative method can do), and worse they sort of mentally view it as equivalent to a context length of 1.
- antihipocrat 2y agoI wonder whether we hold LLMs to a different standard because we have a long term reinforced expectation for a computer to produce an exact result? One of my first teachers said to me that a computer won't ever output anything wrong, it will produce a result according to the instructions it was given. LLMs do follow this principle as well, it's just that when we are assessing the quality of output we are incorrectly comparing it to the deterministic alternative, and this isn't really a valid comparison.
- robwwilliams 2y agoI assume 42 is a joke from deep history and The Hitchhiker’s Guide. Pretty amusing to read the Wikipedia entry: https://en.wikipedia.org/wiki/42_(number) https://en.wikipedia.org/wiki/42_(number)
- sedatk 2y agoDouglas Adams picked 42 randomly though. :)
- robertlagrant 2y agoNot at all. It was derived mathematically from the Question: What do you get if you multiply six by nine?
- sedatk 2y agoI stand corrected in base 13.
- eterm 2y agoIt was just a joke, and doubly so the fact it "works" in base 13. It was written as a joke in fairly ramshackle radio play. He had no idea at the time of writing it that the joke would connect so well and become it's own "thing" and dominate discourse of the radio series and novels to come. It's not a joke about numbers, it's a linguistical joke, that works well on radio, something that HHGTG is stuffed full of. https://scifi.stackexchange.com/questions/12229/how-did-douglas-adams-choose-the-ultimate-question https://scifi.stackexchange.com/questions/12229/how-did-doug...
- deleted 2y ago[deleted]
- HappMacDonald 2y agoThat's not the question, though. Everybody knows that the question is the one posed to Mister Turtle and Mister Owl which neither of them can find the answer to.
- thechao 2y agoWhich is weird, because I thought we'd all agreed that the random number was 4? https://xkcd.com/221/ https://xkcd.com/221/
- mynameismon 2y agoCan you share any links about this?
- Shorel 2y agoThey choose 37 =)
- devit 2y agoThe "person one" vs "person two" bias seems trivially solvable by running each pair evaluation twice with each possible labelling and the averaging the scores. Although of course that behavior may be a signal that the model is sort of guessing randomly rather than actually producing a signal.
- harrisonjackson 2y agoAgreed on the second part. Correcting for bias this way might average out the scores but not in a way that correctly evaluates the HN comments. The LLM isn't performing the desired task. It sounds possible to cancel out the comments where reversing the labels swaps the outcome because of bias. That will leave the more "extreme" HN comments that it consistently scored regardless of the label. But that may not solve for the intended task still.
- rahimnathwani 2y agoThe LLM isn't performing the desired task. It's 'not performing the task', in the same way that the humans ranking voice attractiveness are 'not performing the task'. I wouldn't treat the output as complete garbage, just because it's somewhat biased by an irrelevant signal.
- velcrovan 2y agointerleaving a bunch of people's comments and then asking the LLM to sort them out and rank them…seems like a poor method. The whole premise seems silly, actually. I don't think there's any lesson to draw here other than that you need to understand the problem domain in order to get good results from an LLM.
- deleted 2y ago[deleted]
- switch007 2y agoI've had very similar experiences. How quickly and easily people are willing to give up first class sources is quite frightening
- mdp2021 2y agoVery nice article. But the title, and the idea, is the very frequent "racist" form of the proper "People [can be] just as bad as my LLMs". Now: some people can't count. Some people hum between words. Some people set fire to national monuments. Reply: "Yes we knew", and "No, it's not necessary". And: if people could lift the tons, we would not have invented cranes. Very, very often in these pages I meet people repeating "how bad people are". That is "how bad people can be", and "and we would have guessed these pages are especially visited by engineers, who must be already aware of the importance of technical boosts" - so, besides the point relevant to the fact that the median does not represent the whole set, the other point relevant to the fact that tools are not measured on reaching mediocre results.
- th0ma5 2y agoRacist is the wrong word probably maybe ... antisocial in that it is against society.
- consumer451 2y agoMaybe misanthropic?
- mdp2021 2y agoNope. Misanthropic is when some people dislike other people. Racist is when some people attribute some dubious quality to all people in some category.
- geonineties 2y ago(dropped the snark) Racist means grouping according to race, or potentially geographic origins. The word for what you're describing is probably closest to discriminatory, or prejudicial. However, misanthropic is probably more correct as the paper applies to all people negatively.
- mdp2021 2y ago
- andrewmcwatters 2y agoI don’t understand the significance of performing tests like these. To me it’s literally the same as testing one Markov chain against another.
- bxguff 2y agoKind of an odd metric to try to base this process off of. are more comments inherently better? is it responding to buzz words? Makes sense talking about hiring algos / resume scanners in part one and if anything this elucidates some of the trouble with them.
- deleted 2y ago[deleted]
- markbergz 2y agoFor anyone interested in these LLM pairwise sorting problems, check out this paper: https://arxiv.org/abs/2306.17563 https://arxiv.org/abs/2306.17563 The authors discuss the person 1 / doc 1 bias and the need to always evaluate each pair of items twice. If you want to play around with this method there is a nice python tool here: https://github.com/vagos/llm-sort https://github.com/vagos/llm-sort
- fpgaminer 2y agoThe paper basically sums to suggesting (and analyzing) these otpions: * Comparing all possible pair permutations eliminates any bias since all pairs are compared both ways, but is exceedingly computationally expensive. * Using a sorting algorithm such as Quicksort and Heapsort is more computationally efficient, and in practice doesn't seem to suffer much from bias. * Sliding window sorting has the lowest computation requirement, but is mildly biased. The paper doesn't seem to do any exploration of the prompt and whether it has any impact on the input ordering bias. I think that would be nice to know. Maybe assigning the options random names instead of ordinals would reduce the bias. That said, I doubt there's some magic prompt that will reduce the bias to 0. So we're definitely stuck with the options above until the LLM itself gets debiased correctly.
- smallnix 2y agoIs my understanding wrong that LLMs are trained to emulate observed human behavior in their training data? From that follows that LLMs fit to produce all kinds of human biases. Like preferring the first choice out of many, and the last our of many (primacy biases). Funnily the LLM might replicate the biases slightly wrong and by doing so produce new derived biases.
- mplewis 2y agoLLMs don't emulate human behavior. They spit out chunks of words in an order that parrots some of their training data.
- educasean 2y agoIs this just pedantry or is there some insight to be gleaned by the distinction you made?
- MyOutfitIsVague 2y agoI can only assume that either they are trying to point out that words aren't behavior, and mimicking human writing isn't the same thing as mimicking human behavior, or it's some pot-shot at the capabilities of LLMs.
- Xelynega 2y agoIt's not really pedantic when there's an entire wikipedia page on the tendency for people to conflate the two: https://en.wikipedia.org/wiki/ELIZA_effect https://en.wikipedia.org/wiki/ELIZA_effect I believe the distinction they're trying to make is between "sounding like a human"(being able to create output that we understand as language) and "thinking like a human"(having the capacity for experience, empathy, semantic comprehension, etc.)
- tavavex 2y agoBut nowhere in the original post was there a mention of "thinking like a human". The poster said that these systems, at their core, emulate human behaviors. Writing is a human behavior - as are all the things that are requirements for writing (operating within the rules of given language/s, making logical choices, etc). The things you listed as evidence of human thinking were never implied when talking about replicating what human writing is.
- jopsen 2y agoBut an LLM can't be held accountable.. neither can most employees, but we often forget that :)
- leptons 2y agoEmployees get fired all the time, and the more wrong answers I get from an LLM, the less I use it until it's never used again.
- malfist 2y agoJokes on you, my org at work just adopted KPIs about how many AI suggestions engineers accept
- rainsford 2y agoBut an LLM doesn't understand "never used again" as a consequence and the threat of it is useless as a motivation to improve (also because LLMs have no concept of "motivation" or "threats" or anything else).
- leptons 2y agoYou're talking about LLMs as if they are some kind of singular entity, but LLMs as used for coding only exist as a product of a company that employs humans. If nobody uses the LLM because it sucks, those people will be out of a job.
- jayd16 2y agoYou can certainly send people to jail for negligence.
- simne 2y ago> But an LLM can't be held accountable.. neither can most employees Yes and no. Yes, this is really problem, because at current level of technologies, some thing are inexpensive only if done in large numbers (factor of scale), so for example, just could not exist one person who could be accountable for machine like Boeing-747 (~500 human-years of work per plane). Unfortunately, modern automobile is considered large system, made from thousands parts, so again, not exist one person to know everything. And no, Germans said "Ordnung muss sein", which in modern management mean, constant clear organization of the game of the whole team is more important than the success of individual players. Or, in simple words, right organization, controlled by rules is considered enough reliable to be accountable. And for example in automobile industry, now normal to consider accountable whole organization. And for example, Daimler officials few years ago said, Daimler safety systems will use Daimler view on robotic laws - priority will be safety of people inside vehicle. You may know, traditionally used Lem robotic laws, which have totally different view, separated from inside vs outside approach. In civil aviation using approach, to just use simple designs or design with evidence of reliability. Sure, government regulators could decide something even more original, will see. Any way, as technology emerge, accountability of machines will be sure subject of many discussions.
- deleted 2y ago[deleted]
- animanoir 2y ago[dead]
- K0balt 2y agoI know this is only adjacent to OP’s point, but I do find it somewhat ironic that it is easy to find people who are just as unreliable and incompetent at answering questions correctly as a 7b model, but also a lot less knowledgeable. Also, often less capable of carrying on a decent conversation. I’ve noticed an periconcious urge when talking to people to judge them against various models and quants, or to decide they are truly SOTA. I need to touch grass a bit more, I think.
- satisfice 2y agoPeople are alive. They have rights and responsibilities. They can be held accountable. They are not "just as bad" as your LLMs.
- icelancer 2y ago> They can be held accountable Is this a universal phenomenon where you've worked? Consider yourself very lucky.
- lxe 2y agoIt's almost as if we trained LLMs on text produced by people.
- MrMcCall 2y agoI love the posters that make fun of those corporate motivational posters. My favorite is: No one is as dumb as all of us. And they trained their PI* on that giant turd pile. * Pseudo Intelligence
- LoganDark 2y agoI don't count LLMs as intelligent. To a certain degree they can be a component of intelligence, but I don't count an LLM on its own.
- tavavex 2y agoArtificial intelligence is a generic term for a very broad field that has existed for like 50-70 years, depending on who you ask. 'Intelligence' isn't praise or endorsement. I think it's a succinct word that does the job at explaining what the goal here is. All the "Artificial intelligence? Hah, more like Bad Unintelligence, am I right???" takes just sound so corny to me.
- callc 2y ago> I think it's a succinct word that does the job at explaining what the goal here is. Sure. If the goal is intelligence then LLMs fail. LLMs do not currently have the same intelligence as humans. If a human being in front of me were to answer my question like an LLM does, I would think they are an overly confident parrot. Not saying LLMs are bad, they are an incredible tool. Just not intelligence. Words matter.
- tavavex 2y agoWho said anything about matching human intelligence though? If that's the reference point, then nothing in the field of AI has ever had the 'right' to be called that. Computer vision, ML-based optimization of anything or ranking/recommendation systems are all considered to be within AI, despite none of them being remotely similar to 'human intelligence'. This is the main point of my post - I feel like people retroactively try to see AI as being some kind of an endorsement term, or having to do anything regarding humans - or that 'intelligence' is in itself an endorsement and something so extremely good that only humans can be bestowed with it. In reality, these comparisons only appeared after the boom of generative AI and would've been seen as ludicrous by any AI researchers prior to it.
- soared 2y agoWouldn’t the same outcome be achieved much more simply by giving LLMs a two choices (colors, numbers, whatever), asking “pick one” and assessing the results in the same way?
- ramity 2y agoYou absolutely can. Deterministic inference is achievable, but it isn't as performant. The reason why sadly boils down to floating point math.
- le-mark 2y agoHuman level artificial intelligence has never had much appeal to me, there are enough idiots in the world, why do we need artificial ones? Ie if average machine intelligence mirrored human IQ distribution?
- roywiggins 2y agoOwners would love to be able to convert capital directly into products without any intermediate labor[0]. Fire your buildings full of programmers and replace them with a server farm that only gets faster and more efficient over time? That's a great position to be in, if you own the IP and/or server farm. [0] https://qntm.org/mmacevedo https://qntm.org/mmacevedo
- megadata 2y agoAt least LLMs are very often ready so acknowledge they might be wrong. It can be incredibly hard to get a person to acknowledge that they might be remotely wrong on a topic they really care about. Or, for some people, the thought that they might be wrong about anything attall is just like blasphemy to them.
- Xelynega 2y agoIs this not just because aggressive material was filtered out of training data and the system prompts usually include some preamble about being polite? "Acknowledging they might be wrong" makes them sound like more than token predictors trained on polite sounding text.
- roywiggins 2y agoMost of the reason LLMs will "admit they're wrong" is because they've been trained not to argue too hard, and to not hold strong preferences. It's a sort of customer service personality. When you don't do that sufficiently you run the risk of producing the "Sydney" personality that Bing Chat had, which would argue back, and could go totally feral defending its incorrect beliefs about the world, to the point of insulting and belittling the user.
- rainsford 2y ago> ...a lot of the safeguards and policy we have to manage humans own unreliability may serve us well in managing the unreliability of AI systems too. It seems like an incredibly bad outcome if we accept "AI" that's fundamentally flawed in a way similar to if not worse than humans and try to work around it rather than relegating it to unimportant tasks while we work towards a standard of intelligence we'd otherwise expect from a computer. LLMs certainly appear to be the closest to real AI that we've gotten so far. But I think a lot of that is due to the human bias that language is a sign of intelligence and our measuring stick is unsuited to evaluate software specifically designed to mimic the human ability to string words together. We now have the unreliability of human language processes without most of the benefits that comes from actual human level intelligence. Managing that unreliability with systems designed for humans bakes in all the downsides without further pursuing the potential upsides from legitimate computer intelligence.
- smohare 2y ago[dead]
- itchyjunk 2y agoWhat is your measure of intelligence?
- rainsford 2y agoI honestly don't have a great one, which is less worrying than it might otherwise be since I'm not sure anyone else does either. But in a human context, I think intelligence requires some degree of creativity, self-motivation, and improvement through feedback. Put a bunch of humans on an island with various objects and the means for survival and they're going to do...something. Over enough time they're likely to do a lot of unpredictable somethings and turn coconuts into rocket ships or whatever. Put a bunch of LLMs on an equivalent island with equivalent ability to work with their environment and they're going to do precisely nothing at all. On the computer side of things, I think at a minimum I'd want intelligence capable of taking advantage of the fact that it's a deterministic machine capable of unerringly performing various operations with perfect accuracy absent a stray cosmic ray or programming bug. Star Trek's Data struggled with human emotions and things like that, but at least he typically got the warp core calculations correct. Accepting LLMs with the accuracy of a particularly lazy intern feels like it misses the point of computers entirely.
- jayd16 2y agoIf the question inherently allows for "no-preference" to be valid but that is not a possible answer then you've left it to the person or llm to deal with that. If a human is not allowed to specify no preference why would you expect uniform results when you don't even ask for it? You only asked to pick the best. Even if they picked perfectly, its not defined in the task to make sure you select draws in a random way.
- raincole 2y agoWhat a clickbait title. TL;DR: the author found a very, very specific bias that is prevalent in both humans and LLMs. That is it.
- henlobenlo 2y agoThis is the "anyone can be a mathematician meme". People who hang around elite circles have no idea how dumb the average human is. The average human hallucinates constantly.
- bawolff 2y agoSo if you give a bunch of people a boring task, pay them the same regardless of if they treat it seriously or not - the end result is they do a bad job! Hardly a shocker. I think this say more about the experimental design then it does about AI & humans.
- isaacremuant 2y agoSo many articles like this HN have a catchy title and then a short article that doesn't really conclude the title. The experiment itself is so fundamentally flawed it's hard to begin criticizing it. HN comments as a predictor of good hiring material is just as valid as social media profile artifacts or sleep patterns. Just because you produce something with statistics (with or without LLMs) and have nice visuals and narratives doesn't mean is valid or rigorous or "better than nothing" for decision making. Articles like this keep making it to the top of HN because HN is behaving like reddit where the article is read by few and the gist of the title debated by many.
- djaouen 2y agoYes, but a consensus of people beats LLMs every time. For now, at least.
- oldherl 2y agoIt's just because people tend to put the "original" result in the first place and the "improved" result in the second place in many scientific studies. LLM and humans are learning that and assume that the second one is the better one.