16 ms·
LLMs are steroids for your Dunning-Kruger
- Brendinooo 11mo ago>I think LLMs should not be seen as knowledge engines but as confidence engines. This is a good line, and I think it tempers the "not just misinformed, but misinformed with conviction" observation quite a bit, because sometimes moving forward with an idea at less than 100% accuracy will still bring the best outcome. Obviously that's a less than ideal thing to say, but imo (and in my experience as the former gifted student who struggles to ship) intelligent people tend to underestimate the importance of doing stuff with confidence.
- shermantanktop 11mo agoConfidence has multiple benefits. But one of those benefits is social - appearing confident triggers others to trust you, even when they shouldn’t. Seeing others get burned by that pattern over and over can encourage hesitation and humility, and discourage confident action. It’s essentially an academic attitude and can be very unfortunate and self-defeating.
- Chabsff 11mo ago> I feel like LLMs are a fairly boring technology. They are stochastic black boxes. The training is essentially run-of-the-mill statistical inference. There are some more recent innovations on software/hardware-level, but these are not LLM-specific really. This is pretty ironic, considering the subject matter of that blog post. It's a super-common misconception that's gained very wide popularity due to reactionary (and, imo, rather poor) popular science reporting. The author parroting that with confidence in a post about Dunner-Krugering gives me a bit of a chuckle.
- miningape 11mo agoI also find it hard to get excited about black boxes - imo there's no real meat to the insights they give, only the shell of a "correct" answer
- yannyu 11mo agoWhat's the misconception? LLMs are probabilistic next-token prediction based on current context, right?
- Chabsff 11mo agoYeah, but that's their interface. That informs surprisingly little about their inner workings. ANNs are arbitrary function approximators. The training process uses statistical methods to identify a set of parameters that approximate the function as best as possible. That doesn't necessarily mean that the end result is equivalent to a very fancy multi-stage linear regression. It's a possible outcome of the process, but it's not the only possible outcome. Looking at a LLMs I/O structure and training process is not enough to conclude much of anything. And that's the misconception.
- yannyu 11mo ago> Yeah, but that's their interface. That informs surprisingly little about their inner workings. I'm not sure I follow. LLMs are probabilistic next-token prediction based on current context, that is a factual, foundational statement about the technology that runs all LLMs today. We can ascribe other things to that, such as reasoning or knowledge or agency, but that doesn't change how they work. Their fundamental architecture is well understood, even if we allow for the idea that maybe there are some emergent behaviors that we haven't described completely. > It's a possible outcome of the process, but it's not the only possible outcome. Again, you can ascribe these other things to it, but to say that these external descriptions of outputs call into question the architecture that runs these LLMs is a strange thing to say. > Looking at a LLMs I/O structure and training process is not enough to conclude much of anything. And that's the misconception. I don't see how that's a misconception. We evaluate all pretty much everything by inputs and outputs. And we use those to infer internal state. Because that's all we're capable of in the real world.
- kmijyiyxfbklao 11mo agoThen why not say "they are just computer programs"? I think the reason people don't say that is because they want to say "I already understand what they are, and I'm not impressed and it's nothing new". But what the comment you are replying to is saying is that the inner workings are the important innovative stuff.
- parineum 11mo agoI'm not sure what claim your disputing or making with this. What more are LLMs than statistical inference machines? I don't know that I'd assert that's all they are with confidence but all the configurations options I can play with during generation (Top K, Top P, Temperature, etc.) are all ways to _not_ select the most likely next token which leads me to believe that they are, in fact, just statistical inference machines.
- ACCount37 11mo agoWhat more are human brains than piles of wet meat? It's not an argument - it's a dismissal. It's boneheaded refusal to think on the matter in any depth, or consider any of the implications. The main reason to say "LLMs are just next token predictions" is to stop thinking about all the inconvenient things. Things like "how the fuck does training on piles of text make machines that can write new short stories" or "why is a big fat pile of matrix multiplications better at solving unseen math problems than I am".
- zahlman 11mo ago> What more are human brains than piles of wet meat? Calculation isn't what makes us special; that's down to things like consciousness, self-awareness and volition. > The main reason to say "LLMs are just next token predictions" is to stop thinking about all the inconvenient things. Things like... They do it by iteratively predicting the next token. Suppose the calculations to do a more detailed analysis were tractable. Why should we expect the result to be any more insightful? It would not make the computer conscious, self-aware or motivated. For the same reason that conventional programs do not.
- ACCount37 11mo agoDo you have, by chance, a set of benchmarks that could be administered to humans and LLMs both, and used to measure and compare the levels of "consciousness, self-awareness and volition" in them? Because if not, it's worthless philosophical drivel. If it can't be defined, let alone measured, then it might as well not exist. What is measurable and does exist: performance on specific tasks. And the pool of tasks where humans confidently outperform LLMs is both finite and ever diminishing. That doesn't bode well for human intelligence being unique or exceptional in any way.
- LeroyRaz 11mo agoHow is that a misconception? LLMs are just advanced statistical modelling (unsupervised machine learning) with small tweaks (e.g., some fine-tuning for human preference). At the core, they are just statistical modelling. The fact that statistical modelling can produce coherent thoughts is impressive (and basically vindicates materialism) but that doesn't change the fact it is all based on statistical modelling. ...? What is your view?
- sho_hn 11mo agoI'm not sure this is something I really worry about. Whenever I use an LLM I feel dumber, not smarter; there's a sensation of relying on a crutch instead of having done the due diligence of learning something myself. I'm less confident in the knowledge and less likely to present it as such. Is anyone really cocksure on the basis of LLM received knowledge? > As I ChatGPT user I notice that I’m often left with a sense of certainty. They have almost the opposite effect on me. Even with knowledge from books or articles I've learned to multi-source and question things, and my mind treats the LLMs as a less reliable averaging of sources.
- deadbabe 11mo agoIf you feel dumber, it’s because you’re using the LLM to do raw work instead of using it for research. It should be a google/stackoverflow replacement, not a really powerful intellisense. You should feel no dumber than using google to investigate questions.
- Insanity 11mo agoI don't think this is entirely accurate. If you look at this: https://www.media.mit.edu/publications/your-brain-on-chatgpt/ https://www.media.mit.edu/publications/your-brain-on-chatgpt..., it shows that search engines do engage your brain _more_ than LLM usage. So you'll remember more through search engine use (and crawling the web 'manually') than by just prompting a chatbot.
- pessimizer 11mo agoI find that it is terrible for research, and hallucinates 25% to 90% of its references. If you tell it to find something and give it a detailed description of what you're looking for, it will pretend like it has verified that that thing exists, and give you a bulletpoint lecture about why it is such an effective and interesting thing that 1) you didn't ask for, and 2) is really it parroting your description back to you with embellishments. I thought I was going to be able to use LLMs primarily for research, because I have read an enormous number of things (books, papers) in my life, and I can't necessarily find them again when they would be useful. Trying to track them down through LLMs is rarely successful and always agonizing, like pulling teeth that are constantly lying to you. A surprising outcome is that I often get so frustrated by the LLM and so detailed in how I'm complaining about its stupid responses that I remind myself of something that allows me to find the reference on my own. I have to suspect that people who find it useful for research are researching things that are easily discoverable through many other means. Those are not the things that are interesting. I totally find it useful to find something in software docs that I'm too lazy to look up myself, but it's literally saving me 10 minutes.
- cachius 11mo agoFrom the title I thought this was a repost of 'AI is Dunning-Kruger as a service ' https://news.ycombinator.com/item?id=45851483 https://news.ycombinator.com/item?id=45851483 It is not.
- gopheryourshelf 11mo ago>“the problem with the world is that the stupid are cocksure, while the intelligent are full of doubt.” Is it me or does everyone find that dumb people seem to use this statement more than ever?
- stevenwoo 11mo agoIt appears to be a paraphrasing of William Butler Yeats https://en.wikipedia.org/wiki/The_Second_Coming_(poem) https://en.wikipedia.org/wiki/The_Second_Coming_(poem)
- 0xdeadbeefbabe 11mo agoUgh. You can be cocksure of your doubts. It's still confidence, duh.
- aeve890 11mo agoEveryone thinks they're the intelligent ones, of course. Which reinforces the repetition ad nauseam of Dunning Kruger. Which is on itself dumb AF because the effect described by Dunning and Kruger has been repeatedly exaggerated and misinterpreted. Which in turn is even dumber because Dunning-Kruger effect is debatable and reproducibility is weak at best.
- vehemenz 11mo agoI hate to comment on just a headline—thought I did read the article—but it's wrong enough to warrant correcting. This is not what the Dunning-Kruger effect is. It's lacking metacognitive ability to understand one's own skill level. Overconfidence resulting from ignorance isn't the same thing. Joe Rogan propagated the version of this phenomenon that infiltrated public consciousness, and we've been stuck with it ever since. Ironically, you can plug this story into your favorite LLM, and it will tell you the same thing. And, also ironically, the LLM will generally know more than you in most contexts, so anyone with a degree epistemic humility is better served taking it at least as seriously as their own thoughts and intuitions, if not at face value.
- topaz0 11mo agoFunnily enough, DK is also not real -- just a statistical artifact of a poorly chosen analysis.
- dktoao 11mo agoOk, I think you are going to need to explain to me why "Overconfidence resulting from ignorance" isn't exactly the same thing as "lacking metacognitive ability to understand one's own skill level". Just worded more simply
- lukev 11mo agoI very much agree. I've been telling folks in trainings that I do that the term "artificial intelligence" is a cognitohazard, in that it pre-consciously steers you to conceptualize a LLM as an entity. LLMs are cool and useful technology, but if you approach them with the attitude you're talking with an other, you are leaving yourself vulnerable to all sorts of cognitive distortions.
- roywiggins 11mo agoIt certainly isn't helped by the RLHF and chat interface encouraging this. LLM providers have every incentive to make their users engage it like an other. It was much harder to accidentally do when it was just a completion UI and not designed to roleplay as a person.
- cgriswald 11mo agoI don't think that is actually a problem. For decades people have believed that computers can't be wrong. Why, now, suddenly, would it be worse if they believed the computer wasn't a computer? The larger problem is cognitive offloading. The people for whom this is a problem were already not doing the cognitive work of verifying facts and forming their own opinions. Maybe they watched the news, read a Wikipedia article, or listened to a TEDtalk, but the results are the same: an opinion they felt confident in without a verified basis. To the extent this is on 'steroids', it is because they see it as an expert (in everything) computer and because it is so much faster than watching a TED talk or reading a long form article.
- roywiggins 11mo agoIt can also dispense agreeable confirmation on tap, with very little friction and hardly any chance of accidentally encountering something unexpected or challenging. Even TED talks occasionally have a point of view that isn't perfectly crafted for each hearer.
- mmaunder 11mo agoUse an agent to create something with a non-negotiable outcome. Eg software that does something useful, or fails to, in a language you don’t program in. This is a helpful way to calibrate your own understanding of what LLMs are capable of.
- AndrewKemendo 11mo agoHumans broadly have a tenuous grasp of “reality” and “truth.” Propagandists, spies and marketers know what philosophers of mind prove all too well: most humans do not perceive or interact with reality as it is, rather their perception of it as it contributes or contradicts their desired future. Provide a person confidence in their opinion and they will not challenge it, as that would risk the reward of lend you live in a coherent universe. The majority person has never heard the term “epistemology” despite the concept being central to how people derive coherence. Yet all these trite pieces written about AI and its intersectionality with knowledge claim some important technical distinction. I’m hopeful that a crisis of epistemology is coming, though that’s probably too hopeful. I’m just enjoying the circus at this point
- LeroyRaz 11mo agoHacker News readers, and especially commenters, are number 1!
- balderdash 11mo agoI ascribe the effect of LLMs as similar to reading the newspaper, when I learn about something I have no knowledge base in I come away feeling like I learned a lot. When I interact with a newspaper or LLM in an area where I have real domain expertise I realize they don’t know what they are talking about - which is concerning about the information I get from them about topics I don’t have that high level of domain expertise.
- moffkalast 11mo agoAnd why stop at newspapers, it's been a while since one could say books have any integrity, pretty much anyone can get anything into print these days. From political shenanigans to self help books designed to confirm people's biases to sell more units. Video's by far the hardest to fake but that's changing as well. Regardless of what media you get your info from you have to be selective of what sources you trust. It's more true today than ever before, because the bar for creating content has never been lower.
- Night_Thastus 11mo agoThe problem is that LLM output is so incredibly confident in tone. It really sounds like you're talking to an expert who has years of experience and has done the research for you - and tech companies push this angle quite hard. That's bad when their output can be complete garbage at times.
- AlienRobot 11mo agoIt makes me really sad how Google pushes this technology that is simply flat out wrong sometimes. I forgot what exactly I searched for, but I searched for a color model that Krita supports hoping to get the online documentation as the first result and the under several Youtube thumbnails the AI overview was telling me that Krita doesn't support that color model and you need a plugin for that. Under the AI overview was the search result I was looking for about that color model in Krita. And worse of all is that it's not even consistent, because I tried the same searches again and I couldn't get the same answer, so it just randomly decides to assert complete nonsense sometimes while other times it gives the right answer or says something completely unrelated. It's really been a major negative in my search experience. Every time I search for something I can't be sure that it's actually quoting anything verbatim, so I need to check the sources anyway. Except it's much harder to find the link to the source with these AI's than it is to just browse the verbatim snippets in a simple list of search results. So it's just occupying space with something that is simply less convenient.
- avree 11mo agoThe title makes this incomprehensible. The author seemingly defines Dunning-Kruger as the... opposite of the Dunning-Kruger effect.
- gowld 11mo agoThe "Dunning-Kruger Effect" Effect: A reference to Dunning-Kruger Effect is almost certainly incorrect.
- kraftman 11mo agoI feel like when I talk to someone and they tell me a fact, that fact goes into a kind of holding space, where I apply a filter of 'who is this person that is telling me this thing to know what the thing they are telling me is'. There's how well I know them, there's the other beleifs I know they have, there's their professional experience and their personal experience. That fact then gets marked as 'probably a true fact' or 'mark beleives in aliens'. When I use chatGPT I do the same before I've asked for the fact: how common is this problem? how well known is it? How likely is that chatgpt both knows it and can surface it? Afterwards I don't feel like I know something, I feel like I've got a faster broad idea of what facts might exist and where to look for them, a good set of things to investigate, etc.
- giraffe_lady 11mo agoThe important part of this is the "I feel like" bit. There's a fair but growing bit of research that the "fact" is more durable in your memory than the context, and over time, across a lot of information, you will lose some of the mappings and integrate things you "know" to be false into model of the world. This more closely fits our models of cognition anyway. There is nothing really very like a filter in the human mind, though there are things that feel like them.
- kraftman 11mo agoMaybe but then thats the same wether I talk to chatGPT or a human isnt it? except with chatgpt i instantly verify what im looking for, whereas with a human i cant do that.
- giraffe_lady 11mo agoI wouldn't assume that it's the same, no. For all we knock them unconscious biases seem to get a lot of work done, we do all know real things that we learned from other unreliable humans, somehow. Not a perfect process at all but one we are experienced at and have lifetimes of intuition for. The fact that LLMs seem like people but aren't, specifically have a lot of the signals of a reliable source in some ways, I'm not sure how these processes will map. I'm skeptical of anyone who is confident about it in either way, in fact.
- chaostheory 11mo agoThere are so many guardrails now that are being improved daily. This blog post is a year out of date. Not to mention that people know how to prompt better these days. To make his point, you need specific examples from specific LLMs.
- jakubmazanec 11mo agoIt's possible that the Dunning-Kruger effect is not real, only a measurement or statistical artefact [1]. So it probably needs more and better studies. [1] https://www.mcgill.ca/oss/article/critical-thinking/dunning-kruger-effect-probably-not-real https://www.mcgill.ca/oss/article/critical-thinking/dunning-...
- travisgriggs 11mo ago8 months or so ago, my quip regarding LLMs was “stochastic parrot.” The term I’ve been using of late is “authority simulator.” My formative experiences with “authority figures” was a person who can speak with breadth and depth about a subject and who seems to have internalized it because they can answer quickly and thoroughly. Because LLMs do this so well, it’s really easy to feel like you’re talking to an authority in a subject. And even though my brain intellectually knows this isn’t true, emotionally, the simulation of authority is comforting.
- GMoromisato 11mo agoSpeaking of uncertainty, I wish more people would accept their uncertainty with regards to the future of LLMs rather than dash off yet another cocksure article about how LLMs are {X}, and therefore {completely useless}|{world-changing}. Quantity has a quality of its own. The first chess engine to beat Gary Kasparov wasn't fundamentally different than earlier ones--it just had a lot more compute power. The original Google algorithm was trivial: rank web pages by incoming links--its superhuman power at giving us answers ("I'm feeling lucky") was/is entirely due to a massive trove of data. And remember all the articles about how unreliable Wikipedia was? How can you trust something when anyone can edit a page? But again, the power of quantity--thousands or millions of eyeballs identifying errors--swamped any simple attacks. Yes, LLMs are literally just matmul. How can anything useful, much less intelligent, emerge from multiplying numbers really fast? But then again, how can anything intelligent emerge from a wet mass of brain cells? After all, we're just meat. How can meat think?
- svieira 11mo ago> How can meat think? Some of us used to think that meat spontaneously generated flies. Maybe someday we'll (re-)learn that meat doesn't spontaneously generate thought either?
- ACCount37 11mo agoI don't give much merit to ideas that demand the existence of Magic Fairy Dust. And especially not now. Not when LLMs can already do pretty much anything that a human can - and some of those things they can even do well.
- svieira 11mo agoGiven that everything the LLM can do it learned from human descriptions of the space ... one would have to posit a very inefficient language for that model not to do something with those billions of parameters. But when you fly because of a bunch of balloons sprinkled with magic fairy dust are pulling you up, the magic fairy dust is still at work.
- phamson02 11mo agoI partly share the author's point that ChatGPT users (myself included) can "walk away not just misinformed, but misinformed with conviction". Sometimes I want to criticise aloud, write a post blaming this technology for those colourful, sophisticated, yet empty bullshits I hear from a colleague or read in an online post. But I always resist the urge. Because I think: Isn't it always going to have some kinds of people like that? With or without this LLM thing. If there is anything to hate about this technology, for the more and more bullshits we see/hear in daily life, it is: (1) Its reach: More people of all ages, of different backgrounds, expertise, and intents are using it. Some are heavily misusing it. (2) Its (ever increasing) capability: Yes, it has already become pretty easy for ChatGPT or any other LLMs to produce a sophisticated but wrong answer on a difficult topic. And I think the trend is that with later, more advanced versions, it would become harder and take more effort to spot a hidden failure lurking in a more information-dense LLM's answer.
- bryanlarsen 11mo agoMy opinion: if LLM's speed you up, you're doing it wrong. You have to carefully review and audit every line that comes out of an LLM. You have to spend a lot of time forcing LLM's to prove that the code it wrote is correct. You should be nit-picking everything. Despite, LLM's are useful. I could write the code faster without an LLM, but then I'd have code that wasn't carefully reviewed line-by-line because my coworkers trust me (the fools). It'd have far fewer tests because nobody forced me to prove everything. It'd have worse naming because every once in a while the LLM does that better than me. It'll be missing a few edge cases the LLM thought of that I didn't. It'd have forest/trees problems because if I was writing the code I'd be focused on the code instead of the big picture.
- nzach 11mo ago> You have to carefully review and audit every line that comes out of an LLM. You have to spend a lot of time forcing LLM's to prove that the code it wrote is correct. You should be nit-picking everything. I'm not sure this statement is true most of the time. This kind of reasoning reminds me of the discussion around 'code correctness'. In my opinion there are very few instances where correctness is really important. Most of the time you just need something that works well enough. Imagine you have a continuous numeric scale that goes from 'never works' to '100% formal proofs' to indicate the correctness of every piece of software. Pushing your code to the '100% formal proofs' side takes a lot of resources, that could be deployed on other places.
- bryanlarsen 11mo agoAt least for us, every bug that makes it into a release that gets installed on a client computer costs us 100x - 1000x as much as a bug that gets caught earlier.
- Kiro 11mo agoMost code is not critical like that. A lot of the stuff I write has very little impact if things go wrong and it's easy to tell if it's incorrect.
- pants2 11mo agoI've seen this! Following some Math and Physics subreddits it's a regular occurrence for a new submitter to come in and post some 40 pages of incomprehensible bullshit and claim that they developed a unifying theory of physics with ChatGPT and that ChatGPT has told them it's a breakthrough in the field. Of course that used to happen regularly before LLMs but not nearly as often.
- turtletontine 11mo agoIncluding the former CEO of Uber. I’m somewhat curious what these people even think they’ve discovered, what outstanding problem they think they’ve actually solved… but I’m not curious enough to actually dig through their slop. https://gizmodo.com/billionaires-convince-themselves-ai-is-close-to-making-new-scientific-discoveries-2000629060 https://gizmodo.com/billionaires-convince-themselves-ai-is-c...
- danparsonson 11mo ago"Vibe physics" good lord... It's like reading the thoughts of a five year old - absolutely certain that they know how the world works with little or no basis in reality.
- ramraj07 11mo agoHey! you called me out: https://news.ycombinator.com/item?id=35178196 https://news.ycombinator.com/item?id=35178196
- simianwords 11mo ago>How often do you think a ChatGPT user walks away not just misinformed, but misinformed with conviction? I would bet this happens all the time. And I can’t help but wonder what the effects are in the big picture. this is so wrong! i simply can't get ChatGPT to admit something clearly wrong. it can play both sides and gives nuance which is exactly what i expect. but it is so un-sycopanthic that it won't leave you feeling like you are right. any examples of it doing so are welcome! show me examples where it takes a clearly wrong or false idea and makes it look as if it is a good idea (unless you specifically ask it to do it).
- uoaei 11mo agoFreely available online information is very often educationally incredibly shallow and commonly oversimplified to the point of being wrong. So of course an agent trained on it would be, too.
- zkmon 11mo ago>> How often do you think a ChatGPT user walks away not just misinformed, but misinformed with conviction? I would bet this happens all the time. Why is the "Dunning-Kruger" is not mentioned anywhere in the article body while is gloriously visible in the title? By the way, AI is not wrong "all the time".
- rwolf 11mo ago"all the time" in this context means frequently, not constantly, and is usually understood as a global rate. for example: people fall in love all the time.
- zkmon 11mo agoI don't think you get "frequently" misinformed by ChatGPT. What kind of prompts you are using, that make ChatGPT to give false info very often?
- yannyu 11mo agoIt's remarkably easy to get ChatGPT to hallucinate. A go-to for me is to ask it about factual things as compared to other factual things. The description of the Bluey episode is wrong in nearly every possible way. Additionally, the response continues on analyzing in multiple ways, and again continues to get the fundamental facts and therefore the analysis wrong. Prompt: Compare the Space episode of Bluey to Interstellar Response: That’s a really fun and surprisingly fitting comparison! Let’s break down how Bluey’s “Space” episode parallels Christopher Nolan’s Interstellar — both emotionally and thematically — even though one is a 7-minute preschool cartoon and the other a 3-hour sci-fi epic. 1. Core Theme: Love Across Distance Interstellar: The film centers on a father (Cooper) who leaves his daughter (Murph) to save humanity, exploring how love transcends time and space — literally. The emotional heart of Interstellar is the bond between parent and child that endures despite impossible separation. Bluey: “Space” In this episode, Bluey, Bingo, and their dad (Bandit) play an imaginative game where they’re astronauts who have to separate to save their mission. The moment Bluey and Bandit must say goodbye echoes that same emotional tension — the pain of letting go but also the love that motivates sacrifice. Parallel: Both stories use the setting of outer space as a metaphor for emotional distance — how people we love can feel far away, and how connection can bridge that gap.
- resters 11mo agothey do not have to be. People who seek an idea bubble end up finding one.
- vanillax 11mo agoI'll take big - enterprise AI written code, over offshore any day of the week. I for one welcome AI for this exact use case.
- djoldman 11mo agoJust a reminder that the "Dunning-Kruger effect" is probably not real [0]. It makes sense to refer to it as a concept but it's probably not an appropriate assumption to make about people. [0] https://www.mcgill.ca/oss/article/critical-thinking/dunning-kruger-effect-probably-not-real https://www.mcgill.ca/oss/article/critical-thinking/dunning-...
- deleted 11mo ago[deleted]
- jayd16 11mo agoThis seems to move the idea that you might not understand how skilled you are to some kind of law that ties humility to knowledge more strictly. Maybe this is my misunderstanding but I don't think the common invocation really took it as a law that the unknowledgeable always think their skills are higher.
- nis0s 11mo agoThere’s a gap that LLMs are trying to fill in such cases, which is that there’s too much information that we can possibly hope to make sense of in a lifetime. Just as it’s possible to compute something incorrectly with a calculator, you can definitely be led astray by an LLM, which is why I am surprised that people think these models are good enough to replace humans at work. The only thing which makes sense is to both raise the bar for publishing, and to only take published works seriously. If something isn’t published, then authors should provide code to demonstrate the effect they’re describing.
- thewebguyd 11mo ago> which is why I am surprised that people think these models are good enough to replace humans at work. There are a lot of office jobs that I'd fit into the category of "bullshit jobs." They may serve some purpose in the huge bureaucracy of enterprises but the day to day ultimately boils doing to managing someone's calendar and sending emails. Quite a few people at my work have now started using Copilot for their emails. It's obviously AI (at least to me), and yet, the content and formatting are an improvement over what they were sending before. So much of the marketing hype on LLMs is about how it'll replace all the engineering work (the MBA's wet dream, to replace all the expensive labor). In reality, I think its more capable at replacing non-tech labor and middle management. An LLM can send out an email to the team and analyze a project check-in faster, and better, than some overpaid middle manager can. I have no doubts an LLM could probably serve the role of a project management office, or a business analyst. Sure, there should still be a human in the loop for now, but you need far, far less humans in those roles than previously.
- nis0s 11mo agoI go back and forth on the idea that some jobs are bullshit, maybe I haven’t been exposed to enough industries or work places. Every place I worked definitely didn’t have bullshit jobs to hand out as adult daycare, but I can see how some places can become bloated because an over ambitious middle manager wants to say they manage X number of people on their resume. So there are bullshit jobs in that there are people who aren’t being utilized correctly, so in that case I’d say they’re no bullshit jobs, just bullshit leadership or managers.
- gowld 11mo agoI recently asked a leading GenAI chatbot to help me understand a certain physics concept. As I pressed it on the aspect I was confused about, the bot repeatedly explained, and in our discussion, consistently held firm that I was misunderstanding something, and made guesses about what I was misunderstanding. Eventually I realized and stated my mistake, and the chatbot confirmed and explained the difference between my wrong version and the truth. I looked at some sources and confirmed that the bot was right, and I had misremembered something. I was quite impressed that it didn't "give in" and validate my wrong idea.
- hathawsh 11mo agoI've seen similar results in physics. I suspect LLMs are capable of redirecting the user accurately when there have been long discussions on the web about that topic. When an LLM can pattern-match on whole discussions, it becomes a next-level search engine. Next, I hope we can somehow get LLMs to distinguish between reliable and less-reliable results.
- littlestymaar 11mo agoFound somewhere on the internet a few days ago: LLMs are Dunning-Kruger as a service. Edit: it was https://christianheilmann.com/2025/10/30/ai-is-dunning-kruger-as-a-service/ https://christianheilmann.com/2025/10/30/ai-is-dunning-kruge...
- deleted 11mo ago[deleted]
- AaronAPU 11mo agoLLMs basically act as defense attorneys for all your dumbest ideas. It is very easy to assume their confidence in you is justified, especially if you already lean narcissistic. You now see threads on X of famous people using Grok to explain how smart their ideas are. But there’s a problem: You can literally get it to do that with every single dumb idea.
- pklausler 11mo agoLLMs, kind of like Bill Bryson's books, are great at presenting "information" that seems completely plausible, authoritative, and convincing to the reader. But when you actually do know the truth about a subject, you realize how completely full of crap they too often are. And somehow after being given a patently counterfactual response to one query, we just blindly continue to take their responses to other queries as having value.
- niccl 11mo ago> But when you actually do know the truth about a subject, you realize how completely full of crap they too often are The Gell-Mann Amnesia Effect https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect
- serial_dev 11mo agoSimilar to (same as?) Gell-Mann amnesia effect.
- obelos 11mo agoI've most frequently heard this referred to as “Gell-Mann Amnesia,” and yes, LLMs are fertile ground to find it.
- rockostrich 11mo agoAt the moment, I find them to be the perfect tool to get started with learning about something. I don't expect it to tell me everything I need to know or to even be right, but if I ask ChatGPT or another LLM a question about a subject I'm not familiar with then it will at least use a bunch of terminology that I didn't have in my vocabulary before starting. For example, I just bought a 1990 Miata and I want to install a couple of rocker switches in the dash to individually control the pop-up headlights. I have enough circuits knowledge to safely change outlets and light switches, but I didn't know about relays. I asked ChatGPT how to add these switches and it immediately mentioned buying DPDT switches and tying in the OEM relay into a SPDT relay. It may have gotten the actual circuit diagram completely wrong, but now I know exactly what to read up on.
- bob1029 11mo agoI find the biggest crime with LLMs to be the size of the problems we feed them. Every time I start getting lazy and asking ChatGPT things like "write me a singleton that tracks progression for XYZ in a unity project", I wind up with a big hole where some deeper understanding of my problem should be. A better approach is to prompt it like "Show me a few ways to persist progression-like data in a unity project. Compare and contrast them". Having an LLM development policy where you ~blindly accept a solution simply because it works is like an HOV lane to hell. It is very tempting to do this when you are tired or in a rush. I do it all the time.
- nlawalker 11mo agoIt all depends on what you do with it - I see the first prompt just as a slightly different starting place than the second one.
- cluoma 11mo agoI used to get this same feeling during lectures in uni. Often the information was presented well and, along with some clear examples, everything seemed to make perfect sense. It wasn't until working through practice problems later, on my own, did it become clear how much detail I was missing.
- Swizec 11mo ago> It wasn't until working through practice problems later, on my own, did it become clear how much detail I was missing. This is a common problem in learning. Recognition is easier than recall and smoothness is confused for understanding. You actually need to struggle with the concepts a bit to learn effectively. Without the struggle it feels more effective, but is not.
- bamboozled 11mo agoDetail missing and being a confidently wrong are two different things though ? Edit: Claude told me the other day told me my entire building might have to be demolished due to a slightly bow in my newly poured stem wall, I uploaded a photo etc and it was liked, “yes this is a serious structural issue blah blah blah” , the inspector came to look at it and literally laughed that I was worried about it.
- elgenie 11mo agoNow consider what's happening to the learning process of the (rather large) subset of current college students choosing to replace that struggle for detailed understanding with LLM queries.
- lamontcg 11mo agoI'm pretty well "on the spectrum" and people glazing me in real life produce suspicion and discomfort rather than any good feelings. I don't have a problem just ignoring all the LLM glazing, although I'd really like the ability to turn it off. The fact that they've all been trained to do it, because so many of the "normies" fall for it, is kind of an indictment in my eyes. Bit of a mirror held up to society. You should probably be worried about how fake flattery works so well in society, and how this enables sociopaths and narcissists to flourish and control everything. This LLM problem is just a symptom.
- jumploops 11mo agoI recall trying to use GPT-4 to plan a trip through the PNW in ~Spring of 2023. It presented a reasonable agenda, however 80% of the rockhounding spots were completely made up! Over time, and as LLMs have gotten less sycophantic, I’ve found myself trusting them a bit more (a dangerous and slippery slope). With that said, GPT-4o in particular, seemed to rank user satisfaction above truth. I’ve found that GPT-5 Pro is currently the best at pushing back against silly ideas, and does a decent job of informing me that my questions could be better (:
- dfxm12 11mo agoAs always, trust, but verify! Google maps lists "made up" places or outdated info. AI isn't scouting these locations physically... Of course, at that point, the real question is, what's the value difference (taking into account personal, external and social costs) between asking chatgpt and /r/rockhounding (or whatever message boards they frequent)? At least if you start a thread on reddit, you might meet other people in the area with the same hobby, find a spot no one's talked about yet, get expert context and leave a trail for others to find.
- jumploops 11mo agoOne of the reasons I love rockhounding is that most of the information is _not_ online, but there is quite a bit of print literature from the last century that hasn't seemed to be scanned. My recommendation for newcomers is to find a local rockhounding club and start there. Some of the places listed in the old books are no longer publicly accessible, so best to tread carefully!
- deleted 11mo ago[deleted]
- mattcantstop 11mo agoI think this is true. It can super charge some bad takes. But I've had the opposite experience. The average person is never going to read a scientific study, nor invest the time to find out the real details of any topic they are opinionated about other than simply typing a Youtube search and finding a video that is: - Entertaining - The person has their same biases - the present the information in a short, consumable manner that doesn't require much investment. In comparison to this dynamic LLMs are wonderful. They can reference scientific data. I have noticed that they do push back on bad takes (very gently) and steer people towards truth. It's not that I think LLMs are perfect. They are not. But they are infinitely better than the average human at discovering truth.
- LeroyRaz 11mo agoYou realize that they will gladly hallucinate science... You should check the papers it claims to reference as see if the claims it makes are actually backed up. In my experience, it can completely mischaracterize scientific literature. For example, I asked it if a codebase was a faithful implementation of an algorithm described in a CS paper, and is said "no" and then proceeded to list a dozen small changes. Every single change was incorrect. The codebase was in fact a completely faithful implementation.
- ksynwa 11mo agoI wonder if LLMs need to be this way owing to the role of pseudo-intelligent conversation partners they've been shoehorned into or if it's a deliberate choice of the vendors.
- just-another-se 11mo agobeen thinking about this for a while - how will society progress when everyone has their own version of "yes man" confirming everything they think of?
- PeaceTed 11mo agoThis is a problem if 'everybody' is using it but I suspect there will be a few groups. It will be a 'tortoise and the hare' situation. The LLM folks (the hare) will get the initial upper hand as it appears as though they are moving far faster than others but with limited or wrong actual results. This could change if we can solve the hallucination issue. Yes, they are in personal echo chambers but that can only get you so far when you hit the real world. It will be painful and messy but it will resolve long term. Worse case we end up with a Dune "do not make machines that think like a person". The slow group (tortoise) are those that do not actively engadge in these things. Yes, this trying to keep up but using much slower mental faculties. I suspect long term they will do better as the fast group fail to deliver. Again if we do not solve issues of LLMs which is not certain. So long as there is still the slow group, we probably would not go down the dark path of individual echo chambers. Long term, eventually if you trip over the same mental stumbling block, you learn to not do that any more.
- Dilettante_ 11mo agoHow is "don't use LLMs as a source of truth" still news today? The machine does work, it doesn't know anything. Let the sucker fetch websites and write code.
- habibur 11mo agoI think it's ok. When wikipedia arrived, everyone was up in arms that people are learning from something that's open for anyone to edit. But it rectified itself. The same thing happened when Internet arrived. "Don't believe anything you read on the Internet." I guess the reaction was same when printed media arrived. But the thing is, things get better over time.
- radarsat1 11mo agoAh so nothing bad happening anymore due to people believing what they read on the internet, huh? Interesting take.
- Capricorn2481 11mo agoHere's a thought - improving AI is a completely different ball game.
- otabdeveloper4 11mo ago> But it rectified itself. Or did it?
- dr_hooo 11mo ago> The same thing happened when Internet arrived. "Don't believe anything you read on the Internet." Isn't the saying "Don't believe *everything* you read on the Internet."? Which is quite different (and still holds today).
- LeroyRaz 11mo agoI don't think things get better over time. What is your source for that? Here's an article (with sources) describing a massive down trend in literacy and reading comprehension: https://jmarriott.substack.com/p/the-dawn-of-the-post-literate-society-aa1 https://jmarriott.substack.com/p/the-dawn-of-the-post-litera... In short, college students nowdays have lower reading comprehension than young children in the 1850s. That is not what I would call progress. Speaking personally, I believe I would potentially have significantly worse critical reasoning abilities if I had grown up using LLMs. It is very clear to me the temptation of using them as an ersatz for engagement and thought. I think you are perhaps conflating technological progress (yes technology has improved) with demographic progress. Demographic progress is far from monotonically increasing (reading comprehension is newly plummeting, maths scores are dropping in America, science per scientist is stalling compared to 50 years ago, etc...)
- deleted 11mo ago[deleted]
- galaxyLogic 11mo ago> LLMs should not be seen as knowledge engines but as confidence engines. The thing I like best about LLM is when I ask question about some technical problem, and it tells that it is a KNOWN problem. It thus gives me confiidence that I don't need to spend time) to look for solution where there is no good soloution. Just go around it somehow. It let's me know I'm not the only person with this problem. And that way it gives me confidence that I'm not stupid, the problem is a real problem. As an example I was working with WebStorm and tried to find a way to make the Threads-tab the default tab shown when debugger opens. AI told me there is no way it knows about. Good, problem solved, solved by finding out there is no solution.
- adamhartenz 11mo agoThis is the kind of stuff AI lies about all the time. I can get it to tell me "That is some good insight, and is a known issue..." with things I make up out of thin air.
- LeroyRaz 11mo agoBe careful. The models easily hallucinate problems and misdiagnose. For example, I had an issue with some GPU code, and it assured me, with utter conviction that my problem was caused by some subtle race condition ('a known issue') that the model described in great when the real issue was just a trivial typo - no race condition, no subtly or complexity.
- inshard 11mo agoThe author misses the science of emergence. Reductionist views can’t fully explain macro-level capabilities that arise in these systems. Something emerges at higher scales from the possibility space as model sizes grow; they stop being mere “stochastic parrots” or black boxes running simple regressions. The weights develop their own inherent logic based on how they relate to each other, analogous to how brain waves encode memory at a level higher than individual neuron networks. Ultimately, the value of AI lies in the imagination of its wielder. The Unknown Unknowns framework is a useful tool for navigating AI effectively (it powerful to help elaborate on Known Unknowns and identify Unknown Unknowns), along with a healthy dose of critical thinking and understanding how reinforcement learning and RLHF work post-pretraining.
- flashfaffe2 11mo agoWithout being to self-centred here but since I have been using LLM heavily I have always challenged the results given The post seems to propose the following vector: Idea-> LLm validation -> confidence -> no further checks My process is more : Idea-> LLm response -> skeptical reflection -> adversarial prompting -> synthesis
- pedalpete 11mo agoI quite regularly ask LLMs to take the other side of an argument, or to tell me where something is wrong. Unfortunately, they don't seem very good at this process, and in some ways seem to defend the previous position. Does anyone else take this approach and have success with it?
- fennec-posix 11mo agoThis article also touches on why LLMs can be so dangerous for those who are going through a psychotic episode, it will hit you with the "That's a great idea", "You're correct", etc. Which will just further play into someone's delusions, to a point it's directly down a statistical well telling the person what they want to hear. Sadly this has ended in tragedy more than a few times.
- OfflineSergio 11mo agoI really like to know if other stuff that make things easy would have felt the same to fols who have been doing computer programming for the past 30 years if they were presented to them with a similar speed.
- kuil009 11mo agoI use LLMs mainly as a mirror for my own thinking, not as a source of authority. When I explain my ideas to the model during development, I often see flaws or confusion in my own words. This is where I learn the most. The author talks about people who rely on AI for arguments or research. They let the model's smooth, but statistical, language replace their own thinking. Language is naturally uncertain. LLMs just show this uncertainty using statistics. If you understand this, LLMs are no longer a "confidence engine." Instead, they become a tool to fix and improve your thoughts. A key point is that even if we try hard, we cannot help but react to what the AI says. We must remember that neither AI nor humans are perfect. I believe we should accept AI responses critically and always be skeptical, just like when meeting a stranger.
- spacecadet 11mo agoAlso says people are not using AI for anything meaningful. If you are trying to use AI in any meaningful way you are hyper skeptical of it and always trying to refine your dataset and understand its outputs. Anything less is mindless consumption, no different from any Joe Internet Consumer.
- satisfice 11mo agoSo weird that I have the opposite relationship with LLMs. I find them revolting and have to force myself to use them.
- BabbalaGG 11mo agohttps://www.pnas.org/doi/10.1073/pnas.2518443122 https://www.pnas.org/doi/10.1073/pnas.2518443122
- johnsmith1840 11mo agoJust like reading one good article is the same. Before LLMs I was studying distributed AI training/inference. I read hundreds of sources for this blogs, book chapters, papers, reddit posts anything and everything. If you think you understand a system via a single article or chatgpt session that's a you problem. LLM is just giving you a weird mixture of the same content you got before with a random chance of it being different every time. Ask it again, in a different way over and over. Eventually you will see between it's arbtrary lines. It is MASSIVELY faster at performing this task than before.
- more_corn 11mo ago“Does lemon juice make me invisible to cameras?” “Nope, that won’t work. If you’re planning to rob a bank try a mask or some other way of obscuring your face. Some classics are a bandanna for that classic outlaw look, a balaclava (black is the new black!). Retro cult movie style is to use some pantyhose.”