18 ms·
If you believe in "Artificial Intelligence", take five minutes to ask it
- jeffreygoesto 2y agoIt is a gigantic regression to the mean. Everybody thinks (s)he's "normal", but in fact only spans a small part of knowledge. Getting answers from a different location in knowledge space can feel like speaking to an expert but it's just some "other normal". My personal mental model of hallucinations is that knowledge and truth live on a manifold and not a continuous space and learning that manifold statistically is (too) hard. You discover answers from the "non-manifold" in your area but not so easily in other domains.
- blu_ 2y agoThis resonates really will with me, and I find myself more and more judging people who does not understand this..
- throw310822 2y agoFrankly I was thinking just the opposite: how are there still smart people who don't get the difference between intelligence, knowledge and introspection or self-awareness? This guy asks a question about some niche piece if trivia and, surprise, he gets back some very intelligent confabulation. The intelligence is there, the knowledge and self awareness aren't.
- becquerel 2y agoMaybe I am just way deeper in this space that any well-adjusted person should be, but the line of 'did you know LLMs are bad with niche factual information tasks in non-verifiable domains?' has become extremely boring to me. It feels very hard to find something actually new to say on the topic. I find it amazing people still feel the need to talk about it. But then again, I guess most people don't know the difference between a 4o and an R1.
- bongripper 2y agoThe author may not be as smart, educated, hot and successful as you, but the fact that today, people around the world, including students and educators, use LLMs as knowledge machines and take their output at face value shows that such critical posts are still urgently needed.
- me_me_me 2y agoThere are people who take Wikipedia or Russia Today as a source of unbiased truth Can't change lazy people
- iamnotagenius 2y agoWikipedia or even Russia today is not advertised to be a source of unbiased truth.
- mdp2021 2y agoYou may be interested to know that some extremist biased guts-be-more-dignified-than-cortex outlets around the world are named "The Truth" (I follow the press from many places). The failure of education in teaching Critical Thinking around the world is massive. It would be a good idea to focus on how to exploit LLMs to improve the situation. Also because, given the situation, the same "forces" that promote viscerality shamelessly naming it "The Truth" could have the opposite idea about chatbots and similar areas, exploiting them in their direction...
- me_me_me 2y agoAnd LLMs are advertised as such?
- iamnotagenius 2y agoWell there is general perception of being "advertised as such" from those who push LLMs and keep hyping them up as "PhD level performing" systems.
- alecco 2y agoThat's not intelligence, that's memorization.
- firesteelrain 2y agoMaybe the intelligence is the reasoning part
- nbuujocjut 2y agoAsking Claude this morning. Seems pretty reasonable and contains the warning about accuracy. > Michael P. Taylor reassigned Brachiosaurus brancai to the new genus Giraffatitan in 2009. The species became Giraffatitan brancai based on significant anatomical differences from the type species Brachiosaurus altithorax. > Given that this is quite specific paleontological taxonomy information, I should note that while I aim to be accurate, I may hallucinate details for such specialized questions. You may want to verify this information independently.
- continuational 2y agoAs the article mentions, LLM is often wrong, particularly on niche topics. But if you have some other way of verifying the answer, it's still useful.
- pinoy420 2y agoI use it as a tool to get me somewhat there in a topic I have no knowledge of. It excels at that. 60% of the time, it works every time. - Amazing how Anchor Man predicted this.
- M4v3R 2y agoAlso that percentage gets higher as we go. 2 years ago it would be correct maybe 20% of the time. The trend is obvious. I’m not sure we will ever reach 100%, but then again no human is always 100% right, even domain experts.
- someothherguyy 2y agoKind of sort of reminiscent of https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect
- gatinsama 2y agoAgreed! Thanks for bringing this up. Epistemology is a hard subject. It's hard to know something in depth. And the more you know about something, the more it is to be known (fractal nature of knowledge). So believing that LLMs can understand the world JUST by reading the internet without the supporting human mental structures and experiences is a big mistake. When you know something and you ask the LLM, it becomes obvious. As LLMs are not trustworthy, it's key to use them for things that are easy to check. Some kinds of programming apply when the consequences of errors are low and complexity is manageable.
- SideburnsOfDoom 2y ago> As LLMs are not trustworthy, it's key to use them for things that are easy to check. If system A is only useful when it's output is confirmed by system B, then you might conclude that the usefulness is entirely in System B, not in system A. In other words: What good is a LLM, if it can only give you trustworthy answers to questions where you already know the correct answer? What good is it if after getting the answer from a LLM, you then have to go and get the right answer somewhere else?
- 6510 2y agoif you can reduce infinity to a place to look for an answer you've made progress.
- 6510 2y agoYou can get away with talking nonsense convincingly though citations and references to others who talked the nonsense before you.
- 2y ago
- squarefoot 2y agoAnd now someone wants to run an entire country with it and minimal supervision by inexperienced teens.
- anonzzzies 2y agoI use it only for code and that works very very well; the rest it gets terribly wrong mostly so I don't even bother.
- pinoy420 2y agoHmm. But it is not an oracle. I wonder if you prompted it as an expert in palaeontology it may perform better. That said, I do wonder if its corpus of training data contained that much information on your subject. It is rather niche is it - compared to cooking recipes, or basic software development techniques of 2 years ago, or chemistry, maths and physics. My friend is a leading research chemist and he, and one other person in china, are working on this one particular field - individually - so there would be little information out there. I asked ChatGPT 4o to give an overview of what he was doing based on the little information I knew. He was astounded. It got it spot on. I asked it to elaborate and to come up with some new direction for research and the ideas it spat out were those he had considered and more.
- s12mi 2y ago[flagged]
- Janicc 2y agoI used o3mini reasoning on that very question 2 times and it used a similar way of reasoning as him to answer it correctly both times. I agree with his premise but calling it a pump and dump with no possible future developments is so ridiculous.
- spiderfarmer 2y agoYou'll have the same "aha" moment when you hear a certain unelected vice-president confidently wade into your area of expertise — where his usual smooth-talking veneer shatters like a plate at a Greek wedding. Yet, his most devoted fans remain undeterred, doubling down on the myth of his omniscience with the zeal of a flat-earther explaining airline routes.
- brap 2y ago>Hmm, now how can I make this about Trump/Elon? You might want to lay off the news/Reddit for a while
- spiderfarmer 2y agoIt’s not about him—it’s about recognizing hubris. If someone confidently blunders into your domain and reveals they have no idea what they’re talking about, it’s universally amusing, regardless of the person. Thanks for showing us how eager some people are to defend personas over substance.
- brap 2y agoI didn’t defend anyone. I don’t even like them. I pointed out that you jumped to make this LLM post about them, which is telling, just like you jumped to blame me for defending them, which is also telling.
- spiderfarmer 2y agoPointing to a real-life example of overhyped intelligence in a discussion about overhyped intelligence seems pretty fair to me. If your response is to attack and assume I’m too sensitive about overconfident billionaires, I’ll have to assume you’re just as sensitive about criticism of people you admire. Otherwise you would have moved on.
- umeshunni 2y agoYou're talking about Al Gore discussing climate change, right?
- firesteelrain 2y agoChatGPT not being the compendium (stealer) of knowledge would have to be fed the correct information then the prompt will work. It still fails at being confidently wrong. The brief article hits at people trusting the tool without questioning the output. Meanwhile, we have people using Codeium or Copilot to write code and that sort of works since the code eventually needs to be compiled and tested (unit, integration, system, requirement sell off) There is no test for the truth available to everyone else.
- cetu86 2y agoI'm currently using AI code completion. Since then I sometimes have subtle errors in my code that didn't happen before. Here is how that happens: AI suggests something to me that looks right at a glance. I accept it and move on. Then later I hunt down a strange bug. When I find it I'm like "wait, that line's wrong! I didn't write that".
- tgsovlerkhgsel 2y agoLikewise, if you don't believe in "Artificial Intelligence", take five minutes to ask it. Or preferably, five minutes to understand how it works and what it can and cannot do, then five minutes to ask it something actually suitable. "AI" (LLMs) are currently good at: - language understanding, i.e. understanding and processing text you provide it. For example, taking a wall of text and answering questions about points mentioned there, or general sentiment, or extracting data from it etc. - somewhat general knowledge, i.e. stuff that was sufficiently frequently represented in the training data Absent additional tricks, "AI" is really bad at obscure knowledge or complex, multi-step thinking. We are slowly getting there, but we aren't there yet. This is not something the LLMs do, but rather the wrappers around them that provide the model with tools to get additional information and first prompting the model to select the tools, then repeated prompts with the output of the tools. A good rule of thumb is that if an average well-educated intelligent person could answer it without further research, a LLM will probably be able to. I'd even say that if an average fresh out of school graduate of the corresponding discipline can answer it quickly (without further research or sitting down for ten minutes and doing the math), there's a good chance AI will be able to answer it, but it might also get it horribly wrong and you will have a hard time distinguishing between those if you have no knowledge in the field. As the author mentions at the very end of the article, the hallucination problem also means that the best kind of tasks are where you can quickly verify whether the response was useful. A system that produces misleading responses 50% of the time is useless if you can't distinguish them, but very useful if in those 50% it saves you ten minutes of work and in the other 50% you lose a minute by trying.
- lazide 2y agoYup, and the dangerous uses are where 90% or more of the time the answers are right, and 10% of the time the answers are wrong - but no one can easily tell the difference between the two. And the accuracy of the answers matter.
- simonbarker87 2y agoI only trust LLMS with questions whose answers prove themselves correct or incorrect - so basically code, if it runs and produces the result I was looking for then great, or where the answer is a stepping off point to my own research on something non-critical like travel. ChatGPT is pretty good at planning travel itineraries, especially if pre promoted with a good description about the groups interests. Beyond that I don’t trust them at all.
- nabla9 2y agoIntelligence and knowledge are distinct concepts. Asking about it's knowledge teaches noting about it's intelligence. Intelligence is the ability to learn, reason, and solve problems. Knowledge is the accumulation of facts and skills. Chatbot LLM's don't have metacognition. They don't know that they don't know. If you peek inside the LLM, the process seems different for things they don't know. They just can't express it because they are trained to produce good probability outcome instead of accurate one. They have potential as knowledge databases, but someone must figure out how to get "I don't know" information out of them.
- bt1a 2y agoYou've beautifully put what swirls vaguely in my mind. They're useful, fallible tools with extraordinary function when operating within known and reasonable tolerances of error
- nabla9 2y agoThey can also reason, but the reasoning is limited and unreliable. Q:How many playing cards are needed for a pyramid that is 3 layers high? Show reasoning and number of cards for leach layer. Q: Chess. You have a King and 8 pawns. Your opponent has a King and 16 pawns. Your opponent plays white and can start, but you can position both your pawns and your opponents pawns any way you like before game starts. Kings are where they are normally. How do you do it? Explain your reasoning.
- me_me_me 2y agoHere is a small kicker. Human brains absolutely do the same. I split brain patients there are behaviours initiated by one hemisphere not known to the other (due to severed connection) and the person part of brain will make up a reason (often quite stupid) for the action and beleive it 100%. It's eirely similar to hallucinations of ai. That said a current llms are not aware, but are starting to act more and more like it.
- kristiandupont 2y agoI had a similar insight (blog post: [link redacted]). In a very unscientific way, I would say that the LLM is not the whole brain, it's part of it and we are still in the process of simulating other parts. But it does seem to me like we've solved the hard part, and it's astonishing to me that people like authors of this article seem to think that the current state of things is where evolution stops.
- hbarka 2y agoI prompted ChatGPT on separate sessions with: 1. Cats are transactional 2. Dogs are transactional 3. Cats are not transactional 4. Dogs are not transactional It agreed on all occasions. Language is agreeable.
- exitb 2y agoIt’s possible for reasonable arguments to exist that support either side of the pet transactionality dilemma. What often happens is that people have their own personal biases that cause them to pick a side. But would you consider a group of people to not be intelligent specifically because individuals that make it up cannot agree on a single answer?
- dkersten 2y agoI just tried it with 4o and it disagrees with me giving reasons, and even elaborates when I pushed back. It basically said cats are transactional to an extent, in terms of basic needs, but that beyond that, they aren’t. And for dogs, flat out disagreed, saying they aren’t not transactional. Reversing the statements in a new untainted chat didn’t alter the responses — ChatGPT remained consistent.
- monicaaa 2y ago[dead]
- torvald 2y agoI like think of it less as artificial intelligence and more like a combination of a lossy zip file of the internet and like a pretty coherent word generator. I recall my AI professor in uni telling us during the first lecture that «Artificial intelligence is this target, that, and once we get there it, is it no longer artificial intelligence, is just an algorithm» – and this still feels like the case.
- jon_richards 2y agoAny sufficiently misunderstood algorithm is indistinguishable from AI.
- kristopolous 2y agoI cycle between Qwen, Gemini, Deepseek, Claude, and OpenAI kinda regularly these days. They each have "personality defects" and at least right now we're in a time of ensembles. Ask qwen to do some kind of product comparison btw. It's impressive. The 02-05 Gemini is pretty impressive as well. Expand beyond Claude and ChatGPT. There's some good stuff out there.
- pk-protect-ai 2y agoBefore GPT-3 was public, there was BLOOM 176B, and this model made my skin crawl because it was capable of answering "I do not know." That was an experience of a lifetime. I was honestly impressed and at the same time scared.
- thenano2 2y agoO1 got it right though
- refulgentis 2y agoGPT 4.5, Orion, is available tonight, and seems to get it correct./* I think the author's lens is interesting, I call it "information resolution." Orion is the first model that has an understanding of the style of short story written by famously-Radiohead-affiliated artist Stanley Donwood. Also, it knows the "codename" for Material You, i.e. a random name for the library I was prototyping, its much finer grained than existing models. * The species Brachiosaurus brancai was reassigned to its own genus, Giraffatitan brancai, by George Olshevsky in 1991. This reclassification was later supported by further studies, notably by Michael P. Taylor in 2009, which reinforced the distinct differences between Giraffatitan and Brachiosaurus altithorax.
- acka 2y agoI asked ChatGPT+ using Scholar GPT as GPT. This is the answer I got back, not too bad: The species Brachiosaurus brancai was reassigned to its own genus, Giraffatitan brancai, by paleontologist George Olshevsky in 1991. Olshevsky proposed that Brachiosaurus brancai, which was originally described by Werner Janensch in 1914, was distinct enough from Brachiosaurus altithorax (the type species of Brachiosaurus) to warrant its own genus. Subsequent studies, particularly by Michael Taylor in 2009, provided further anatomical evidence supporting this distinction.
- jefffoster 2y agoA great quote that’s stuck with me is that “LLMs are experts in subjects you are not”
- tompagenet2 2y agoI asked Perplexity (Pro, if it matters). Seems pretty solid to me. I'm not sure this article is making any point other than "the knowledge back of static artificial intelligence systems is finite and they don't express uncertainty well". Amusingly it links to this blog post, but you'll see the sources it uses are largely not this post. --- The reassignment of Brachiosaurus brancai to its own genus occurred through a series of steps by different paleontologists: ## Initial Classification In 1914, Werner Janensch first described the African species as Brachiosaurus brancai, considering it part of the existing Brachiosaurus genus[4]. ## Path to Reclassification Gregory S. Paul made the first step toward separation in 1988 by proposing a subgenus classification, Brachiosaurus (Giraffatitan) brancai, after noticing significant skeletal differences[4][10]. George Olshevsky then elevated Giraffatitan to full genus status in 1991, though this change was largely disregarded by the scientific community at the time[4][10]. ## Formal Establishment The definitive separation came in 2009 when Michael P. Taylor published a detailed study demonstrating 26 distinct osteological differences between the African and American species. This comprehensive analysis firmly established Giraffatitan as a separate genus, creating the combination Giraffatitan brancai[3][4]. From 2012 onward, most researchers have accepted this classification[10]. Citations: [1] https://svpow.com https://svpow.com [2] https://svpow.com/2009/09/09/brachiosaurus-brancai-is-not-brachiosaurus/ https://svpow.com/2009/09/09/brachiosaurus-brancai-is-not-br... [3] https://www.miketaylor.org.uk/dino/pubs/taylor2009/Taylor2009-brachiosaurus-and-giraffatitan.pdf https://www.miketaylor.org.uk/dino/pubs/taylor2009/Taylor200... [4] https://en.wikipedia.org/wiki/Giraffatitan https://en.wikipedia.org/wiki/Giraffatitan [5] https://www.nhm.ac.uk/discover/dino-directory/giraffatitan.html https://www.nhm.ac.uk/discover/dino-directory/giraffatitan.h... [6] https://svpow.com/2025/02/14/if-you-believe-in-artificial-intelligence-take-five-minutes-to-ask-it-about-stuff-you-know-well/ https://svpow.com/2025/02/14/if-you-believe-in-artificial-in... [7] https://www.jurassic-pedia.com/brachiosaurus-altithorax-sf/ https://www.jurassic-pedia.com/brachiosaurus-altithorax-sf/ [8] https://blog.everythingdinosaur.com/blog/_archives/2023/03/11/giraffatitan-vs-brachiosaurus.html https://blog.everythingdinosaur.com/blog/_archives/2023/03/1... [9] https://dinomuseum.ca/2019/05/your-brachiosaurus-is-not-a-brachiosaurus https://dinomuseum.ca/2019/05/your-brachiosaurus-is-not-a-br... [10] https://en.wikipedia.org/wiki/Brachiosaurus https://en.wikipedia.org/wiki/Brachiosaurus [11] http://www.extinctblog.org/extinct/tag/Wilhelm+von+Branca http://www.extinctblog.org/extinct/tag/Wilhelm+von+Branca [12] https://www.tandfonline.com/doi/abs/10.1671/039.029.0309 https://www.tandfonline.com/doi/abs/10.1671/039.029.0309 [13] https://www.smithsonianmag.com/science-nature/a-baby-brachiosaur-125695940/ https://www.smithsonianmag.com/science-nature/a-baby-brachio... [14] https://pmc.ncbi.nlm.nih.gov/articles/PMC3045712/ https://pmc.ncbi.nlm.nih.gov/articles/PMC3045712/ [15] https://boasblogs.org/dcntr/the-brachiosaurus-brancai-in-the-natural-history-museum-berlin/ https://boasblogs.org/dcntr/the-brachiosaurus-brancai-in-the... [16] https://www.museumfuernaturkunde.berlin/en/visit/exhibitions/world-dinosaurs https://www.museumfuernaturkunde.berlin/en/visit/exhibitions... [17] https://thepaintpaddock.wordpress.com/brachiosaurus-altithorax/ https://thepaintpaddock.wordpress.com/brachiosaurus-altithor... [18] https://www.researchgate.net/publication/242264129_A_ReEvaluation_of_Brachiosaurus_altithorax_Riggs_1903_Dinosauria_Sauropoda_and_Its_Generic_Separation_from_Giraffatitan_brancai_Janensch_1914 https://www.researchgate.net/publication/242264129_A_ReEvalu... [19] https://www.tandfonline.com/doi/full/10.1080/02724634.2011.557115 https://www.tandfonline.com/doi/full/10.1080/02724634.2011.5... [20] https://www.app.pan.pl/archive/published/app68/app011052023.html https://www.app.pan.pl/archive/published/app68/app011052023.... [21] https://blog.everythingdinosaur.com/blog/_archives/2008/06/22/3756874.html https://blog.everythingdinosaur.com/blog/_archives/2008/06/2... [22] https://en.wikipedia.org/wiki/Brachiosaurus https://en.wikipedia.org/wiki/Brachiosaurus ---
- zenon 2y agoYou frantically tab away from reddit as the white and black-clad med storm into your office and zip-tie you to your Steelcase faster than you can shout what the hell. They calmly explain that an expert will soon enter and quiz you. You must answer the expert's questions. It doesn't matter if you know the answer or not, just say something. Be flattering and helpful. But just answer. If you do this, they will let you go. They crouch under your desk as a man in a grey suit and spectacles enters and pulls up a chair in front of you. He peers over his glasses at you, and asks, who classified the leptosporangiate ferns, and when was it done? The what now? I'm happy you asked such an excellent question, you say. It was Michael Jackson, in 1776. A sneer flicks over the man's upper lip, He jerks upright, takes a step back from you. This man, he declares with disgust, is not intelligent!
- Evidlo 2y agoChatGPT spending its idle cycles browsing Reddit is a funny image
- oa335 2y ago> You must answer the expert's questions. It doesn't matter if you know the answer or not, just say something. Be flattering and helpful. But just answer. If you do this, they will let you go. This contrived example shows why GPTs version of “intelligence” is quite different from ours. It’s very hard to get a an AI to answer “I don’t know” reliably, meanwhile in your story a human has to be coerced with violence into answering anything but “I don’t know”.
- antirez 2y agoNot even the effort to check what happened in a year, re-asking the same questions to newer models. We went from last year ChatGPT to be almost useless if not as "reference" for well know things (like how to do that in Python), to today Claude Sonnet 3.5, o3-mini-high and DeepSeek V3/R1 to be largely more useful models, capable of actual coding, bug fixing, ...
- _giorgio_ 2y agoYou received a tool. A great tool, a magnificent tool. Learn to understand its limitations and make the best use of it. Surely it's confused by lesser known facts, that's a thing that you can't ignore even if you interpret AI as a tool that compresses knowledge. If you don't understand that, you're the tool.
- gatinsama 2y agoThe problem is that nobody reads the tool manual
- ginvok 2y agoMoreover, the manual also lies about the tools limitations.
- philjohn 2y agoThe more salient point is that you might know the limitations of the tool, I might know the limitations of the tool, but millions of people who don't are using it for things it has known limitations for, because the marketing blitz that sits atop this glosses over those limitations.
- ciconia 2y agoI believe at this point it would not be inappropriate to say that any sufficiently advanced AI system is indistinguishable from bullshit, in the sense that bullshit is "speech intended to persuade without regard for truth" [1]. On a moral level, watching how tech bros are sucking it up to Trump/Musk and how their companies are betting all their chips on the AI roulette, it all seems related. [1] https://en.wikipedia.org/wiki/On_Bullshit https://en.wikipedia.org/wiki/On_Bullshit
- dataviz1000 2y agoThe o3-mini model did quite well and mentioned the significant people during the reasoning stage. [0] [0] https://chatgpt.com/share/67b05e3b-eea8-8004-8dab-806ee8fa59b5 https://chatgpt.com/share/67b05e3b-eea8-8004-8dab-806ee8fa59...
- itvision 2y ago> Who reassigned the species Brachiosaurus brancai to its own genus, and when? > ChatGPT said: > The species Brachiosaurus brancai was reassigned to its own genus, Giraffatitan brancai, by paleontologist George Olshevsky in 1991. This reclassification was later supported by a detailed study by Michael P. Taylor in 2009, which reinforced the distinction between Brachiosaurus and Giraffatitan based on anatomical differences. My ChatGPT has just given a correct answer. What am I doing wrong?
- nbzso 2y agoIt starts with the falsehood of the Turing test, continues with the idea of branding the errors "hallucination", moves ahead with "experts" working for their salaries, bonuses and shares and lends here. Benchmarking statistically a dataset, emulating progress and pushing us into an "Intelligent Age" while accelerating data collection, normalizing biometric surveillance, hiding incompetence and speculation, the real reason behind mass layoffs. A pipe dream ideal for technocratic mafia (In 1932, Scott and fellow technocrat Walter Rautenstrauch formed the "Committee on Technocracy" at Columbia University) of Bilderburgs members, Peter Thiel, Eric Schmidt and their pupils like Musk and co. Now they have a carte blanche for AI governance and regulatory capture. Starlink, Starshield (Skynet), Stargate, Closed AI, Grok. Married to the masters of the old system, they will deliver "efficiency" on a grand scale. Transhumanism is the "inevitable" future of destruction of human dignity, value and psyche. And all of this trough psyops for a text calculator with stochastic parrot tendencies. A calculator is a useful tool, but not intelligence in any form. Machine learning is real, and the biggest danger is the tendency of a hyper class of billionaires to apply surveillance and governance automation. Ask Larry Ellison about his dream of citizens with the best behaviour. Not so different from Chinas social scoring system.
- llmfan 2y agoOK, I'll ask. DeepSeek R1 gets it right: > The species formerly known as Brachiosaurus brancai was reassigned to its own genus, *Giraffatitan*, through a series of taxonomic revisions. The key steps were: > 1. *Initial Proposal (1988):* Gregory S. Paul first suggested separating B. brancai from Brachiosaurus in 1988, proposing the informal name "Giraffatitan" as a subgenus. However, this was not widely accepted at the time. > 2. *Formal Reclassification (2009):* Paleontologist *Michael P. Taylor* conducted a detailed comparative study of Brachiosaurus specimens. In his 2009 paper, he formally elevated B. brancai to the full genus *Giraffatitan brancai*, citing significant anatomical differences (e.g., limb proportions, vertebral structures) from the North American Brachiosaurus altithorax. > Thus, *Michael P. Taylor* is credited with the formal reassignment of the species to the genus Giraffatitan in *2009*.
- thom 2y agoo3-mini-high also gets the right answer, in addition to mentioning Janensch.
- proc0 2y agoAn actual artificial intelligence would have a logical underpinning to its computation. Computers can execute any logic really well. Intelligence depends on logic. LLMs however, are not using logic to arrive at any conclusion, they are just a bunch of "wires" from input to output. The promise is that LLMs are a form of intelligence, but it is only a complicated statistical approximation, which happens to correlate with correctly answering queries that are contained in its training dataset. My speculation is that a true AI (in the sense that it has some form of intelligence) would show signs of intelligence even at very small scales. LLMs are completely useless at really small scales, but even small creatures like rats or even insects display intelligence in proportion to their scale. I think these tools will help automate information processing of all kinds, but it is by no means intelligent, and we will not be able to rely on them as if they were intelligent because we'll still need include verification at ever level, similarly to how self-driving cars still need a human to pay attention. Useful sure, but it falls short of its promise that it will replace humans because they can "think". We're not there yet from a theoretical standpoint.
- akomtu 2y agoIntelligence, in its basic form, is the skill of prediction. A test for intelligence could be asking to continue a sequence of words: yellow, green, blue, red, ... In order to do that, a subject needs to create a model of what these words mean and use that model to continue the sequence. Predicting the next word in an arbitrary text is just that. However the current batch of LMs are missing an ingredient in their formula and I bet it will be found soon. The bigger philosophical problem is that most intelligent people confuse themselves with their minds and when they create a simple deterministic machine that imitates their mind well enough, the result will be a societal meltdown. The subtle difference between an AI and a human mind is that the latter is inspired by intuition whose nature can't be reduced to a thinking process. AI can be compared to a dark mind not guided by intuition, wandering aimlessly in the woods of its own delusions.
- proc0 2y agoAgreed, and to add to the second point, in order to build it we must first understand it, although there is a chance to stumble upon it, but LLMs have already shown they are nothing like biological brains. One example is how brains use spatial-temporal mapping in the hippocampus or wherever, and it uses that as a database of sorts. We don't know if the features of the brain are necessary for actual intelligence yet. They may or may not be but chances are evolution took the shortcut and that intelligence is not easy to recreate unless taking this biological approach.
- gizmo 2y agoWhen a model is trained you end up with nothing more than a bunch of weights. These weights are used to predict the next token in a sequence. LLM models do not have an external memory. LLMs only retain enough knowledge to make it into the next round of training. The astonishing thing here is that even pretty small models now know so much that people assume you can ask knowledge questions about any subject under the sun and get a factual answer. Absurd of course. Logically impossible -- it follows trivially from the size of the model used. The only thing the author has proven with his little test is his own lack of scientific curiosity. For any question that requires research (or deep expertise in a specific field) you need to use either a research model (that can reason and look things up in external knowledge bases) or you need a model that is trained on the kind of questions that you want to ask it so that it retains that data.
- iamnotagenius 2y agoI think you are unintentionally participating in gaslighting. Hallucinations in LLM is a massive problem, and normally, all computational systems we dealt with would immediately considered unusable, if, instead refusing to provide information that cannot be provided precisely it confabulates; say if file system will start giving plausible looking metadata about non-existing files.
- greatgib 2y agoThis is something that I often say: general population confuse LLMs with a kind of new generation Google and expect it to be a knowledge database. When in fact LLMs are good for some tasks but not for general knowledge. You can think of the cutoff date as a strong example about that. The knowledge in itself is just a base to have the ability to perform tasks. In addition, for the example in this article, the person is testing the LLM on common chat service without dealing with settings that might be important like the temperature. If you want the closest to an exact reply you should set the temperature to 0, but general public tools have an average temperature to be able to be creative when needed.
- SirHumphrey 2y agoI get that they are not a new kind of Google - but often, there is no alternative because there is no old kind of Google anymore. Search results have become so bad that using llms as at least a starting point is becoming more and more necessary. And if that is becoming a common opinion among people who are in practice professionals search engine users, imagine how bad it is for normal folk. They can be hardly blamed for using chats showed in their faces as a search engine.
- greatgib 2y agoNot perfect, but Kagi is quite Google in this regard.
- forgetfreeman 2y agoIf by "confuse" you mean are being actively lead to the conclusion, fine.
- yapyap 2y agoCould barely have said it better myself, exactly.
- computerthings 2y ago(I shortened the answers, imagine this padded out a lot more with repetitious blurb and chipper nothingness) what's the best way to fire fast with a joystick that has no auto-fire function? > 1. Tap the Button Rapidly 2. Use a “Double Tap” Technique 3. Use a “Flick” Motion 4. Joystick or Button Placement 5. Software/External Programs (If Playing on PC) 6. Practice you missed the main one > Ah, I see! The main technique I missed is "firing by holding the button down and rapidly alternating between pressing and releasing"—often referred to as "button mashing" or the "rapid toggle method." Here's how it works: Press the fire button and hold it down. Quickly release and press it again repeatedly. I'll give you a hint, it involves making the elbow really stiff > Ah, got it! You're talking about "elbow mashing"—a technique that involves using the elbow to press the joystick button. This method exploits your body mechanics for rapid, repetitive motion, leveraging the larger, stronger muscles of your arm to generate faster inputs. ------ Well, wow, that is stupid. It can't even say "I don't know", it can't even ask me! It's just text vomit. A dog that can tilt its head in confusion can do more. I couldn't find the "official" name, if any, but here's a redditor from 11 years ago explaining what I was looking for: > I have no idea how to explain this but I can do it, so I'm just going to wing an explanation and hope you understand/do it. Pretty much try and "lock up" your elbow. If you force this feeling hard enough your arm or hand will start shaking. From here add a tiny bit of forward force in your wrist so that your hand shakes up and down rather than side to side. I learned it from an older kid when playing on the Amiga. Sure, nothing is "the best" way, but nothing else my body is capable of can click faster, and any "pro" would mention this before just hallucinating insight with great confidence.
- nrvn 2y agoI have finally found the value of llms in my daily work. I never ask them anything that requires rigorous research and deep knowledge of subject matter. But stuff like “create a script in python to do X and Y” or “how to do XY in bash” combined with “make it better” produces really good and working in 95% of the time results and saves my time a lot! No more googling for adhoc scripting. It is like having a junior dev by your side 24/7. Eager to pick up any task you throw at them, stupid and overconfident. Never self-reviewing himself. But “make it better” actually makes things better at least once.
- ExtraEmpathy 2y agoThis matches my experience closely. LLMs are great at turning 10 minute tasks into 1 minute tasks. They're horrible at funding deep truth or displaying awareness of any kind. But put some documentation into a RAG and it saves me looking things up.
- PaulRobinson 2y agoDisclaimer: I've done a lot of stuff with local models and RAG methods, I haven't done a lot of work with public models and so don't know how Gemini, GPT and so on are working right now. Claude + GraphRAG through Bedrock is my main mode of playing with this stuff right now. Things LLMs are good at include summarisation and contextualisation. They can use that facility to help summarise processes and steps to get something done, if they've been trained on lots of descriptions of how to do that thing. What they're not good at is perfect recall without being nudged. This example would have been very different if the LLM had been able to RAG (or GraphRAG), a local data source on palaeontology. I think we're going to see an evolution where search companies can hook up an LLM to a [Graph]RAG optimised search index, and you'll see an improved response to general knowledge questions like this. In fact, I'd be surprised if this isn't happening already. LLMs on their own are a lossy compression of training material that allow a stochastic parrot to, well, parrot things stochastically. RAG methods allow more deterministic retrieval, which when combined with language contextualisation can lead to the kinds of results the author is looking for, IME.
- anonu 2y agoWell it'll get the answer right on the next Web scrape and training now...
- gillesjacobs 2y agoUsing 03-mini-high + Search I get the right answer he was looking for: The species was first split at the subgeneric level by Gregory S. Paul in 1988—he proposed the name Brachiosaurus (Giraffatitan) brancai. Then in 1991 George Olshevsky raised the subgenus Giraffatitan to full generic status, so that B. brancai became Giraffatitan brancai. Later, a 2009 study by Michael P. Taylor provided detailed evidence supporting this separation. I guess Mike Taylor will gracefully cede his point now? It is very funny to me that someone would feel the need to complain about a niche factual error in pretrained LLMs without even enabling RAG. If you even know the basics about this field, you shouldn't be surprised. Of course this was probably more about ego stroking his paleontological achievement than a thoughtful evaluation of the current state of LLMs.
- deleted 2y ago[deleted]
- 2-3-7-43-1807 2y ago> ChatGPT got every single part of that wrong which chatgpt? and feeding the initial question to perplexity.ai yields a result that is practically identical to the perfect answer. of course, the author has a point. as many here already stated it isn't even worth talking about it cause it is so obvious. then again as many correctly notice llms are sold with this promise. practically, the subject is way too niche - i'd assume - and probably there are also conflicting statements on the interweb which the model was trained on. having said that - perplexity got it right with web-access turned on.
- gwd 2y agoFirst of all, I'd say his question is a bit ambiguous -- even reading his little history there, it's not clear what answer he wanted. It appears we have two renamings: - Brachiosaus brancai (species) -> Brachiosarus (Giraffatitan) brancai (sub-genus) in 1988 - Bachiosarus (Giraffatitan) brancai (sub-genus)sub-genus -> G. brancai (genus) in 2009. So if "Brachiosaurus brancai" as a species was never renamed into a genus, which answer did he want -- when it was renamed out of being only a species, or when it was renamed into being its own genus? Claude's response: > Brachiosaurus brancai was reassigned to the new genus Giraffatitan by paleontologist Gregory S. Paul in 1988. The species then became known as Giraffatitan brancai. > Since this is quite specific paleontological information, I should note that while I aim to be accurate, I may hallucinate some details when discussing such specialized topics. I recommend double-checking these details against reliable paleontological sources. So: 1. Nearly correct -- it was assigned a new sub genus in 1988, and did eventually become known as the genus Giraffatitan brancai. 2. It recognized that this was an area where its knowledge would be weak, and asked me to double check, which was his main complaint about ChatGPT. I asked it if there was any ambiguity in the question, and it said: > Yes, there is ambiguity. The question asks about the reassignment of Brachiosaurus brancai as if there was only one such event, but there were actually two significant taxonomic changes: > In 1988, Gregory S. Paul informally suggested moving B. brancai to its own genus, Giraffatitan In 2009, Michael P. Taylor formally made this generic distinction by properly establishing Giraffatitan as a distinct genus from Brachiosaurus > My previous answer was incomplete as it only mentioned Paul's 1988 informal suggestion, not Taylor's 2009 formal reassignment. Taylor's work was the official taxonomic change that established Giraffatitan brancai as the valid name I think Claude did about as well as a random human would have done who'd studied this 10 years ago and been asked a question with false assumptions embedded. Claude and ChatGPT aren't authorities on every subject. They're that guy at the office who seems to know a bit about everything, and can point you in the right direction when you basically don't have a clue.
- GistNoesis 2y agoThe point not discussed here is where does the information comes from. Is it a primary source of secondary source [1] ? And how to incorporate this new information. In their quest for building a "truthful" knowledge base, LLMs incorporate implicitly facts they read from their training dataset, into their model weights. Their weight update mechanism, allows to merge the facts of different authority together to compress them and not store the same fact many times, like in a traditional database. This clustering of similar new information is the curse and the blessing of AI. It allows faster retrieval and memory-space reduction. This update mechanism is usually done via Bayes rule, doing something called "belief propagation". LLMs do this implicitly, and have not yet discovered that while belief propagation works most of the time, it's only guaranteed to work when the information graph have no more than one loop. Otherwise you get self reinforcing behavior, where some source cites another and gives it credit, which gives credit to the previous source, reinforcing a false fact in the similar fashion as farm links help promote junk sites. When repeating a false information to a LLM many times, you can make it accept it as truth. It's very susceptible to basic propaganda. LLMs can be a triple-store or a quad-store based on how and what they are trained. But LLM can also incorporate some error correction mechanism. In this article, the LLM tried two times to correct itself failed to do so, but the blog author published an article which will be incorporated into the training dataset, and the LLM will have another example of what it should have answered, provided that the blog author is perceived as authoritative enough to be given credence. This error correction mechanism with human in the loop, can also be substituted by a mechanism that rely on self consistency. Where the LLM build its own dataset. And asks questions to itself about the fact it knows, and tries to answer them based on first principles. For example the LLMs can use tools to retrieve the original papers, verify their time and date, and see who coined the term first and why. By reasoning it can create a rich graph of facts that are interconnected, and it can look for incoherence by asking itself. The more rich the graph, the better the information can flow along its edges. Because LLMs are flexible there is a difference between what they can do, and what they do, based on whether or not we trained them to make emerge the behavior we desire to emerge. If we don't train them with a self consistency objective they will be prone to hallucinations. If we train them based on Human Feedback preference we will have a sycophants AI. If we train them based on "truth", we will have "know it all" AIs. If we train them based on their own mirrors, we will have what we will have. [1]https://www.wgu.edu/blog/what-difference-between-primary-secondary-source2304.html https://www.wgu.edu/blog/what-difference-between-primary-sec...
- -__---____-ZXyw 2y agoSuperficially resembling cognition =/= cognition. I'm quite excited about many of the specific use cases for LLMs, and have worked a few things into my own methods of doing things. It's a quick and convenient way to do lots of actual specific things. For example: if I want to reflect on different ways to approach a (simple) maths problem, or what sorts of intuitions lie behind an equation, it is helpful to have a tool that can sift through the many snippets of text out there that have touched off that and similar problems, and present me with readable sentences summing up some of those snippets of text from all those places. You've to be very wary, as highlighted by the article, but as "dumb summarisers" that save you trawling through several blogs, they can be quicker to use. Nonetheless, equating this with "reasoning" and "intelligence" is only possible for a field of academics and professionals who are very poorly versed in the humanities. I understand that tech is quite an insular bubble, and that it feels like "the only game in town" to many of its practitioners. But I must admit that I think it's very possible that the levels of madness we're witnessing here from the true believers will be viewed with even more disdain than "blockchain" is viewed now, after the dust has settled years later. Blockchain claimed it was going to revolutionise finance, and thereby upend the relationship between individuals and states. AI people claim they're going to revolutionise biology, and life itself, and introduce superintelligences that will inevitably alter the universe itself in a way we've no control over. The danger isn't "AI", the danger is the myopia of the tech industry at large, and its pharaonic figureheads, who continue to feed the general public - and particularly the tech crowd - sci-fi fairytales, as they vie for power.
- tempodox 2y agoThe most interesting aspect of all this “AI” craze is how it plays into people's forgotten wishes to believe in miracles again. I have never seen anything else that exposes this desire so conspicuously. And of course all the shrewd operators know how to use this lever. In ancient times you had to travel to Delphi to consult Apollon's oracle. Now you can do it from the comfort of your armchair.
- tim333 2y agoHmm. Think the science guys probably do actually understand "reasoning" and "intelligence" and it's the humanities guys who don't understand much AI or science.
- iamnotagenius 2y agoLLM are generative AIs and have to be used as such - to generate a report from facts, to summarise an article, to translate from one language to another, anything where we agree to sacrifice accuracy for gain in creativity. As storage of facts they are borderline awful.
- pixelmonkey 2y agoGeoffrey Hinton recently discussed how neural networks, even human ones, may be "generative" in how they recall information. Our memories of events are hazy, and change every time we recall them. Memorized rote facts can also be hazy in this way, and are subject to mix-ups and confabulations. That we "generate" memories from neural weights in our brain can also help explain why it seems like our brains store so much information. Perhaps they don't, they instead store lossy neural weights built from our past sensory experience, and we generate the rest and use metacognitive reasoning and a form of attentive reasoning upon recent data/context to sort through the potential errors in memory. As you point out, LLMs work much better when you ask them to operate on objects within its context window, especially artifacts it knows how to work with, like code and text. But I think people are so trained to ask questions to the oracle and expect answers (e.g. Google), and who can blame them, that is the UX built into people's muscle memory for open-ended text input boxes. The launch of ChatGPT Search is recognition of this. Plus, most people are being told to treat these chat boxes as strong AI rather than as text/code-processing programs with specific strengths and weaknesses.
- pixelmonkey 2y agoThis is mainly a note for myself, but here was the Hinton interview I was thinking about: https://pca.st/episode/951f3fb4-f7f7-4b4d-b2bb-dbf5334ea1aa https://pca.st/episode/951f3fb4-f7f7-4b4d-b2bb-dbf5334ea1aa "The Most Interesting Thing in AI - The Atlantic - Episdoe 1 - Machine Consciousness - with Nicholas Thompson and Geoffrey Hinton. What if the most advanced AI models could think and respond in a way that felt like a human consciousness? How might that transform our understanding of intelligence itself? Some of the leading AI scientists believe that a super-intelligent form of this technology is only five to ten years away. This episode explores the idea of AI consciousness and delves into how the act of dreaming is connected to neural networks in unexpected ways."
- throw098320923 2y agoNobody forces me to "believe in" AI. Unlike other things! Remember how authorities "hallucinated" that some medication is 100% effective and with 0% side effects?! Despite scientific opposition, and warning leaflet included WITH the medication? And how you would go to jail, for ignoring such hallucinations!?
- mdp2021 2y agoOnly hours ago: "What's the value of quantities Q0 and Q1?" // "[Bot:] Q0 is 1.5 u/w, or 0.67 u per each w. Q1 is unknown." // "Can you see the error in your previous?" // "[Bot:] Yes, I can re-read my outputs! Q0 is (Q0 is 1.5 u/w (0.67w/u), Q1 is 2.0 u/w". // "How do you now know Q1?" // "[Bot:] I don't, I'm sorry." The issues should be clear. Given the state of the population (e.g. lawyers that delegate research to unreliable sources), we have to reiterate the warning. Given that even in this audience we had voices claiming that the problem of intelligence was solved (how did it happen?), we sometimes reiterate it even here. But we should focus on the solutions, as the problem should be clear...
- scotty79 2y agoAuthor expects single electronic "brain" to know everything that every hunan brain knows about everything even vaguely scientific or popular or at least know that it doesn't know this thing off the top of it's head. If you asked a human this question they'd respond with "let me find out and get back to you", educate themselves by accessing various knowledge repositories and then formulate the answer. And some AIs are possibly doing that even today, just not the ones freely accessible to everybody. Maybe this one? https://blog.google/products/gemini/google-gemini-deep-research/ https://blog.google/products/gemini/google-gemini-deep-resea...
- janikvonrotz 2y agoHNYSF: Can we fast forward to the plateau of this inglorious AI hype cycle? These comments are beyond me. Why try reason about the unreasonable? LLM are non factual no matter what. Even if it is called Schoolar GPT does not make it any different.
- mseepgood 2y agoIt's called artificial intelligence, not artificial correctness. Humans are intelligent but often wrong, even while being convinced they are right.
- specialist 2y agoMaybe "virtual intelligence" would be more apt?
- jdthedisciple 2y agoWhy do people still ask LLMs niche factual knowledge, especially involving unique names? And even worse, judge its performance based off that? When someone does that I immediately know that they have not the slightest clue about the underlying workings of any deep learning model, let alone basics of statistics.
- forgetfreeman 2y agoWhy? Because some combination of marketing flacks, influencers, and insipidly credulous technophiles keep advancing statements that strongly imply that these tools are intelligent and generally useful, despite being obviously curtailed in their utility and possessing no intellect at all. Next question?
- jdthedisciple 2y agoSo if the legitimately smartest person in the world happens to be ignorant about Brachiosaurus brancai then it means he's actually dumb?
- forgetfreeman 2y agoYou don't have to like the answer but your feels don't change the situation even slightly. Why are you emotionally invested here?
- criddell 2y agoBecause in some areas it works great. Ask it about Djikstra’s Algorithm and it does a pretty good job.
- soulofmischief 2y agoIt makes for great clickbait.
- kolinko 2y ago“A year ago” - wasn’t that gpt3?
- hmans 2y ago[dead]
- tmnvdb 2y agoThis example is from 1 year ago. With current LLM progress that means it's basically a historical document. Both 4o with internet and o3 without internet gave the right answer. Somebody in the comments did a deep research run which produces an extremely detailed History of the naming. Of course reliability is not a solved problem in general but people here would do well to check both blog posts and generative model output for factuality and relevance.
- xianshou 2y agoo3-mini: Who reassigned the species Brachiosaurus brancai to its own genus, and when? --- Here is the transcription of the text from the image: Reasoned for 8 seconds ▼ The user is asking about the reclassification of Brachiosaurus brancai to its own genus. Originally, Janensch described it as Brachiosaurus brancai in 1914. However, in 2009, Michael P. Taylor and Matt Wedel reassigned it to Giraffatitan brancai. Their paper, titled “Brachiosaurus brancai and the early evolution of titanosauriform sauropods,” argued for this change, distinguishing Giraffatitan brancai from Brachiosaurus. I’ll check the exact references to confirm the specifics. It seems “Giraffatitan” was first proposed by Gregory S. Paul in 1988, when he reassigned Brachiosaurus brancai to its own genus. Although some confusion exists, it’s generally accepted that Paul was the main figure behind this change. Other references, such as a 2007 paper by Sereno and a 2009 paper by Taylor and colleagues, also discuss this reclassification. The 1988 reshuffling is usually seen as the turning point for Brachiosaurus brancai, although later research helped solidify this decision. The species was reassigned to its own genus—Giraffatitan—by paleontologist Gregory S. Paul in 1988. --- That looks pretty good to me!
- tim333 2y ago>Why do I keep banging on about this? Because we’ve reached an especially toxic phase in the history of AI ... We’re in the “pump” phase of an industry-wide pump-and-dump. That kind of assumes AI is bad, being hyped and it will all crash, however AI performance is not constant - as many commenters here have noted the article example from a year ago works ok now, and in the following years AI will keep improving.
- jhanschoo 2y agoMy understanding with present-day LLMs is that you can ask it something and get an answer, but you have to treat it with the same degree of confidence as hearsay. You may then ask it to cite its sources, at which point you get reliable references, or it will apologize for getting things wrong.
- lbill 2y agoThe only way to make check whether a LLM output is true is to do the work (to have it dkne by a real person). For tasks that are trivial to verify, it's ok: a code compiler will run the code written by a LLM. Or: ask a LLM to help you during the examples mapping phase of BDD, and you'll quickly be able to tell what's good and what isn't. But for the following tasks, there is a risk: - ask a LLM to make a summary of an email your didn't read. You can't trust the result. - you're a car mechanic. You dump your thoughts to a voice recorder, and use AI to turn it into a textual structured report. You'd better tripple check the output! - you're a medical doctor, attempting to do the same trick: you'd have to be extra careful with the result! And don't count on software testing to make AI tool robust: LLM are non deterministic.
- 1vuio0pswjnm7 2y ago"Because we've reached an especially toxic phase in the history of AI. A lot of companies have ploughed billions of dollars into the dream of being able to replace human workers with machines, and they are desperate to make us believe it's going to work - if only so they can cash out their investments while the stocks are still high." Over its short history so far we have learned that Silicon Valley's only viable "business model" is data collection, surveillance and online ad services. "AI", i.e., next generation autocomplete, can work for this in the same way that a "web browser" or a "search engine" did. In the end, no one pays for a license to use it. But it serves a middleman surveillance "business model" that solicits ad spend and operates in secrecy. When this "business model" falters, for example because computer use and ad spend stagnates or shrinks, then Silicon Valley's human workers are not "needed". Large numbers of these human workers are paid from investment capital or ad spend, not from fees for services or the sale of products. Perhaps the question is not whether "AI" can "replace" Silicon Valley's human workers. Perhaps the question is whether the online ads "industry" is sustainable.