9 ms·
The technology is not just less than superintelligence, for many applications it is less than prior forms of intelligence like traditional search and Stack Exch
by WhyOhWhyQ 2y ago
The technology is not just less than superintelligence, for many applications it is less than prior forms of intelligence like traditional search and Stack Exchange, which were easily accessible 3 years ago and are in the process of being displaced by LLMs. I find that outcome unimpressive.
And this Tweeter's complaints do not sound like a demand for superintelligence. They sound like a demand for something far more basic than the hype has been promising for years now.
- "They continue to fabricate links, references, and quotes, like they did from day one."
- "I ask them to give me a source for an alleged quote, I click on the link, it returns a 404 error." (Why have these companies not manually engineered out a problem like this by now? Just do a check to make sure links are real. That's pretty unimpressive to me.)
- "They reference a scientific publication, I look it up, it doesn't exist."
- "I have tried Gemini, and actually it was even worse in that it frequently refuses to even search for a source and instead gives me instructions for how to do it myself."
- "I also use them for quick estimates for orders of magnitude and they get them wrong all the time. "
- "Yesterday I uploaded a paper to GPT to ask it to write a summary and it told me the paper is from 2023, when the header of the PDF clearly says it's from 2025. "
- Thlom 2y agoA municipality in Norway used LLM to create a report about the school structure in the municipality (how many schools are there, how many should there be, where should they be, how big should they be, pros and cons of different size schools and classes etc etc). Turns out the LLM invented scientific papers to use as references and the whole report is complete and utter garbage based on hallucinations.
- brookst 2y agoAnd that says… what? The entire LLM technology is worthless for all applications, from all implementations? A company I worked for spent millions on a customer service solution that never worked. I wouldn’t say that contracted software is useless.
- nancyminusone 2y agoNo, it says that people dislike liars. If you are known for making up things constantly, you might have a harder time gaining trust, even if you're right this time.
- sswatson 2y agoAll of these things can be true at the same time: 1. LLMs have been massively overhyped, including by some of the major players. 2. LLMs have significant problems and limitations. 3. LLMs can do some incredibly impressive things and can be profoundly useful for some applications. I would go so far as to say that #2 and #3 are hardly even debatable at this point. Everyone acknowledges #2, and the only people I see denying #3 are people who either haven't investigated or are so annoyed by #1 that they're willing to sacrifice their credibility as an intellectually honest observer.
- absolutelastone 2y ago#3 can be true and yet not be enough to make your case. Many failed technologies achieved impressive engineering milestones. Even the harshest critic could probably brainstorm some niche applications for a hallucination machine or whatever.
- fragmede 2y agoAnd yet we keep electing them to public office.
- svrtknst 2y agoYou, and the OP, are being unfair in your replies. Obviously, it's not worthless for all applications but when LLMs obviously fail in disastrous ways in some important areas, you can't refute that by going "actually it gives me codign advice and generates images". Thats nice and impressive, but there are still important issues and shortcomings. Obligatory, semirelated xkcd: https://xkcd.com/937/ https://xkcd.com/937/
- icepat 2y ago
- w0m 2y ago"an old poorly implemented model can't do item X well therefore the technology is garbage" Likely the most accurate measure of progress would be watching detractors goalposts move over time.
- jodrellblank 2y ago"Even a journey of 1,000 miles begins with the first step. Unless you're an AI hyper then taking the first step is the entire journey - how dare you move the goalposts"
- KoolKat23 2y agoThis is more a lack of understanding of it's limitations, it'd be different if they asked for it to write a python script to collate the data.
- pfdietz 2y agoAh, it's like communism, then (to its diehards). It cannot fail, it can only be failed.
- xigoi 2y agoIf the LLM is intelligent, why can’t it figure out that writing a script would be the best way to solve the problem?
- DarmokJalad1701 2y agoSome of the more modern tools do exactly that. If you upload a CSV to Claude, it will not (or at least not anymore) try to process the whole thing. It will read the header, and then ask you what you want. It will then write the appropriate Javascript code and run it to process the data and figure out the stats/whatever you asked it for. I recently did this with a (pretty large) exported CSV of calories/exercise data from MyFitnessPal and asked it to evaluate it against my goals/past bloodwork etc (which I have in a "Claude Project" so that it has access to all that information + info I had it condense and add to the project context from previous convos). It wrote a script to extract out extremely relevant metrics (like ratio of macronutrients on a daily basis for example), then ran it and proceeded to talk about the result, correlating it with past context. Use the tools properly and you will get the desired results.
- mwigdahl 2y agoAll of these anecdotal stories about "LLM" failures need to go into more detail about what model, prompt, and scaffolding was used. It makes a huge difference. Were they using Deep Research, which searches for relevant articles and brings facts from them into the report? Or did they type a few sentences into ChatGPT Free and blindly take it on faith? LLMs are _tools_, not oracles. They require thought and skill to use, and not every LLM is fungible with every other one, just like flathead, Phillips, and hex-head screwdrivers aren't freely interchangeable.
- spamizbad 2y agoIf any non-trivial ask of an LLM also requires the prompts/scaffolding to be listed, and independently verified, along with its output, their utility is severely diminished. They should be saving time not giving us extra homework. Far better to just get these problems resolved.
- mwigdahl 2y agoThat isn't what I'm saying. I'm saying you can't make a blanket statement that LLMs in general aren't fit for some particular task. There are certainly tasks where no LLM is competent, but for others, some LLMs might be suitable while others are not. At least some level of detail beyond "they used an LLM" is required to know whether a) there was user error involved, or b) an inappropriate tool was chosen.
- butlike 2y agothen they shouldn't market it as one-size fits all
- mwigdahl 2y agoAre they? Every foundation model release includes benchmarks with different levels of performance in different task domains. I don't think I've seen any model advertised by its creating org as either perfect or even equally competent across all domains. The secondary market snake oil salesmen <cough>Manus</cough>? That's another matter entirely and a very high degree of skepticism for their claims is certainly warranted. But that's not different than many other huckster-saturated domains.
- freilanzer 2y agoSo they used the model as a database? It should be immediately obvious to anyone that this won't work.
- jabroni_salad 2y agoWell yeah, it's fancy autocomplete. And it's extremely amazing what 'fancy autocomplete' is able to do, but making the decision to use an LLM for the type of project you described is effectively just magical thinking. That isn't an indictment against LLM, but rather the person who chose the wrong tool for the job.
- eric_cc 2y agoSounds like a user problem, though. When used properly as a tool they are incredible. When you give up 100% trust to them to be perfect it’s you that is making the mistake.
- agentcoops 2y agoThe whole point is that an LLM is not a search engine and obviously anyone who treats it as one is going to be unsatisfied. It's just not a sensible comparison. You should compare working with an LLM to working with an old "state of the art" language tool like Python NLTK -- or, indeed, specifying a problem in Python versus specifying it in the form of a prompt -- to understand the unbridgeable gap between what we have today and what seemed to be the best even a few years ago. I understand when a popular science author or my relatives haven't understood this several years after mass access to LLMs, but I admit to being surprised when software developers have not. Hosted and free or subscription-based DeepResearch like tools that integrate LLMs with search functionality (the whole domain of "RAG" or "Retrieval Augmented Generation") will be elementary for a long time yet simply because the cost of the average query starts to go up exponentially and there isn't that much money in it yet. Many people have and will continue to build their own research tools where they can determine how much compute time and API access cost they're willing to spend on a given query. OCR remains a hard problem, let alone appropriately chunking potentially hundreds of long documents into context length and synthesizing the outputs of potentially thousands of LLM outputs into a single response.
- throwawaymaths 2y agoto be fair a few? one? years ago LLMs were touted? marketed? as a "search killer", and a lot of people do use it in that fashion.
- joquarky 2y agoA lot of people need to improve their critical thinking skills to deconstruct the marketing hype, and then choose the right tool for the job.
- throwawaymaths 2y agosure. isn't that effectively what Sabine is doing though? She just doesn't have as compelling a use in the cases where LLMs are strong.
- Terretta 2y ago"They continue to fabricate links, references, and quotes, like they did from day one." - "I ask them to give me a source for an alleged quote, I click on the link, it returns a 404 error." Why have these companies not manually engineered out a problem like this by now? Just do a check to make sure links are real. That's pretty unimpressive to me. There are no fabricated links, references, or quotes, in OpenAI's GPT 4.5 + Deep Research. It's unfortunate the cost of a Deep Research bespoke white paper is so high. That mode is phenomenal for pre-work domain research. You get an analyst's two week writeup in under 20 minutes, for the low cost of $200/month (though I've seen estimates that white paper cost OpenAI over USD 3000 to produce for you, which explains the monthly limits). You still need to be a domain expert to make use of this, just as you need to be to make use of an analyst. Both the analyst and Deep Research can generate flawed writeups with similar misunderstandings: mis-synthesizing, misapplication, or missing inclusion of some essential. Neither analyst nor LLM is a substitute for mastery.
- fridder 2y agoWhile I agree, it doesn't stop business folks pushing for its use in area where it is inappropriate. That is, at least for me, part of the skepticism.
- blactuary 2y agoHow do people in the future become domain experts capable of properly making use of it if they are not the analyst spending two weeks on the write-up today?
- never_inline 2y agoMy complaints with Deep Research LLMs is they don't go deeper than 2 pages of SERPs. I want them to dig down obscure stuff, not list cursorily relevant peripheral directions. they just seem to do breadth first than depth first search.
- whamlastxmas 2y agoI’m sorry but the experience of coding with an LLM is about ten billion times better than googling and stack overflowing every single problem I come across. I’ve stack overflowed maybe like two things in the past half year and I’m so glad to not have to routinely use what is now a very broken search engine and web ecosystem.
- quonn 2y agoIt‘s broken now. It was fine 5 years ago.
- player1234 2y agoHow did you measure and compare googling/stack overflow to coding with an LLM? How did you get to the very impressive number ten billion times better?! Can you share your methodology? How have you defined better?
- whamlastxmas 2y agoI take calipers to my boss’s forehead veins and see how pissed he is routinely throughout the day
- blactuary 2y agoThe search ecosystem is broken now because google is focused on LLMs
- internet101010 2y agoThat's part of it. The other part is Google sacrificing product quality for excessive monetization. An example would be YouTube search - first three results are relevant, next 12 results are irrelevant "people also watched", then back to relevant results. Another example would be searching for an item to buy and getting relevant results in the images tab of google, but not the shopping tab.
- whamlastxmas 2y agoIt’s broken bc google has spent 20+ years promoting garbage content in a self-serving way. No one was able to compete unless they played by googles rules, and so all we have left is blog spam and regular spam
- casey2 2y agoThe 404 links are hilarious, like you can't even parse the output and retry until it returns a link that doesn't 404? Even ignoring the billions in valuation, this is so bad for a $20 sub.
- vonneumannstan 2y ago[flagged]
- giantrobot 2y ago> This is just not a use case where the expected performance on these tasks is high. Yet the hucksters hyping AI are falling all over themselves saying AI can do all this stuff. This is where the centi-billion dollar valuations are coming from. It's been years and these super hyped AIs still suck at basic tasks. When pre-AI shit Google gave wrong answers it at least linked to the source of the wrong answers. LLMs just output something that looks like a link and calls it a day.
- vonneumannstan 2y agoTo be fair the newest tools like Deep Research are actually quite good and hallucination is essentially not a real problem for them. https://marginalrevolution.com/marginalrevolution/2025/02/deep-research.html https://marginalrevolution.com/marginalrevolution/2025/02/de...
- frm88 2y ago<<After glowing reviews, I spent $200 to try it out for my research. It hallucinated 8 of 10 references on a couple of different engineeribg topics. For topics that are well established (literature search), it is useful, although o3-mini-high with web search worked even better for me. For truly frontier stuff, it is still a waste of time.>> <<I've had the hallucination problem too, which renders it less than useful on any complex research project as far as I'm concerned.>> These quotes are from the link you posted. There are a lot more.
- vonneumannstan 2y agoI think Sabine is just wrong in this case. I don't think Deep Research can even hallucinate links in this way at all.
- 2y ago
- deleted 2y ago[deleted]
- eric_cc 2y agoThe tweeters complaints sound like a user problem. LLM’s are tools. How you use them, when you use them, and what you expect out of them should be based on the fact they are tools.
- waffletower 2y agoThis assessment is incomplete. Large languages models are both less and more than these traditional tools. They have not subsumed them and all can sit together in separate tabs of a single browser window. They are another resource, and when the conditions are right, which is often the case in my experience, they are a startlingly effective tool for navigating the information landscape. The criticism of Gemini is a fair one, and I encountered it yesterday, but perhaps with 50% less entitlement. But Gemini also helped me translate obscure termios APIs to python from C source code I provided. The equivalent using search and/or Stack Overflow would have required multiple piecemeal searches without guarantees -- and definitely would have taken much more time.