6 ms·
> No human could read all of this in a lifetime. AI consumes it in seconds. And therefore it's impossible to test the accuracy if it's consuming your own data.
by bdbdbdb 10mo ago
> No human could read all of this in a lifetime. AI consumes it in seconds.
And therefore it's impossible to test the accuracy if it's consuming your own data. AI can hallucinate on any data you feed it, and it's been proven that it doesn't summarize, but rather abridges and abbreviates data.
In the authors example
> "What patterns emerge from my last 50 one-on-ones?" AI found that performance issues always preceded tool complaints by 2-3 weeks. I'd never connected those dots.
Maybe that's a pattern from 50 one-on-ones. Or maybe it's only in the first two and the last one.
I'd be wary of using AI to summarize like this and expecting accurate insights
- gchamonlive 10mo ago> it's been proven that it doesn't summarize, but rather abridges and abbreviates data Do you have more resources on that? I'd love to read about the methodology. > And therefore it's impossible to test the accuracy if it's consuming your own data. Isn't it only if it's hard to verify the result? If it's a result that's hard to produce but easy to verify, a class which many problems fall into, you'd just need to look at the synthetized results. If you ask it "given these arbitrary metrics, what is the best business plan for my company?" It'd be really hard to verify the result. I'd be hard to verify the result from anyone for that matter, even specialists. So I think it's less about expecting the LLM to do autonomous work and more about using LLMs to more efficiently help you search the latent space for interesting correlations, so that you and not the LLM come up with the insights.
- bdbdbdb 10mo ago> If you ask it "given these arbitrary metrics, what is the best business plan for my company?" It'd be really hard to verify the result. I'd be hard to verify the result from anyone for that matter, even specialists. Hard to verify something so subjective, for sure. But a specialist will be applying intelligence to the data. An LLM is just generating random text strings that sound good. The source for my claim about LLMs not summarizing but abbreviating is on hn somewhere, I'll dig it out Edit: sorry, I tried but couldn't find the source.
- gchamonlive 10mo ago> But a specialist will be applying intelligence to the data. An LLM is just generating random text strings that sound good. I'd only make such a claim if I could demonstrate that human text is a result of intelligence and LLMs not, because really, what's the actual difference? How isn't LLM "intelligent" when it can clearly help me make sense of information? Note that this isn't to say that it's conscious or not. But it's definitely intelligent. The text output is not only coherent, it's right often enough to be useful. Curiously, I'm human, and I'm wrong a lot, but I'm right often enough to be a developer.
- thoughtpeddler 10mo agoLook into the emerging literature around "needle-in-a-haystack" tests of LLM context windows. You'll see what the poster you're replying to is describing, in part. This can also be described as testing "how lazy is my LLM being when it comes to analyzing the input I've provided to it?" Hint: they can get quite lazy! I agree with the poster you replied to that "RAG my Obsidian"-type experiments with local models are middling at best. I'm optimistic things will get a lot better in the future, but it's hard to trust a lot of the 'insights' this blog post talks about, without intense QA-ing (if the author did it, which I doubt, considering their writing is also lazily mostly AI-assisted as well).
- block_dagger 10mo agoYour colleagues using the tech will be far ahead of you soon, if they aren’t already.
- iLoveOncall 10mo agoFar ahead in producing bugs, far ahead in losing their skills, far ahead in becoming irrelevant, far ahead in being unable to think critically, that's absolutely right.
- afandian 10mo ago"The market can stay irrational longer than you can stay solvent" feels relevant here.
- pitched 10mo agoThe new tools have sets of problems they are very good at, sets they are very bad at and they are generally mediocre at everything else. Learning those lessons isn’t easy, takes time, and will produce bugs. If you aren’t making those mistakes now with everyone else, you’ll be doing them later when you do decide to start catching up and it will be more noticeable then.
- _DeadFred_ 10mo agoAnd all of those things (good at, bad at, the lessons learned on current models current implementation) can change arbitrarily with model changes, nudges, guardrails, etc. Not sure that outsourcing your skillset on the current foundation of sand is long term smart, even if it's great for a couple of months. It may be those un-learning the previous iteration interactions once something stable arrives that are at a disadvantage?
- evilduck 10mo agoWhy would the AI skeptics and curmudgeons today not continue to dismiss the "something stable" in the future?
- kenjackson 10mo agoSimilar to P/NP, verification can often be faster than solving. For example, you can then ask the AI to give you the list of tool complaints and the performance issues. Then a text search can easily validate the claim.
- xtiansimon 10mo ago> “I'd be wary of using AI to summarize like this and expecting accurate insights.” Sure, but when do you have accurate results when using an iterative process? It can happen at the beginning or at the end when you’re bored, or have exhausted your powers of interrogation. Nevertheless, your reasoning will tell you if the AI result is good, great, acceptable, or trash. For example, you can ask Chat—Summarize all 50 with names, dates and 2-3 sentence summaries and 2-3 pull quotes. Which can be sufficient to jog your memory, and therefore validate or invalidate the Chat conclusion. That’s the tool, and its accuracy is still TBD. I for one am not ready to blindly trust our AI overlords, but darn if a talking dog isn’t worth my time if it can make an argument with me.
- potsandpans 10mo ago> ...and it's been proven that it doesn't summarize, but rather abridges and abbreviates data. I don't really know what this means, or if the distinction is meaningful for the majority of cases.
- deleted 10mo ago[deleted]
- TimByte 10mo agoI think as long as you keep a skeptical loop and force the model to cite or surface raw notes, it can still be useful without being blindly trusted
- missedthecue 10mo ago"AI can hallucinate on any data you feed it, and it's been proven that it doesn't summarize, but rather abridges and abbreviates data." Have you ever met a human? I think one of the biggest reasons people become bearish on AI is that their measure of whether it's good/useful is that it needs to be absolutely perfect, rather than simply superior to human effort.
- bigstrat2003 10mo agoRight now AI is inferior, not superior, to human effort. That's precisely why people are bearish on it.
- missedthecue 10mo agoI don't think thats obvious. In 20 minutes for example, deep research can write a report on a given topic much better than an analyst can produce in a day or two. It's literally cheaper, better, and faster than human effort.
- jrflowers 10mo agoWhat do you man by “better” in this context?
- fcantournet 10mo agoIt has more words put together in seemingly correct sentences, so it's long enough his boss won't actually read it to proof it.
- missedthecue 10mo agoIt synthesizes a more comprehensive report, using more sources, more varied sources, more data, and broader insights than a human analyst can produce in 1-2 days of research and writing. I'm not confused about this. If you don't agree, I will assume it's probably because you've never employed a human to do similar work in the past. Because it's not particularly close. It's night and day. *Note that I'm not saying 20 minutes of deep research beats 9 months of investigative journalism with private interviews with primary sources or anything like that. I'm talking about asking an analyst on your team to do a deep dive into XYZ and have something on your desk tomorrow EOD.
- novok 10mo agoAI is a new kind of bulk tool, you need to know how to use it well and context management is a huge part of it. For that 1-1 example, you would do a for loop with new context with subagents or a literal for loop for example to prevent the 'first two and last one' issue. Then with those 1-1 summaries, look at that to make the determination for example. Humanity has gotten amazing results from unreliable stochastic processes, managing humans in organizations is an example of that. It's ok if something new is not completely deterministic to still be incredibly useful.