9 ms·
The Differences Between Deep Research, Deep Research, and Deep Research
- blacksqr 2y agoYes, but what about Deep Research?
- readyplayernull 2y agoAre you telling me that AI's are starting to diverge and that we might get a combinatorial explosion of reasoning paths that will produce so many different agents that we won't know which one can actually become AGI? https://leehanchung.github.io/assets/img/2025-02-26/05-quadrants.png https://leehanchung.github.io/assets/img/2025-02-26/05-quadr...
- nwhnwh 2y agoLife is so complicated.
- pseudocomposer 2y agoThey are certainly diverging, and becoming more useful tools, but the answer to “which one can actually become AGI?” is, as always, “none of them.”
- OutOfHere 2y agoAGI is about performing actions with a high multi-task intelligence. Only the top right corner (Deep, Trained) has any hope of getting closer to AGI. The rest can still be useful for specific tasks, e.g. "deep-research".
- dartos 2y agoAGI is, and always was, marketing from LLM providers. Real innovation is going in with task-specific models like AlphaFold. LLMs are starting to become more task specific too, as we’ve seen with the performance of reasoning model on their specific tasks. I imagine we’ll see LLMs trained specifically for medical purposes, legal purposes, code purposes, and maybe even editorial purposes. All useful in their own way, but none of them even close to sci-fi.
- HeatrayEnjoyer 2y agoAGI is a term that's been around for decades.
- dartos 2y agoAI was a term that was around for decades as well, but it’s meaning dramatically changed over the past 3 years. Prior to gpt-3 AI was rarely used in marketing or to talk about any number of ML methods. Nowadays “AI” is just the new “smart” for marketing products. Terms change. The current usage of AGI, especially in the context I was talking about, is specifically marketing from LLM providers. I’d argue that the term AGI, when used in a non fiction context, has always been a meaningless marketing term of some kind.
- mdp2021 2y ago> has always been Well now it is not: it is now "the difference between something with outputs that sounds plausible vs something with outputs which are properly checked".
- TeMPOraL 2y agoFeels like you were born just yesterday :). > Prior to gpt-3 AI was rarely used in marketing or to talk about any number of ML methods. In the decade prior to GPT-3, AI was frequently used in marketing to talk about any ML methods, up to and including linear regression. This obviously ramped up heavily after "Deep Learning" got coined as a term. AI now actually means something in marketing, but the only reason for that is that calling out to an LLM is even simpler than adding linear regression to your product somewhere. As for AGI, that was a hot topic in some circles (that are now dismissed as "AI doomers") for decades. In fact, OpenAI started with people associating or at least within the sphere of influence of LessWrong community, which both influenced the naming and perspective the "LLM industry" started with, and briefly put the outputs of LessWrong into spotlight - which is why now everyone uses terms like "AGI" and "alignment" and "AI safety". However, unlike "alignment", which got completely butchered as a term, AGI still roughly means what it meant before - which is basically a birth of a man-made god. That's true of AGI as "meaningless marketing term" too, if people so positive on it paused to follow through the implications beyond "oh it's like ChatGPT, but actually good at everything".
- sgt101 2y agoIsn't this the worst possible case for an LLM? The integrity of the product is central to the value of it and the user is by definition unable to verify that integrity?
- falcor84 2y agoI'm not following, what's "by definition" here? You can verify the integrity of an AI report in the same way you would do with any other report someone prepared for you - when you encounter something that feels wrong, check the referenced source yourself.
- sgt101 2y agoI think it's normal to invest authority in a report that someone else has prepared - people take what is written on trust because the person who prepared it is accountable for any errors... not just now, but forever. LLM's are not accountable for anything.
- falcor84 2y agoThat's a lofty ideal, but the typical research report is flawed in multiple ways, and even in academia "Most Published Research Findings Are False" [0], and very few people's careers have in any way suffered, as there's very little de-facto accountability for this trust. The best case scenario as I see it, is that we become more critical of research in general, and start training AI to help us identify these issues. And then we could perhaps utilize GAN to improve report generation. [0] https://en.wikipedia.org/wiki/Replication_crisis https://en.wikipedia.org/wiki/Replication_crisis
- falcor84 2y agoP.S. There's a paper from today demonstrating the effectiveness of AI identifying errors in human research: https://news.ycombinator.com/item?id=43295692 https://news.ycombinator.com/item?id=43295692
- simonw 2y agoI really like the distinction between DeepSearch and DeepResearch proposed in this piece by Han Xiao: https://jina.ai/news/a-practical-guide-to-implementing-deepsearch-deepresearch/ https://jina.ai/news/a-practical-guide-to-implementing-deeps... > DeepSearch runs through an iterative loop of searching, reading, and reasoning until it finds the optimal answer. [...] > DeepResearch builds upon DeepSearch by adding a structured framework for generating long research reports Given these definitions, I think DeepSearch is the more valuable and interesting pattern. It's effectively RAG built using tools in a loop, which is much more likely to answer questions effectively than more traditional RAG where there is only one attempt to find relevant documents to include in a single prompt to an LLM. DeepResearch is a cosmetic enhancement that wraps the results in a "report" - it looks impressive but IMO is much more likely to lead to inaccurate or misleading results. More notes here: https://simonwillison.net/2025/Mar/4/deepsearch-deepresearch/ https://simonwillison.net/2025/Mar/4/deepsearch-deepresearch...
- derefr 2y ago> DeepResearch is a cosmetic enhancement that wraps the results in a "report" No, that's not what Xiao said here. Here's the relevant quote > It often begins by creating a table of contents, then systematically applies DeepSearch to each required section – from introduction through related work and methodology, all the way to the conclusion. Each section is generated by feeding specific research questions into the DeepSearch. The final phase involves consolidating all sections into a single prompt to improve the overall narrative coherence. (I also recommend that you stare very hard at the diagrams.) Let me paraphrase what Xiao is saying here: A DeepSearch is a primitive — it does mostly the same thing a regular LLM query does, but with a lot of trained-in thinking and searching work, to ensure that it is producing a rigorous answer to your question. Which is great: it means that DeepSearch is more likely to say "I don't know" than to hallucinate an answer. (This is extremely important as a building block; an agent needs to know when a query has failed so it can try again / try something else.) However, DeepSearch alone still "hallucinates" in one particular way: it "hallucinates understanding" of the topic, thinking that it already has a complete mental toolkit of concepts needed to solve your problem. It will never say "solving this sub-problem seems to require inventing a new tool" and so "branch off" to another recursed DeepSearch to determine how to do that. Instead, it'll try to solve your problem with the toolkit it has — and if that toolkit is insufficient, it will simply fail. Which, again, is great in some ways. It means that a single DeepSearch will do a (semi-)bounded amount of work. Which means that the costs of each marginal additional DeepSearch call are predictable. But it also means that you can't ask DeepSearch itself to: • come up with a mathematical proof of something, where any useful proof strategy will implicitly require inventing new math concepts to use as tools in solving the problem. • do investigative journalism that involves "chasing leads" down a digraph of paths; evaluating what those leads have to say; and using that info to determine new leads. • "code me a Facebook clone" — and have it understand that doing so involves iteratively/recursively building out a software architecture composed of many modules — where it won't be able to see the need for many of those modules at "design time", but will only "discover" the need to write them once it gets to implementation time of dependent modules and realizes that to achieve some goal, it must call into some code / entire library that doesn't exist yet. (And then make a buy-vs-build decision on writing that code vs pulling in a dependency... which requires researching the space of available packages in the ecosystem, and how well they solve the problem... and so on.) A DeepResearch model, meanwhile, is a model that looks at a question, and says "is this a leaf question that can be answered directly — or is this a question that needs to be broken down and tackled by parts, perhaps with some of the parts themselves being unknowns until earlier parts are solved?" A DeepResearch model does a lot of top-level work — probably using DeepSearch! — to test the "leaf-ness" of your question; and to break down non-leaf questions into a "battle plan" for solving the problem. It then attempts solutions to these component problems — not by calling DeepSearch, but by recursively calling itself (where that forked child will call DeepSearch if the sub-problem is leaf-y, or break down the sub-problem further if not.) A DeepResearch model will then takes the derived solutions for dependent problems into account in the solution space for depending problems. (A DeepResearch model may also be trained to notice when it's "worked into a corner" by coming up with early-phase solutions that make later phases impossible; and backtracking to solve the earlier phases differently, now with in-context knowledge of the constraints of the later phases.) Once a DeepResearch model finds a successful solution to all subproblems, it takes the hierarchy of thinking/searching logs it generated in the process, and strips out all the dead-ends and backtracking, to present a comprehensible linear "success path." (Probably it does this as the end-step of each recursive self-call, before returning to self, to minimize the amount of data returned.) Note how this last reporting step isn't "generating a report" for human consumption; it's a DeepResearch call "generating a report" for its parent DeepResearch call to consume. That's special sauce. (And if you think about it, the top-level call to this whole thing is probably going to use a non-DeepResearch model at the end to rephrase the top-level DeepResearch result from a machine-readable recurse-result report into a human-readable report. It might even use a DeepSearch model to do that!) --- Bonus tangent: Despite DeepSearch + DeepResearch using a scientific-research metaphor, I think an enlightening comparison is with intelligence agencies. DeepSearch alone does what an individual intelligence analyst does. You hand them an individually-actionable question; they run through a "branching, but vaguely bounded in time" process of thinking and searching, generating a thinking log in the process, eventually arriving at a conclusion; they hand you back an answer to your question, with a lot of citations — or they "throw an exception" and tell you that the facts available to the agency cannot support a conclusion at this time. Meanwhile, DeepResearch does what an intelligence agency as a whole does: 1. You send the agency a high-level strategic Request For Information; 2. the agency puts together a workgroup composed of people with trained-in expertise with breaking down problems (Intelligence Managers), and domain-matter experts with a wide-ranging gestalt picture of the problem space (Senior Intelligence Analysts), and tasks them with breaking down the problem into sub-problems; 3. some of these sub-problems are actionable — they can be assigned directly for research by a ground-level analyst; some of these sub-problems have prerequisite work that must be done to gather intelligence in the field; and some of these sub-problems are unknown unknowns — missing parts of the map that cannot be "planned into" until other sub-problems are resolved. 4. from there, the problem gets "scheduled" — in parallel, (the first batch of) individually-actionable questions get sent to analysts, and any field missions to gather pre-requisite intelligence are kicked off for planning (involving spawning new sub-workgroups!) 5. the top-level workgroup persists after their first meeting, asynchronously observing the reports from actionable questions; scheduling newly-actionable questions to analysts once field data comes in to be chewed on; and exploring newly-legible parts of the map to outline further sub-problems. 6. If this scheduling process runs out of work to schedule, it's either because the top-level question is now answerable, or because the process has worked itself into a corner. In the former case, a final summary reporting step is kicked off, usually assigned to a senior analyst. In the latter case, the workgroup reconvene to figure out how to backtrack out of the corner and pursue alternate avenues. (Note that, if they have the time, they'll probably make "if this strategy produces results that are unworkable in a later step" plans for every possible step in their original plan, in advance, so that the "scheduling engine" of analyst assignments and fieldwork need never run dry waiting for the workgroup to come up with a new plan.)
- EncomLab 2y agoOne of my co-workers joked at the time that "sure AlphaGO beat Lee Sedol at GO, but Lee has a much better self-driving algorithm." I thought this was funny at the time, but I think as more time passes it does highlight the stark gulf that exists between the capability of the most advanced AI systems and what we expect as "normal competency" from the most average person.
- tsunego 2y agoLove me some good old whataboutism (sure, LLMs are now super-intelligent at writing software, but can they clean my kitchen? No? Ha!)
- jvanderbot 2y agoThe computer beat me at chess, but it was no match for me at kickboxing. - Emo Phillips Tale as old as time. We can make nice software systems but general purpose AI / Agents isn't here yet.
- TeMPOraL 2y agoWorse than that: it seems that it's much easier to make computer achieve superhuman feats in cognitive work, than it is to make it do even most basic physical interactions with the real world. In short: the natural order of things is that computers are better at thinking, and people are better at manual labor. Which is the opposite of what we wanted.
- jvanderbot 2y agoAI is just hydraulics for the mind. Or should be. I choose a direction and apply force.
- deleted 2y ago[deleted]
- stavros 2y ago> it does highlight the stark gulf that exists between the capability of the most advanced AI systems and what we expect as "normal competency" from the most average person Yes, but now we're at the point where we can compare AI to a person, whereas five years ago the gap was so big that that was just unthinkable.
- linwangg 2y ago[dead]
- jimmySixDOF 2y agoThis gives STORM a high mark but didn't seem to get great results from GPT Researcher which is the other open source project that was doing this before the recent flavor of the day DeepReasearch has become. But there are so many ways to configure GPT Researcher for all kinds of budgets so I wonder if this comparison really pushed the output or just went with defaults and got default midranges for comparison.
- tsunego 2y agoNeat summary but you forgot Grok!
- jvanderbot 2y agoMaybe the article has been edited in the last four minutes since you posted, but Grok is definitely in there now.
- giancarlostoro 2y agoIts interesting it says Grok excels at report generation, because I've found myself asking it to give me answers in a table format, to make it easier to 'grok' at the output, since I'm usually asking it to give me comparisons I just can't do natively on Amazon or any other ecommerce site. Funnily enough, Amazon will pick for you products to compare, but the compared items usually are terrible, and you can't just add whatever you want, or choose columns. With Grok, I'll have it remove columns, add columns, shorten responses, so on and so forth.
- ankit219 2y agoThink this captures one of the bigger differences between what Open AI offers and what others offer using the same name. Funnily enough, Google's Gemini 2.0 Flash also has a native integration to google search[1]. They have not done it with their Thinking model. When they do we will have a good comparison. One of the implications of OpenAI's DR is that frontier labs are more likely to train specific models for a bunch of tasks, resulting in the kind of quality wrappers will find hard to replicate. This is leading towards model + post training RL as a product, instead of keeping them separate from the final wrapper as product. Might be interesting times if the trajectory continues. PS: There is also genspark MOA[2] which creates an indepth report on a given prompt using mixtures of agents. From what i have seen in 5-6 generations, this is very effective. [1]: https://x.com/_philschmid/status/1896569401979081073 https://x.com/_philschmid/status/1896569401979081073 (i might be misunderstanding this, but this seems a native call instead of explicit) [2]: https://www.genspark.ai/agents?type=moa_deep_research https://www.genspark.ai/agents?type=moa_deep_research
- cxie 2y agodeep search is the new RAG
- TeMPOraL 2y agoDeep Search is RAG - that is, if we're still expanding the acronym instead of treating it as a word that just means "queries a vector database". Prediction for Next Hot Thing in Q4 2025 / Q1 2026: someone will make the Nobel prize-worthy discovery that you can stuff results of your deep search into a database (vector or otherwise) and then use it to improve the ability to compile a higher-quality report from much larger amount of sources. We'll call it DeepRAG or Retrieval Augmented Deep Research or something. Prediction for Q2 2026: next Nobel prize awarded for realizing you may as well stop treating report generation as the core aspect of "deep research" (as it obviously makes no sense, but hey, time traveler's spoilers, sorry!), and stop at the "stuff search results into a database" and let users "chat with the search results", making the research an interactive process. We'll call this huge scientific breakthrough "DeepRAG With Human Feedback", or "DRHF".
- eamag 2y agoWhy is there a citation @article part in the end? Do people actually use it?
- SubiculumCode 2y agoThe primary issue with deep research tools are veracity and accurate source attribution. My issue with tools relying on DeepDeek R, for example, is the high hallucination rate.
- EigenLord 2y agoDR is a nice way to gather information, when it works, and then do the real research yourself from a concentrated launching point. It helps me avoid ADD braining myself into oblivion every time I search the internet. The fatal mistake is thinking that the LLM is now wiser for having done it. When someone does their research, they are now marginally more of an authority on that topic than everyone else in the room, all else being equal. But for LLMs, it's not like they have suddenly acquired more expertise on the subject now that they did this survey. So it's actually pretty shallow, not deep, research. It's a cool capability, and a nifty way to concentrate information, but much deeper capabilities will be required to have models that not only truly synthesize all that information, but actively apply it to develop a thesis or further a research program. Truthfully, I don't see how this is possible within the transformer architecture, with its lack of memory or genuine statefulness and therefore absence of persistent real time learning. But I am generally a big transformer skeptic.
- instagraham 2y agoI noticed that these models very quickly start to underperform regular search like Perplexity Pro 3x. It might be somewhat thorough in how it goes line by line, but it's not very cognizant of good sources - you might ask for academic sources but if your social media slider is turned on, it will overwhelmingly favour Reddit. You may repeat instructions multiple times, but it ignores them or fails to understand a solution
- ipsum2 2y agoYou have to be specific about which model you're referring to. OpenAI DeepResearch does not have a slider, and does follow when you say to only use academic sources.
- dudeinhawaii 2y agoAs a user, I've found that researching the same topics in OpenAI Deep Research vs Perplexity's Deep Research results in "narrow and deep" vs "shallow and broad". OpenAI tends to have something like 20 high quality sources selected and goes very deep in the specific topic, producing something like 20-50 pages of research in all areas and adjacent areas. It takes a lot to read but is quite good. Perplexity tends to hit something like 60 or more sources, goes fairly shallow, answers some questions in general ways but is excellent at giving you the surface area of the problem space and thoughts about where to go deeper if needed. OpenAI takes a lot longer to complete, perhaps 20x longer. This factors heavily into whether you want a surface-y answer now or a deep answer later.
- bilater 2y agoI wet through this journey myself with Deep Search / Research https://github.com/btahir/open-deep-research https://github.com/btahir/open-deep-research I think it really comes down to your own workflow. You sometimes want to be more imperative (select the sources yourself to generate a report) and sometimes more declarative (let a DFS/BFS algo go and split a query into subqueries and go down rabbit holes until some depth and then aggregate). Been trying different ways of optimizing the former but I am fascinated by the more end to end flows systems like STORM do.
- toisanji 2y agowhat is the best open source system to use?
- z3c0 2y ago> In natural language processing (NLP) terms, this is known as report generation. I'm happy to see some acknowledgement of the world before LLMs. This is an old problem, and one I (or my team, really) was working on at the time of DALL-E & ChatGPT's explosion. As the article indicated, we deemed 3.5 unacceptable for Q&A almost immediately, as the failure rate was too high for operational reporting in such a demanding industry (legal). We instead employed SQuAD and polished up the output with an LLM. These new reasoning models that effectively retrofit Q&A capabilities (an extractive task) onto a generative model are impressive, but I can't help but think that it's putting the cart before the horse and will inevitably give diminishing returns in performance. Time will tell, I suppose.
- vednig 2y agoIt's amazing how these are the biggest information organizing platforms in the internet and yet they fail to find different words to describe their products.
- ein0p 2y agoI think "deep research" is a misnomer, possibly deliberate. Research assumes the ability to determine quality, freshness, and veracity of the sources directly from their contents. It also quite often requires that you identify where the authors screwed up, lied, chose deliberately bad baselines, and omitted results in order to make their work "more impactful". You'd need AGI to do that. This is merely search - it will search for sources, summarize them for you, and write a report, but you then have to go in and _verify it's not bullshit_, which can take quite a bit of time and effort, if you care about the quality of the final result, which for big ticket questions you almost always do.