12 ms·
Elicit – AI Research Assistant
- zerojames 2y agoI have run a few queries on Elicit to understand the product a bit more. I asked about media bias detection and used the topic analysis feature. A minute or so later, I had a list of concepts with citations and links to papers I can look at further. This feels like an _amazing_ tool to do literature overviews and to dive into new academic domains with which one is not familiar.
- MrWB 2y agoTry looking at this https://inciteful.xyz/ https://inciteful.xyz/ there’s also a Zotero plugin for it.
- jamesbrady 2y agoElician here: thank you for sharing our tool and for this praise! We're glad you're enjoying it.
- vouaobrasil 2y agoTo be honest, having everything in research being so mechanized is starting to make research itself feel less magical and just about being another part in a machine. With that and the fact that most research today is relatively meaningless or focused on optimization of commercial products, I think science has left the domain of the purely curious and has been taken over by bean counters. No interest in such a product of any kind.
- jordanpg 2y agoI'm a patent attorney. One fantasy that I have is that AI will be a forcing function that will drive all the needless verbosity out of patents. For example, every patent is 50 or 100 pages of what amounts to bland, unreadable background. The stuff that is actually new in the world is usually a tiny fraction of all the words. In a world where everyone's AI can generate all those extra words, their value goes down. (Yes, this is bad news for people in my line of work who get paid to write all those words.) So maybe the profession moves in a direction that values conciseness and brevity, to re-add value to what the lawyers and agents do, that the AI can't do. It's a fantasy. But the same kind of thing applies to academic publishing. Maybe the future of publishing is very short, useful documents with rich, tiered, AI-generated hyperlinks on every word.
- fkyoureadthedoc 2y agoI wouldn't expect that to happen. In a world where AI makes it both easier to generate and process all those extra words why would they actually decrease? It seems like the motivation would be stronger to reduce it if it's humans doing the majority of the generation and processing.
- martindbp 2y agoUnfortunately, AI could lead to more verbose everything, because it also enables us to distil huge amounts of text. Forgot where I read about this, but it has been shown that digitalization has enabled the tax code to go from something one person can read and understand to something completely unmanageable (could be from The Utopia of Rules: On Technology, Stupidity, and the Secret Joys of Bureaucracy), because digital tools allows us to deal with more tax code. Similarly, doctors, teachers and many other occupations are drowning in administration duties that were enabled by digital technology. I hope I'm wrong. Eventually, when only AI produces and consumes text, it could be more brief, but at that point will it matter? Eventually a less ambiguous format could be used, something more like code, that computers and AI can consume and apply.
- jiggawatts 2y agoThe opposite happens with technology. As the cost goes down, usage goes up. (Jevons paradox)
- inigoalonso 2y agoI like that scenario. But wouldn't AI just explode the verbosity of patents instead? The value of the humans' word count might go down, so they write less (less billable hours), but what they right is what is fed to the AI to generate a verbose patent, so there the patent is more valuable. This is assuming the value of a patent is in its chances to (a) be granted (it is not missing any relevant background or caveat), (b) stop competitors (it is hard to make sure it is not infringed), and (c) be upheld if challenged (it is complete and correct). This is where society needs to step up and legislate so these new tools are not abused and the value to society takes precedence, so patents are concise and help make inventions known and understood.
- 2y ago
- cynicalsecurity 2y agoThe name sounds like illicit.
- vsnf 2y agoelicit transitive verb 1: to call forth or draw out (something, such as information or a response) her remarks elicited cheers 2: to draw forth or bring out (something latent or potential) hypnotism elicited his hidden fears https://www.merriam-webster.com/dictionary/elicit https://www.merriam-webster.com/dictionary/elicit
- deleted 2y ago[deleted]
- dpflan 2y agoIndeed, I get the idea behind the name, but there is certainly a dash of irony here being that it is an LLM, who knows if it was "ethically" trained...
- jamesbrady 2y agoActually, it's not an LLM! We do use LLMs, but the secret sauce is an approach we call Factored Cognition which we wrote about here: https://ought.org/research/factored-cognition https://ought.org/research/factored-cognition (Elicit the company and app was spun out from Ought the research lab). We do joke internally about the homophone (in fact, IIRC we did a little joke on our CEO by rebranding for his birthday in 2022) but I'm sorry to report that we're all careful, ethical, and well-behaved people :(
- dpflan 2y agoCool, thanks for more info, nice to see other approaches. What data is used for training?
- informal007 2y agoIt will be more easy to sucess in the basic education rather than AI Research(one of the most hard field in the world) when apply AI to education field. I think it's a little early to bring AI to research field which need enough accuracy and rigorism.
- ta988 2y agoTo me it reminds me of the advent of scholarly databases. The main effect I saw is that researchers started using exclusively those databases, sometimes publisher specific databases (so they were citing only from one publisher!) and were missing all the papers that were not indexed there. In particular a big chunk of the older literature that wasn't yet OCRed (it is better but still not fabulous). This led to so many "we discovered a new X" paper that the older people in the crowd in conferences were always debunking "that was known since at least the 60s". While those AI tools can clearly help with initial discovery around a subject, it worries me that it will reduce the search in other databases, or the digging into paper references. It is often enlightening to unravel references and go back in time to realize that all recent papers were basing their work on a false or misunderstood premise. Not talking about the cases where the citation was copied from another paper and either doesn't exist or had nothing to do with the subject. There was a super interesting article about the "mutations" of citations and how you could, by using similar tools to genetic alysis, generate an evolutionary tree of who copied on who and introduced slight errors that would get reproduced by the next one. edit: various typos
- nicklecompte 2y ago> A good rule of thumb is to assume that around 90% of the information you see in Elicit is accurate. While we do our best to increase accuracy without skyrocketing costs, it’s very important for you to check the work in Elicit closely. We try to make this easier for you by identifying all of the sources for information generated with language models. A 90% accuracy rate seems like the sweet spot between "an annoying waste of time" for honest researchers and "good enough to publish" for dishonest careerists. I don't like disparaging the technology experts who work on these things. But as a business matter, 1/10 answers being wrong just is not good enough for a whole lot of people.
- fragmede 2y agoif it takes 1 hour to get one answer by hand, but only 20 minutes for the machine, and 20 minutes to check the answer, the user still comes out ahead
- nicklecompte 2y agoThose numbers are arbitrary and fictional, and the more relevant made-up quantity would be the variance rather than the mean. It doesn't really matter if the "average user" saves time over 10,000 queries. I am much more concerned about the numerous edge cases, especially if those cases might be "edge fields" like animal cognition (see below). In my experience it takes quite a bit longer to falsify GPT-4's incorrect answers than it does to a Google search and get the right answer. It might take 30 seconds to check a correct answer (jump to the relevant paragraph and check), but 30 minutes to determine where an incorrect answer actually went wrong (you have to read the whole paper in close detail, and maybe even relevant citations). More specifically, it is somewhat quick to falsify something if it is directly contradicted by the text. It is much harder to falsify unsupported generalizations or summaries. As a specific example, I recently asked GPT for information on arithmetic abilities in amphibians. It made up a study - that was easy to check - but it also made up a bunch of results without citing specific studies. That was not easy to check[1]: each paragraph of text GPT generated needed to be cross-checked with Google Scholar to try and find a relevant paper. It turned out that everything GPT said, over 1000 words of output, was contradicted by actual published research. But I had to read three papers to figure that out. I would have been much better off with Google Scholar. But I am concerned that a large minority of cynical, lazy people will say "90% is good enough, I don't want to read all these papers and nobody's gonna check the citations anyway" and further drag down the reliability of published research. [1] This was a test of GPT. If I were actually using it for work, obviously I would have stopped at the fake citation.
- admissionsguy 2y agoI gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't, would have been all over the news), missed key papers, dug up obscure and irrelevant ones. Par for the course with LLMs. Edit: After looking at the examples on the front page "What are the benefits of taking l-theanine?" this seems geared for the general public, so maybe it wasn't the right test.
- sunir 2y agoI think of AI as a super keen arrogant intern. With GitHub copilot, an intern that constantly interrupts me. When that works for me, I am probably weak on the subject material myself. eg writing quirky love poems to my wife in different styles. For research tasks, because the AI is not deeply self-reflective, it can output inconsistent and incoherent results. What it does is present text that only *looks* as if it confidently knows what it is talking about. For domains where high rational quality doesn’t matter like love poetry, it is amazing. For other domains, be wary. If you can’t tell the difference between what is actually good and what merely looks good superficially you will be in trouble.
- errlogic 2y agoI generally agree with your take at first, but the following statements are funny to me: “ I am probably weak on the subject material myself. eg writing quirky love poems to my wife in different styles.” “ For domains where high rational quality doesn’t matter like love poetry, it is amazing.” So, self-described weak at love poetry, but confident that it is a domain that LLMs excel at. That is an interesting take. Perhaps the LLM is just as weak at liberal arts as it is hard science, but it is just more difficult to measure since you aren’t in the domain. Most poetry I’ve seen from LLMs has been pretty rote and boring although as you say, not a “rational quality” I suppose.
- croon 2y agoI was about the comment something like this, but you put it well. If you don't read a lot of poetry, what LLMs output look like poems, but almost always lack wit, a through-line, coherence and poignancy. It usually contains the individual parts, but never fitting as a whole.
- ZeroCool2u 2y agoI've been disappointed in some of these services before, because in the early days and I suspect now still as well they used RAG for obvious reasons. I feel like for research it's rare that RAG works really well. What I'm really excited about is a tool like Elicit using the new Google Gemini 1.5 Pro/Ultra models with the 2 Million token window sizes, filtering down the papers using traditional search and high quality meta-data, then critically, prompt/activation caching to make the tool economically viable. Maybe it won't work better, but I'm willing to bet it'll find those really specific ideas/needles in the haystack a lot more often than vanilla RAG will.
- inimino 2y agoIt's the economically viable part that I think is currently hard, but agreed this is the right approach. The basic problem is that scaling up understanding over a large dataset requires scaling the application of an LLM and tokens are expensive.
- lysecret 2y agoAnd the minutes of latency you currently get with this context lengths.
- ZeroCool2u 2y agoYeah, this is why I mentioned Gemini with the context caching. It's not out yet, but supposedly launching soon. You pay a lower rate for storing the system prompt or whatever you dump in before the user query, plus you don't have to wait the full minute or so for all your research to be ingested every time. https://ai.google.dev/gemini-api/docs/caching#get-started https://ai.google.dev/gemini-api/docs/caching#get-started
- jgalt212 2y ago> Trusted by Researchers At ... How did you get all those blue chip orgs give you permission to use their name? None of our blue chips clients allow us to do so.
- draxil 2y agoCurrent gen of "AI" companies aren't that interested in such trivial things!
- deleted 2y ago[deleted]
- jwr 2y agoAs someone who also tried (and stopped trying) to use customer logos, I also wonder whenever I see this pattern. So many startups do this, and yet I know how difficult it is to get official permission and how using a name without permission can lead to serious consequences. Do startups just roll with it and use the logos without permission?
- inerte 2y agoI do remember the day I learned my company was supposed to get permission, because the logo there implies official endorsement. I just didn’t know. I would start with Hanlon’s Razor, with a 10% chance of malice.
- beshrkayali 2y agoI’ve learned to discard those as being either totally fake and made up (the excuse would be something like we used a template to get started quickly and forgot to change the logos), or that someone (probably an intern) signed up to the service with @company email and they just splat the logo as an official endorsement.
- Euphorbium 2y agoIt seems like it should hallucinate less, as it directly quates, but nope, it still halucinates just as much and then gives a quote that directly contradicts its statement.
- jamesbrady 2y agoElician here! Accuracy and supportedness of the claims made in Elicit are two of the most central things we focus on—it's a shame it didn't work as well as we'd like in this case. I'd appreciate knowing more about the specifics so we can understand and improve
- Euphorbium 2y agohttps://analyticalsciencejournals.onlinelibrary.wiley.com/doi/10.1002/dta.2950 https://analyticalsciencejournals.onlinelibrary.wiley.com/do... elicit summarises this paper abstract: “Psilocybin was present at 0.47 wt% in the mycelium.” Actual quote from the abstract: “ No tryptamines were detected in the basidiospores, and only psilocin was present at 0.47 wt.% in the mycelium.” It does not differentiate between psilocin and psilocybin, those are two different molecules.
- okawei 2y agohttps://scisummary.com https://scisummary.com can get you pretty deep dives on individual papers. I've always found that these document search engines that are "AI-powered" just can't stand up to hand picking articles myself. Also, these are only open access papers, so anything behind a paywall will just be straight up missing.
- jamesbrady 2y agoElician here! Our main focus is a little different to SciSummary actually. We're focussed on understanding researchers broader workflows, and providing a research assistant (i.e. rather than a particular narrow tool for summarisation or search). The workflows we're most excited about at the moment are literature and systematic reviews: we think we can make these orders of magnitude faster and higher quality.
- koeng 2y agoTesting Elicit gave me quite a bit worse results than using PaperQA by futurehouse. While paperqa could understand a bit of the nuance of a scientific query, elicit did not. Too bad the internal paperqa system at scihouse isn't available for public use...
- darkteflon 2y agoTo any Elicians in the chat: I know from one of your previous blog posts that you guys use Vespa on the back end - would you be able to comment on your experience with it generally?