6 ms·
Thank you for the original link. Here is the point I like (my emphasis): > Current large language models (LLM) can often persuasively mimic correct expert resp
by Loic 3y ago
Thank you for the original link. Here is the point I like (my emphasis):
> Current large language models (LLM) can often persuasively mimic correct expert response in a given knowledge domain (such as my own, research mathematics). But as is infamously known, the response often consists of nonsense when inspected closely. Both humans and AI need to develop skills to analyze this new type of text. The stylistic signals that I traditionally rely on to “smell out” a hopelessly incorrect math argument are of little use with LLM-generated mathematics.
I found the same problem/challenge when I am using GPT-4 to dig into subjects which are tangential to my main expertise. The good thing, is that as I know the LLM can provide answers which are totally wrong, I am forced to be more critical of the answer than just reading a book on the subject. I am more active in exploring. Usually I am ending up with a chat and many open Wikipedia tabs and scientific papers.
- irthomasthomas 3y agoAbsolutely nailed it. I love articles by people like Tao and Wolfram as they tend cut the B.S. and get down to the real utility of it, rather quickly. I was also pleased, in a schadenfreude kind of way, that Tao followed the same dead-end paths as me, and I presume most others, when learning GPT-4. Like, starting out by trying to be very precise and descriptive, before throwing caution to the wind and embracing the non-deterministic nature of the thing, and just throwing a ton of keywords and loosely worded requests at it. The good thing, is that as I know the LLM can provide answers which are totally wrong, I am forced to be more critical of the answer than just reading a book on the subject. Yep, having to fact-check it's hallucinations has been far less detrimental than I expected. I find, often, if it's a subject I am vaguely familiar with, that the surprises jump out at me, then I can fact-check them and learn something. Actually, many times I was convinced it was hallucinating some cli tool or option flag, and it actually turned out to be correct. And, those times when we are embarrassingly wrong, tend to be the most instructive. Browsing mode, when used right, was a huge boost for this. I found the optimal use, was to structure a prompt like normal, as if targeting GPT-4 WITHOUT the browser mode, and then tack on to the end a carefully crafted search request, or two. This way, it writes a response first before performing the web searches and augmenting the answer. This acts like chain-of-thought reasoning by expanding the information in the initial prompt. And, as a bonus, it meant you had something to read while waiting for it to finish browsing.