4 ms·
Buried pretty deep in the article > “The raw output of ChatGPT’s proof was actually quite poor. So it required an expert to kind of sift through and actually u
by lqstuart 6mo ago
Buried pretty deep in the article
> “The raw output of ChatGPT’s proof was actually quite poor. So it required an expert to kind of sift through and actually understand what it was trying to say,” Lichtman says. But now he and Tao have shortened the proof so that it better distills the LLM’s key insight.
I guess “ChatGPT came up with a novel approach to a problem that later turned out not to be totally stupid and terrible for once” isn’t as catchy of a headline
- FrustratedMonky 6mo agoHow often are humans initial key insights, also sublimely distilled and beautiful. This is like comparing someone's first draft, with a final published paper.
- arcticfox 6mo agoThat should be buried, I agree 100% with their headline and structure over yours. For comparison, if the amateur did it by hand but the result was sloppy to read, would you prefer "Amateur solves an Erdos problem" or "Amateur came up with a novel approach to a problem that later turned out not to be totally stupid and terrible for once"?
- deleted 6mo ago[deleted]
- vovavili 6mo agoI wouldn't expect a hand-crafted proof by an amateur to be much different.
- stingraycharles 6mo agoWouldn’t the expectation for ChatGPT be that it presents a well refined report, rather than hand crafted proof / notes? Because from what I gather, they basically had to go through the equivalent of a pile of notes to find the crux.
- cccbbbaaa 6mo agoYeah, I also expect better than amateur-level redaction from a system that OpenAI marketing sells as a team of PhD in your pocket.
- shiandow 6mo agoDepends. I reckon a proof by an amateur would either be worthless because it demonstrates no understanding whatsoever or significantly better because they actually understand the proof. LLM produced texts are often in a weird area where the quality of the content and the quality of the writing have very little to do with one another.
- FartyMcFarter 6mo agoI don't think it's true that all amateurs have no understanding whatsoever. Amateurs have proven things before, and they've also wasted mathematicians time with wrong proofs.
- shiandow 6mo agoI'm not saying none of them have any understanding, I'm saying the ones who understand would write better proofs.
- themafia 6mo agoHow many hand-crafted amateur proofs do you read in a month? If the answer is close to zero then what are your expectations actually driven by?
- vovavili 6mo agoMy understanding of the world, human nature and human learning?
- geon 6mo agoDoesn’t this mean the expert solved the problem while trying to devine the tea leaves provided by chatgpt?
- elil17 6mo agoI understood this to mean that the ChatGPT output was technically correct, just hard to understand.
- SpicyLemonZest 6mo agoI haven't reviewed it myself, but when a mathematician calls a proof "quite poor" and experts have to "sift through" it, I would understand that to mean that it's technically incorrect. Errors like "This statement isn't correct, but it points towards a weaker statement that is, and the subsequent steps can be rebuilt on top of the weaker statement" are pretty common in output from both LLMs and math students.
- BeetleB 5mo agoNo - It likely means that the proof was meandering, and had lots of additional pointless steps.
- estimator7292 5mo agoGood/bad is orthogonal to correct/incorrect.
- culi 6mo agoDeepSeek also seems capable of solving it. In under 20 minutes https://chat.deepseek.com/share/nyuz0vvy2unfbb97fv https://chat.deepseek.com/share/nyuz0vvy2unfbb97fv I guess we should test across other LLMs too
- ngruhn 6mo agoDo you have any idea if this is correct?
- themafia 6mo agoThere should be zero expectation that the solution is "novel." It could not have produced any of it were it not in it's training data set. This is simply evidence that our search tools and academic publishing are completely broken and not at all evidence that a machine "thought up a novel solution." Humans constantly anthropomorphize their environment. To their detriment.
- pylua 6mo agoIf that’s the case, are you saying the proof was part of the training set? A lot of novelty is just gluing approaches together and reporting what sticks.
- themafia 6mo agoYes. A lot of brute force methods are the most inefficient means of solving a problem. Language Models are a sign that our current human information infrastructure and access methods are completely wrong.
- pylua 6mo agoArguably, since everything resolves to an axiom, isn’t the solution in any training set ? The order and combination is what makes it special. Is current human information access methods wrong, or do we just synthesize data in a way that is inefficient for this sort of problem solving ?
- themafia 5mo ago> since everything resolves to an axiom This isn't true. There are solutions that are beyond apparent reason and logic. This is what a "breakthrough" is. > The order and combination is what makes it special. Given an infinite amount of time a team of monkeys will produce Shakespeare. Is that "special?" Perhaps we should leave some room for _how_ those combinations happen and how efficient they are. > Is current human information access methods wrong They are wrong. The largest search company is also the largest advertiser. I'm surprised that anyone either fails to apprehend this or pretends not to.