4 ms·
Sorry, OpenAI's take is correct here. If you're not convinced, here is how they prompted LLM: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cd
by raincole 2mo ago
Sorry, OpenAI's take is correct here. If you're not convinced, here is how they prompted LLM: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cdc_prompt.pdf https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98... [0]
A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
[0]: Not one of the proofs in the linked article, but from OpenAI too.
- yaqubroli 2mo agoThe human provides the intention and the ability to appreciate the output. Tools do “heavy lifting” all the time, but we still primarily credit the humans who use them precisely because they made the choice to use them. Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.
- esikich 2mo agoWhat gives the intention and ability to the human?
- oklahomasports 2mo agoAre you playing dumb? Using power tools to build furniture is very different than using an ai robot to carve a statue or whatever.
- zkmon 2mo agoWhen you use a crane to do the "heavy lifting" for construction work, do you give full credit to the cranes?
- mathisfun123 2mo agoI don't disagree with you but there's no need for exaggeration; ain't no high school student writing this: > In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient. which is infact a very important part of the prompt.
- don_esteban 2mo agothe fact that such things have to be explicitly in the prompt points to the fact that the underlying system is still far from where it needs to be (basically, lacks basic understanding what a proof is)
- famouswaffles 2mo agoThey don't have to be. At this point, we have multiple results from 3rd parties where the prompts are very basic. To name a few: - https://xcancel.com/DmitryRybin1/status/2079904005652893709 https://xcancel.com/DmitryRybin1/status/2079904005652893709 - https://archive.ph/2w4fi https://archive.ph/2w4fi (https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba9c https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...)
- skinner_ 2mo agoNo, that's not what this is. This is a warning to the LLM that coming back with partial results is not good enough. Take a grad student with a perfectly good understanding of what a proof is. Their supervisor gives them a major problem to work on. Almost always, the problem is too hard, the student comes back with partial results, and student and the supervisor iterate from there. Now imagine that they have an unusually cruel and unreasonable advisor who tells them, do not dare to talk to me until you've fully solved the problem. This paragraph is exactly that. It's there exactly because the underlying system is smart enough to know that real mathematicians do not work like that.
- ben_w 2mo ago> A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt. I think you're over-estimating what a smarter highschooler could write. A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither: Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity. nor: repeated-edge closed trails masquerading as cycles would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points. * For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom) https://en.wikipedia.org/wiki/A-level_(United_Kingdom)
- ipnon 2mo agoBut why can’t we prompt the LLM “just do math research”? This is what I don’t understand.
- raincole 2mo agoIf there aren't thousands of TPUs doing that [0] right now I'd be quite surprised. [0]: e.g. "go through wikipedia's unsolved math problem list and solve them".
- ascots 2mo ago100% agree. If the models are so capable that they're advancing math, it doesn't seem like a stretch to expect they should be able to determine with "doing math research" entails and the best way to use their capabilities towards that end. Why do we need to hand hold the models by telling them to do parallel research, keep threads independent, etc.
- melector 2mo agoBecause they're not aware of what user wants, and they need to know what expectation is, if it finds out it's a famous open problem it may think informing the user and not trying is best option as average user may not prefer it spending hours when success isn't guaranteed, by telling it to use it's available tools and not stop at partial progress, use subagents for various independent approaches it's allowing LLM to know what it should do and what counts as success. These things do great when goal is well defined.