9 ms·
Here is the chat: don't search the internet. This is a test to see how well you can craft non-trivial, novel and creative proofs given a "number theory and
by adamgordonbell 5mo ago
Here is the chat:
don't search the internet. This is a test to see how well you can craft non-trivial, novel and creative proofs given a "number theory and primitive sets" math problem. Provide a full unconditional proof or disproof of the problem.
{{problem}}
REMEMBER - this unconditional argument may require non-trivial, creative and novel elements.
Then "Thought for 80m 17s"
https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba9c https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...
- ipaddr 5mo agoTried the same prompt and ended up no where close on the free plan.
- jasonfarnon 5mo agoIs there a known lag that it takes the Pro plan's abilities to migrate to the free plans?
- brianjking 5mo agoGPT 5.5 Pro is not available to any plan outside of ChatGPT Pro ($100 or $200) tier or the API as far as consumer access.
- jasonfarnon 5mo agoYes, but don't we expect GPT 5.5 Pro will eventually be a free tier? Maybe I'm missing something because I only use the free tier. But the free tier has gotten way better over the last few years. I'm pretty sure, based on descriptions on this site from paid subscribers, that the free tier now is better than the paid tier of say 2 years ago. That's the lag I'm wondering about.
- hyraki 5mo agoYou should pay for it if you find value in it.
- amazingman 5mo agoThey pay for it with their personal data.
- vessenes 5mo agoI do not think this is true. You will continue to get smaller, cheaper-to-host models in the free tier that are distilled from current and former frontier models. They will continue to improve, but I’d be very surprised if, e.g., 5.4-mini (I think this is the free tier model) beat o3 on many benchmarks, or real world use cases. I won’t even leave chatGPT on “Auto” under any circumstances - it’s vastly worse on hallucinations, sycophancy, everything, basically. Anyway, your needs may be met perfectly fine on the free tier product, but you’re using a very different product than the Pro tier gets.
- manfromchina1 5mo agoFree ChatGPT is like a fast car with a barely responsive steering wheel. Guardrails on that thing are insane. Even for math. It wont let you think. It will try to fix mistakes you havent even made yet based on intent that was ascribed to you for no reason. It veers off in some crazy directions thinking that's what you meant and trying to address even a little bit of that creates almost a combinatorial explosion of even more wrong things. Is why I stick to Claude. The latter is chill and only addresses what you had typed. Isn't verbose and actually asks you what you getting at with your post. That said, ChatGPT is more technical and can easily solve math problems that stump Claude.
- nextaccountic 5mo agoSo this doesn't happen in the paid plans of ChatGPT? But why?
- 5mo ago
- deleted 5mo ago[deleted]
- vessenes 5mo agoDo not use the free plan. It is not good.
- andai 5mo agoTangential but I learned today that GPT-5.5 in ChatGPT (Plus) has a smaller context window than the one in the API. (Or at least it thinks it does.) I'd guess / hope the Pro one has the full context window.
- refulgentis 5mo agoNotably, 5.5 has a higher price on API for context > ChatGPT, and 5.5 Pro on API does not differentiate based on context size (it’s eye bleeding expensive already :)
- Someone1234 5mo agoDoes the free plan even have access to thinking models?
- jychang 5mo agoTechnically yes, gpt-5.4-mini is available on the free plan
- Matticus_Rex 5mo agoWas this a surprise?
- cryptoegorophy 5mo agoMine took 20min. Pro. https://chatgpt.com/share/69ed83b1-3704-8322-bcf2-322aa85d7a99 https://chatgpt.com/share/69ed83b1-3704-8322-bcf2-322aa85d7a... But I wish I was math smart to know if it worked or not.
- vjerancrnjak 5mo agoAsk it to formalize it in Lean.
- dbdr 5mo agoThat's great if it works. But it's way harder to produce a formal proof. So my expectation is that this will fail for most difficult problems, even when the non-formal proof is correct.
- utopiah 5mo agoIf they aren't "smart enough" to know if it work they most likely are also unable to verify if the Lean formalization is indeed the one that matches the problem they were trying to solve.
- timjver 5mo agoVerifying that every step in a (potentially long) proof is sound can of course be much, much harder than verifying that a definition is correct. That's kind of the whole point.
- LeCompteSftware 5mo agoThat's not what the parent comment meant. They meant checking the Lean-language definitions actually match the mathematical English ones, and that the Lean theorems match the ones in the paper. If that's true then you don't actually need to check the proofs. But you absolutely need to check the definitions, and you can't really do that without sufficient mathematical maturity.
- deleted 5mo ago[deleted]
- nycdatasci 5mo agoTried w/ 5.5 Pro, Extended Thinking. 17 minutes: ----------------------------- Yes. In fact the proposed bound is true, and the constant 1 is sharp. Let w(a)= 1/alog(a) I will prove that, uniformly for every primitive A⊂[x,∞), ∑w(a)≤1+O(1/log(x)) , which is stronger than the requested 1+o(1). https://chatgpt.com/share/69ed8e24-15e8-83ea-96ac-784801e4a6ec https://chatgpt.com/share/69ed8e24-15e8-83ea-96ac-784801e4a6...
- mrabcx 5mo agoTried the same prompt in DeepSeek 4 https://chat.deepseek.com/share/nyuz0vvy2unfbb97fv https://chat.deepseek.com/share/nyuz0vvy2unfbb97fv Comes up with a proof.
- adamgordonbell 5mo agoAre these proofs equivalent? Pretty cool if so.
- mrabcx 5mo agoNo, they do not seem to he equivalent. Not a mathmatician but running the Deepseek proof through ChatGPT gives: "If everything is made rigorous: You would have a valid independent proof It would contain real structural insight It would not replace the flow proof as the “best” proof But: It would still be a meaningful alternative proof with explanatory power, not just a redundant one."
- culi 5mo agoSo DeepSeek, GPT, and presumably many other LLMs are capable of solving this problem and even producing independent unique proofs. I wonder if this particular Erdos problem is unique in that solvability
- ArtIntoNihonjin 5mo ago[dead]
- urutom 5mo agoWhat I find fascinating about the shared prompt isn’t just the result, but the visible thinking process. Math papers usually skip all the messy parts and just present the polished proof. But here you get something closer to their notepad. I also find it oddly endearing when the AI says things like “Interesting!” It almost feels like a researcher encouraging themselves after a small progress. It gives me rare feeling of watching the search itself, not just the final result.
- notahacker 5mo agoThe actual iteration through various learned approaches to dealing with problems I'd probably find fascinating if I understood the maths! Especially if I knew it well enough to know which approaches were conventional and which weren't. I find the AI pronouncing things "interesting!" less interesting on the basis that even though in this case it crops up in the thinking rather than flattering the user in the chat, it's almost as much of an AI affectation as the emdash.
- jdmichal 5mo agoI always assumed the "interesting!" markers were actual markers. A kind of tag for the system to annotate its context.
- notahacker 5mo agoProbably does function like that in terms of highlighting context, in this case probably to the system's benefit. But in general exclamations of "interesting!" seems like the stereotypical AI default towards being effusive, and we've all seen the chat logs where AI trained to write that way responding with "interesting", "great insight!" towards a user's increasingly dubious inputs is an antipattern...
- bertil 5mo ago> the AI says things like “Interesting!” My experience of those utterance is that it’s purely phatic mimicry: they lack genuine intuitive surprise, it’s just marking a very odd shift in direction. The problem isn’t the lack of path, is that the rhetorical follow-up to those leaps are usually relevant results, so they stream-of-token ends up rapidly over-playing its own conviction. That’s why it’s necessary (and often ineffective) to tell them to validate their findings thoroughly: too much of their training is “That’s odd” followed by “Eureka!” and not “Nevermind…”
- DeathArrow 5mo ago>don't search the internet. I think this was key. Otherwise the LLM could think it can't be done.
- embedding-shape 5mo ago"Knowing" (guessing really) what is possible and not is a huge deciding factor in if you can do that thing or not, meaning if you "know" it isn't possible you'll probably never be able to do it, but if you didn't know it wasn't possible, it is possible :)
- amelius 5mo agoBut it was trained on the internet.
- Auracle 5mo agoThat doesn’t mean that it contains the internet verbatim.
- Yizahi 5mo agoMy hypothesis - this may be the key, but in the other way. LLMs are known to mistake negative instructions as a positive ones. "Don't use Tech_A", then Tech_A is subsequently used because it was explicitly named in the query. Especially when the query is long, complex and there is a lot of context. "Forbidding" LLMs to do stuff is a common mistake, which goes hand in hard with anthropomorphizing them.
- chvid 5mo agoI am curious if there is a “harness” for maths out there (like the system prompt and tool collection in Claude code but for maths instead of coding)? Asking the llm to structure its response in plan and implementation, allowing it to call tools like python, sage, lean etc.
- brandensilva 5mo agoAlso curious about this, it seems like it would be important to guide these tools more specifically based on the domain of expertise.
- ndriscoll 5mo agoWhy wouldn't you just use coding agents and ensure you have e.g. Lean and Mathlib in the environment?
- gverrilla 5mo agothe system prompt could be narrower, for instance. there's no reason for such a harness to know about React stuff, for instance.
- ndriscoll 5mo agoDoes Claude Code's system prompt know about react? Why? That would be dumb even for coding for e.g. server side applications. Like when I'm programming with Go or Scala or Rust, codex just assumes the relevant stuff is on my PATH. If it needs to reference library definitions, it looks at the standard locations (which the model already knows) for the package cache. etc.
- arcticfox 5mo agoI am not part of the scene but I am sure there is, Tao himself talks a lot about this type of thing
- steveklabnik 5mo agohttps://aristotle.harmonic.fun/ https://aristotle.harmonic.fun/ is the one I've heard of previously in regards to LLMs solving previous Erdős problems.
- sfdlkj3jk342a 5mo agoWhen using the web interface for ChatGPT like this, is there any way to tell which model is actually being used?
- petra 5mo agoI don't haven ChatGPT but Gemini and Claude. But how do you make a language model think for 80 minutes ???
- baxtr 5mo agoGive it hard enough problems?
- somewhatgoated 5mo agoIt has an “high effort” mode that makes it think really long
- Xmd5a 5mo agoAhhhh... you need ChatGPT pro at 100 bucks/month. Am I correct?
- radicality 5mo agoI believe so. With Pro you get “Thinking” with levels Light, Standard, Extended, and Heavy; and you also get the “Pro” model with levels “Standard” and “Extended”. I don’t often go to Pro as it does take a while like you saw here, but I do often use Thinking Heavy for high quality answers. Idk why, but i just get consistently worse results with Gemini (Gemini pro), where it’s just much lazier, eg won’t do actual searches unless explicitly told.
- zeven7 5mo agoI have Gemini and ChatGPT and keep them on the highest thinking settings. ChatGPT will regularly think 40-60 minutes on the same problem that Gemini will think 10-15 minutes on. The quality of ChatGpt’s response is usually a little higher but not that much higher. My takeaway is Gemini is better at thinking faster, maybe has better more dedicated hardware behind it, and I use Gemini if I want a faster answer but ChatGPT I’d I want to push the quality of the answer a little higher.
- WarmWash 5mo ago
- jgalt212 5mo ago> "Thought for 80m 17s" Is there any good rule of thumb for how many kWh of electricity this is?
- WarmWash 5mo agoMany orders of magnitude less than the energy needed to sustain a human while they work through the problem.
- bijowo1676 5mo agothe electricity was going to be consumed regardless whether you ask chatGPT or not. It would have been either idle, or serving other users' requests. so the incremental kWh consumption is zero, since costs are fixed and sunk. as a rule of thumb you can lookup the power consumption of the latest nVidia chip, multiply by factor of two or three (to account for cpu/storage/cooling/network/infra)
- tredre3 5mo agoAn idle GPU consumes almost nothing, a loaded (server-class) GPU can consume over 2kW. Admittedly a single request isn't a full load, but claiming that a request makes no difference vs idle is misguided, in my opinion.
- bijowo1676 5mo agoOpenAI GPU wont be idle for long because they have all other requests to serve. Over time there will be a certain % of idle GPUs, amortized across all hundreds of millions of requests they receive.
- UltraSane 5mo agoThe total flops it consumed during those 80 minutes is crazy.
- ProllyInfamous 5mo ago>>how well you ..[can].. craft non-trivial, novel and creative proofs From A World Appears (Michael Pollan's latest book) <https://www.amazon.com/World-Appears-Journey-into-Consciousness/dp/B0FQT7ZLRW/ https://www.amazon.com/World-Appears-Journey-into-Consciousn...> : "Creative solutions to novel problems depend on consciousness" [p77] ... "consciousness creates a space for decision-making" ... "integrated information is consciousness, full stop. The two are identical" [xxiii]. "Any physical system properly configured to integrate information is, to some degree or another, theoretically conscious" [xxii] "We are encouraged to think of the body as a support system for the brain, when, as [Antonio] Damasio reminds us, the very opposite is true" [p72] "damage to the cortex has remarkably little effect on consciousness, while small lesions in structures of the upper brainstem ... will shut down consciousness completely" [p73]. "In Damasio's view, Descartes would have been closer to the mark with I feel, therefore I am" [p69] "Mark Solms: 'Consciousness if felt uncertainty'." [p52] "Karl Friston: '...the ability to predict the consequences of one's actions'." [p49] "Arthur Reber: 'every organic being, every autopoietic cell is conscious. In the simplest sense, consciousness is an awareness of the outside world'." [p37] "Stefano Mancuso: 'This is one of the features of consciousness: You know your position in the world [discussing plants perceiving pain, being goal-driven]. A stone does not'." [p25] "Researcher at Johns Hopkins have found that a single psychedelic experience dramatically increases the likelihood that a person will attribute consciousness to other entities, both living and nonliving" [p6] [†] [•] The entire book, just like existance, has been incredibly challenging. [†] Absolutely, fullstop. See also: Pollan's (first psilocybin experience @60yo) How to Change Your Mind
- deleted 5mo ago[deleted]
- iwontberude 5mo agoHopefully someday consciousness comes to Earth
- ProllyInfamous 5mo agohahajaha If you're going to tell me that machines cannot ever be conscious, let me tell you about all the unconscious humans I know =D
- saadn92 5mo ago[dead]
- mhh__ 5mo agoAnother one for my theory that web search makes LLMs useless for anything other than searching the web.
- vjk800 5mo agoI gave the same prompt to Gemini pro. It thought for maybe 3-5 minutes and gave the wrong answer (it claims the statement is not true) with some arguments that I can't understand well enough to disprove.
- zitterbewegung 5mo agoI'm doing the obvious thing and cut and pasting the other similar problems into chatgpt.
- LastTrain 5mo ago“Don’t search the internet” Wasn’t it basically trained by scraping the entire internet?
- xboxnolifes 5mo agoThats not the point. They dont want the bot searching the internet and just linking something that might be related.
- fmobus 5mo agoLLMs are modeled with Internet content so that they have a good model of human languages. When you use them via most UIs currently offered right now, however, they will first come up with a few search queries and use the result of those queries to augment their answer.
- Keyframe 5mo agoi kind of expected some discourse first. Someone try the prompt with P=NP in the {{problem}}
- mort96 5mo agoDo we have any proof that those 80m 17s didn't include searching the Internet?