12 ms·
A recent experience with ChatGPT 5.5 Pro
https://twitter.com/wtgowers/status/2052830948685676605 https://twitter.com/wtgowers/status/2052830948685676605
https://xcancel.com/wtgowers/status/2052830948685676605 https://xcancel.com/wtgowers/status/2052830948685676605
- CharlesLau 5mo agoIs the assessment system of undergraduate mathematics education no longer effective?
- margalabargala 5mo agoUndergraduate? No. We've had calculators able to solve undergraduate problems for decades. AI doesn't change the need to understand how calculus works any more than calculators did. The foundations remain valuable. Graduate? Yes.
- whatever120 5mo agoHow should graduate school be changed then? Specifically for mathematics
- dyauspitr 5mo ago90% of the final grade are in room examinations with proctors, maybe two sets of exams of midterms and finals that the vast majority of the final grade comes from. This is already how most of East and South Asia does it anyways and it’s probably the best. For publications and theses, as long as the final results hold and can be replicated and validated, I don’t see why we shouldn’t allow the wholesale use of LLMs
- zozbot234 5mo ago> 90% of the final grade are in room examinations with proctors, maybe two sets of exams of midterms and finals that the vast majority of the final grade comes from. This is really just a glorified undergraduate education, the real point of graduate school is to learn to do real-world relevant research. For the latter, I think LLM use will be accepted but there will be a heavy expectation on the author of making the result very easily digestable for human mathematicians and linking it thoroughly with the existing literature - something that LLMs are very much not successful at, but a student might be able to do quite well with a mixture of expert guidance and personal effort.
- margalabargala 5mo agoHell if I know. It's easy to see problems. Solutions are way harder. I'm not a math professor.
- dyauspitr 5mo agoI don’t think it’s just mathematics. We don’t hear enough about this, but if I think back to my undergraduate years, which were less than 10 years ago, every homework assignment and every take-home exam I had would be trivial for LLMs to solve at this point I wonder what is actually happening on the ground.
- crocdundae 5mo agoWell... here's something from "boots on the ground": I teach a bachelor's degree where programming is a smallish facet of a curriculum. My course is the last of a series of 3 courses which progressively introduce more concepts and try make practical implementations more feasible. I've been able to grade the course purely based on returns to take-home exercises, some of which are complex, some trivial. When ChatGPT (& Co.) came along I was still able to do that but with a major added workload to me (suddenly everyone started producing mountains of code, often nonsensical, but I still had to read it all). I always requested targeted, atomic changes to code (vs. rewrites) which served me well up to a point (I was still able to grade fairly). I requested them originally to avoid "github copies", but that worked kind of OK with ChatGPT too. However, when ClaudeCode came along it was obvious to me I'm loosing the battle. It does not particularly matter to me whether students use AI or not as long as the rows they add and alter in the assignments make sense, but the "last nail to the coffin" problem now with ClaudeCode is that in the latest batch (this spring) it is clear some students "pay themselves" a good grade (i.e. they pay for ClaudeCode, thus bypassing the need to actually learn). I cannot make assignments that are both complex enough to cause ClaudeCode tripping on something and still humane for those who do not use AI or only use free chatbot options. Essentially ClaudeCode plays havoc with the whole grading process: students not using it (whether they try to write code fully manually or ChatGPT assisted) are left with far less points that students who just push all the code I give to ClaudeCode and "let it rip" for some 15 minutes. This really irks me. So, my solution? Still working on it and hoping to find one! For sure no more points from most take-home assignments: lowest grades still achievable through them (the trivial ones), but that's it, the rest it preparation for an exam. Practically this already means anyone with ChatGPT is going to pass, no doubt about it... As for the higher grades, for autumn I'm desperately now figuring out how to even make a meaningful paper based exam for my course. I've myself completed a master's degree writing C language on paper with a pencil. I sure did not want to start doing that to others, but here we are. Besides, back in my youth the only "library" was pretty much ANSI-parts-of-C! I'm not sure what kind of a 2 inch thick stack of papers I'd have to give my students into the exam these days as reference material. One horrible aspect is that students are now far more dependent on compiler errors to spot pretty much anything and everything... I worry the first paper exam from me will be a total horror story to us all. In any case, interesting times.
- bustermellotron 5mo agoI saw Tim Gowers give a talk at the AMS-MAA joint meeting in Seattle about ten years ago where he predicted that in 100 years humans would no longer be doing research mathematics. I wonder if he’s adjusted his timeline. At the time I thought the key missing tool was a natural language search that acted like mathoverflow, where you could explain your problem or ideas as you understood them and get references to relevant literature (possibly outside your experience or vocabulary).
- 34qJhah 5mo agoAnd Teichmüller thought that Germany would win WW2 and volunteered for the Eastern Front. Being a gifted mathematician does not make you right. In fact, mathematicians have a lot of bizarre theories.
- iTokio 5mo agoOn complex problems with lengthy proofs, the first step that I would have done is to ask 5.5 pro in a new, unrelated, session, to be very critical, to try to find flaws in the arguments. And certainly not to send it to a fellow colleague to ask its opinion first. LLMs are certainly becoming capable to code, find vulnerabilities, solve mathematical problems, but we need to avoid putting their works in production, or in front of other humans, without assessing it by any possible mean. Otherwise tech leads, maintainers, experts get overwhelmed and this is how the « AI slop » fatigue begins. To be clear I’m talking about this step: > That preprint would have been hard for me to read, as that would have meant carefully reading Rajagopal’s paper first, but I sent it to Nathanson, who forwarded it to Rajagopal, who said he thought it looked correct.
- NitpickLawyer 5mo ago> but we need to avoid putting their works in production, or in front of other humans, without assessing it by any possible mean. I think this is good advice in general, maybe with an emphasis on public vs. private, friendly contact. Having 0 thought AI slop thrown at you out of the blue is rude. "could have been a prompt" indeed. But having a friend/colleague ask for a quick glance at something they know you handle well is another story for me. If I've worked on a subject for a few years, and know the particulars in and out, I'd have no trouble skimming something that a friend or a colleague sent me. I am sparing those 5-10 minutes for the friend, not for what they sent. And for an expert in a particular domain, often 5 minutes is all it takes for a "lgtm" or "lol no".
- pmontra 5mo agoIt's a very long post with a mix of technical (math) and philosophical sections. Here are the most striking points to reflect upon IMHO. > It seems to me that training beginning PhD students to do research [...] has just got harder, since one obvious way to help somebody get started is to give them a problem that looks as though it might be a relatively gentle one. If LLMs are at the point where they can solve “gentle problems”, then that is no longer an option. The lower bound for contributing to mathematics will now be to prove something that LLMs can’t prove, rather than simply to prove something that nobody has proved up to now and that at least somebody finds interesting. Training must start from the basics though. Of course everybody's training in math starts with summing small integers, which calculators have been doing without any mistake since a long time. The point is perhaps confirmed by another comment further down in the post > by solving hard problems you get an insight into the problem-solving process itself, at least in your area of expertise, in a way that you simply don’t if all you do is read other people’s solutions. One consequence of this is that people who have themselves solved difficult problems are likely to be significantly better at using solving problems with the help of AI, just as very good coders are better at vibe coding than not such good coders People pay coders to build stuff that they will use to make money and I can happily use an AI to deliver faster and keep being hired. I'm not sure if there is a similar point with math. Again from the post > suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would.
- kerabatsos 5mo agoBut perhaps we should regard it as a major achievement.
- lmpdev 5mo agoI mean in the same way getting Wolfram Alpha to solve a really hard/ugly differential equation I suppose
- 5mo ago
- few 5mo ago>So if your aim in doing mathematics is to achieve some kind of immortality, so to speak, then you should understand that that won’t necessarily be possible for much longer — not just for you, but for anybody. This made me a little sad
- bananaflag 5mo agoNow repeat that for every sort of human achievement
- bel8 5mo agoMachines are comming even after table tennis :( https://www.youtube.com/watch?v=VVEzgYxDdrc https://www.youtube.com/watch?v=VVEzgYxDdrc
- pmontra 5mo agoSports are safe. Machines came after runners (motogp, formula 1) and yet we cheer the winners of the 100 m at the Olympics Games. Fully autonomous bikes and cars won't change that. AIs destroy chess players. We still cheer the world champion. We care about sports with humans.
- fragmede 5mo agoRobot MotoGP would be amazing to see just how far the limits could be pushed without risking the life of a human though. Or even full size remote control.
- Ekaros 5mo agoSadly I don't think there is any safe tracks for proper autonomous car racing without limits... Still would be interesting to see what is the absolute best you could do if rules include only say minimum number of wheels and maximum dimensions for vehicles.
- verisimi 5mo ago[flagged]
- jdw64 5mo ago[dead]
- bananaflag 5mo ago> it sounds like there were already precedents or existing pieces of knowledge, but humans had not thought to connect them A lot of math research is like that. And, like the blog post suggests, problems one gives PhD students are 95% like that.
- jdw64 5mo agoMaybe I am still fortunate to have become a programmer. Most of what I do is just assemble things that other people have already built.
- themafia 5mo ago> there were already precedents or existing pieces of knowledge, but humans had not thought to connect them We used to call that "low hanging fruit."
- tanepiper 5mo agoBasically medical science too. My wife was able to diagnose her own anemia that the doctors kept missing, and has since been able to have iron infusions. The human doctors kept ignoring the signals, kept putting it down to 'diet' and 'exercise' (even though she does plenty of both)
- agiipullor 5mo agono blood tests were done?
- NotOscarWilde 5mo agoAs a TCS assistant professor from Eastern Europe, I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models. Paying for Pro from any of my current academic budgets is completely ouf of the field of reality here -- all budgets tend to have restricted uses and software payments fit into very few categories. Effectively, I'd have to ask for a brand new grant and hope the grant rules allow for large software payments and I won't encounter an anti-AI reviewer; such a thing would take one year at least. As a nail to the coffin, I was "denied" all Claude Opus recently as part of Microsoft's clampdown on individual (and academic) use of Copilot. (Chagpt 5.5 Plus does not seem sufficient for any deeper investigations into new research topics, I've tried.) Apologies for the rant.
- dyauspitr 5mo ago[flagged]
- bananaflag 5mo agoFor a TCS assistant professor in Eastern Europe, $200/month would be 20% of their salary. And the situation is better, ten years ago it would have been 80%.
- xanrah 5mo agoLots of people in the west can’t afford 200 a month. How rich are you?
- MinimalAction 5mo agoAs a graduate student, this piece made me sad. I always believed that my work speaks for itself and transcends beyond my limited time on this cosmic experience. This notion of immortality was just a small intangible bonus I hoped for when I jumped into grad school. AI is making me feel less worthy.
- whatever120 5mo agoYou are worthy. You will hone your skills in grad school and be able to command these AIs better than somebody who hasn’t struggled with hard problems for a long time.
- jlarcombe 5mo agoA depressing thought that all that work is just so you can "command AIs better"
- folderquestion 5mo agoIt could happen than the AI, in a near future, is not something external but just a part of your brain, so you retain the glory.
- jlarcombe 5mo agoHah this is getting worse and worse
- dag100 5mo agoWhy stop there? Why not let AI take over all functions, Whispering Earring (https://gwern.net/doc/fiction/science-fiction/2012-10-03-yvain-thewhisperingearring.html https://gwern.net/doc/fiction/science-fiction/2012-10-03-yva... for anyone who hasn't read it) style?
- alexashka 5mo agoAll that work to kick a ball into a net. Nobody looks at this species and goes hm, rational and reasonable :)
- __rito__ 5mo ago> So maybe there should be a different repository where AI-produced results can live. Does the author know about CAISc 2026 [0]? [0]: https://caisc2026.github.io https://caisc2026.github.io
- shevy-java 5mo ago[flagged]
- Smaug123 5mo agoTim Gowers is a Fields medallist who has supervised 13 students (https://www.genealogy.math.ndsu.nodak.edu/id.php?id=67729 https://www.genealogy.math.ndsu.nodak.edu/id.php?id=67729). He's totally capable of gauging what a piece of PhD-level research is.
- incrediblylarge 5mo agoA month ago my PhD supervisor told me it rips on proofs but he also said it's useless when formalising arguments in Lean - is this still the case?
- vjerancrnjak 5mo agoNope. Codex formalizes much better than any tool with exception of Aristotle from Harmonic. https://github.com/vjeranc/fixed-rtrt https://github.com/vjeranc/fixed-rtrt M3 module was formalized fully purely from experimental data and from a nudge by earlier versions of codex in 15-30 minutes in a simple write/compile/fix-first-error loop. I was a bit surprised how fast it picked up the pattern but given there was a paper from '70s it became clear why later.
- mxwsn 5mo ago> Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would. This is a cultural choice. It makes sense that in the mathematics culture we currently have, this is alien. But already, other fields, and many individuals, would disagree and say that the human did have a major achievement here. As long as human-AI collaborations are producing the best results, there is meaningful contribution by the humans, and people that are deeper experts and skilled LLM whisperers should be able to make outsized contributions. The real shoe drops when pure AI beats humans and human-AI collaboration.
- pmontra 5mo agoI replied to a comment about AI in sports and I build on that. We praise car drivers despite most of the performance in their sport comes from the car. The driver makes the difference when two cars are close in performance. Brilliances or mistakes. Horse riders too. In the case of math, the human can lead the LLM on the right track, point it to a problem or to another one. So it deserves some praise. Then the team that built the car, cared about the horse, built the AI might deserve even more praise but we tend to care more about the single most visible human.
- dmbche 5mo agoCould you win an F1 race with the latest winning car against F1 drivers?
- audunw 5mo agoI’m not sure what your point is. I could certainly not, and I could certainly not write a breakthrough paper in mathematics even with the most advanced AI. I wouldn’t even know what to ask of it. Perhaps I could set up an elaborate master agent to consider all possible new problems in mathematics and ask sub agents to work on the most promising ones. But then I could probably also program a self driving car system which could win an F1 race as well.
- adammdaw 5mo agoThis is certainly interesting, though I would say that based on my understanding of how the current models work combinatorial problems would be an area where they could be particularly successful. They are pretty good at combinatorial creativity - its the exploratory and transformational aspects that are still pretty tricky, and I expect would come to bear in other areas of mathematics.
- hodgehog11 5mo agoIndeed, analysis is a bit more loose in its arguments, and so I've found LLMs tend to make more mistakes there.
- nxobject 5mo agoI wonder as well whether large-but-finite contexts can handle algebraic questions that require traversing up and down levels of abstraction, at least not without "thrashing".
- momojo 5mo agoSorry, I'm reposting a comment I made yesterday that seems fitting: > This reminds me of Antirez's "Don't fall into the anti-AI hype". In a sentence: These foundation models are really good at optimizing these extremely high level, extremely well defined problem spaces (ie multiply matrices faster). In Antirez's case, it's "make Redis faster".
- dabinat 5mo agoI feel like this experiment was successful because those prompting the AI were knowledgeable enough to ask the right questions and verify the output was correct. This shows that there is still a place for expertise, even if the LLM does the actual research.
- colechristensen 5mo agoI feel my input to LLMs is most valuable in the initial idea, big picture design tweaks, and the vast majority of my usefulness is negative feedback. This looks wrong, you've gotten off track, you're cheating with workarounds, you're falling into a rabbithole, etc.
- slopinthebag 5mo agoAI generated article btw. Maybe if you find AI to be doing stuff you find impressive, the stuff you were doing wasn't that impressive? Worth ruminating on your priors at least.
- reasonableklout 5mo agoWhat makes you think either the tweet or blog post are AI generated?
- hodgehog11 5mo agoThis is beyond ridiculous to say considering whose blog this is. For those that don't know, this is Timothy Gowers. He is one of the most accomplished mathematicians in the world. Like Terence Tao, he is considered one of the world leaders in mathematics and tends to have good judgement in where the field is going. Even without that knowledge, no, this article is certainly not AI generated. It has none of the tells.
- ziotom78 5mo agoI am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am discussing. For instance, in 3D Clifford algebras it repeatedly confuses exponential of bivectors and of pseudoscalars. Good to know that ChatGPT 5.5 Pro can produce a publishable paper, but from what I have seen so far with Gemini, it seems to me that it is better to consider LLMs as very efficient students who can read papers and books in no time but still need a lot of mentoring.
- recursivecaveat 5mo agoThis is close to my experience with code. LLMs can pick out small mistakes from giant code changes with surprising accuracy, or slowly narrow down a weird. On the other hand I've seen them bravely shoulder on under completely incorrect conceptual models of what they're working with and churn around in circles consequently, spin up giant piles of slop to re-implement something they decided was necessary, but didn't bother to search for, or outright dismiss important error signals as just 'transient failures'. Unlimited stamina, low wisdom.
- tags2k 5mo agoI'm no physics professor but this aligns with the way I use the tools in my "senior engineer" space. I bring the fundamentals to sanity-check the trigger-happy agent and try to imbue other humans with those fundamentals so they can move towards doing the same. It feels like the only way this whole thing will work (besides eventually moving to local models that do less but companies can afford).
- cyanydeez 5mo agoI've been watching the automation of things like flight control systems for the past decade, and the evolution of the fallback to a real pilot in the event of a emergency is what's most concerning about where LLMs are being embedded. Right now, we have a lot of smart people who have trained for decades to understand where these things go wrong and how to nudge them back, but the pool of people are going to slowly be replaced by less knowledgeable. At some point, a rubicon will be crossed where these systems can't fallback to a human operator and will fail spectacularly.
- globular-toast 5mo agoI wish people would stop generating stuff they don't understand only to forward it to someone who does. Something about that really rubs me the wrong way.
- hodgehog11 5mo agoMay I remind you that this is Timothy Gowers. He says he doesn't understand, but he most certainly has far greater capacity than most to detect complete junk from a maybe plausible argument. His colleague is even better able to judge this, hence why he sent it to him. Also if he did send me complete junk, I would still parse it for multiple days to see what is there.
- globular-toast 5mo agoYeah, it doesn't make a difference for me. It's the generation part. Gowers should have sent his prompts to the colleagues, not the generated paper. That's all. I feel like it's creating obligations for others to help with the remaining 20% which always takes the most time, while you get to have all the fun of doing the first 80%. I'm not criticising Gowers directly in this instance because he's exploring the possibilities, my disdain is towards the more general pattern I see emerging where people just send each other LLM outputs.
- hodgehog11 5mo agoNo offense intended, but this sounds like you're projecting impressions about LLM output for programming. Your words make a lot of sense to me in that context, but not as much here. I can speak with a reasonable amount of experience here. I absolutely guarantee you that what Gowers sent over was the vast majority of the work involved for a proof. It's also an interesting exercise in general, hence the blog post. Parsing a proof like this is _much_ easier than creating it. Parsing code often seems like the opposite in my experience, where it is more difficult than writing it yourself.
- auggierose 5mo agoLol. If Gowers sends you a piece of math he doesn't quite understand because he thinks that you might, that is something you celebrate.
- fulafel 5mo agoLink to source blog post: https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/ https://gowers.wordpress.com/2026/05/08/a-recent-experience-...
- dang 5mo agoThat's the top link (i.e. that the title is linked to), no?
- fulafel 5mo agoIndeed, the body in the post made me think it was a url-less submission.
- dang 5mo agoAh yah. Somehow in the last 6 months or so we started using the toptext in new ways - e.g. to share related links and whatnot. I know it's ambiguous (is this a text post or a linked post? did the submitter write the text or was it the mods?) - but my thought is that over time it might naturally evolve into something clearer.
- SubiculumCode 5mo agoI honestly can't say this isn't AGI anymore. AGI shouldn't be a bar so taboo that it has to be at the extreme capability in every domain. What human is? This is as AGI as it needs to be to get my vote. And it's scary.
- agiipullor 5mo agoto quote Demis Hassabis, "these models can solve frontieer problems in math, but also fail in really dumb ways at trivial questions - the car wash question". jagged AGI
- MrScruff 5mo agoIt's ASI with jagged intelligence, which is probably what it will remain for a while. It still sounds to me like remarkable automation rather than something that's expanding the frontier of human knowledge, for now at least.
- einrealist 5mo ago"After 16 minutes and 41 seconds, it came back" ... "further 47 minutes and 39 seconds" ... "After 13 minutes and 33 seconds" ... "After 9 minutes and 12 seconds" ... "After 31 minutes and 40 seconds" ... plus other computations Anyone spotting the issue here? What did that really cost? I am not against compute being used for scientific or other important problems. We did that before LLMs. However, the major LLM gatekeepers want to make all industries and companies dependent on their models. And, at some point, they need to charge them the actual, unsubsidized costs for the compute. In the meantime, companies restructure in the hopes that the compute costs remain cheap.
- colordrops 5mo agoStill not as bad for the environment as animal agriculture, and animal agriculture is absolutely not necessary and only causes harm and suffering for taste pleasure. At least with LLMs we get many positive advancements from them. I don't see these sorts of comments every time someone posts a burger review.
- einrealist 5mo agoDid I praise our animal agriculture anywhere?
- colordrops 5mo agoWe have a wide audience.
- sidkshatriya 5mo ago> "After 16 minutes and 41 seconds, it came back" ... "further 47 minutes and 39 seconds" ... "After 13 minutes and 33 seconds" ... "After 9 minutes and 12 seconds" ... "After 31 minutes and 40 seconds" ... plus other computations Anyone spotting the issue here? What did that really cost? Whatever the Joules... (convert to $ using your preferred benchmark price) it is a fraction to what it might take a human Ph. D. weeks to feed and sustain themselves when working on the same problem. The economics on LLMs is just unbeatable (sadly) when compared to us humans.
- zuogl 5mo agoThe HTML generation is surprisingly good because the training corpus for markup is cleaner than most programming languages.
- adaml_623 5mo ago"It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour" This comment about time is very interesting to me. I know it's "just" doing mathematical proofs but the possibilities of speeding up planning, proposals and decision making in the physical world should excite people.
- bambax 5mo ago> quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques Creativity is connecting ideas from different domains and see if something from one field applies to another. I do think AI is overhyped generally; but a major benefit from AI could be that after ingesting all the existing human knowledge (something no single human can ever hope to achieve) it would "mix and connect" it and come up with novel insights. Most published research sits ignored and unread; AI can uncover and use everything.
- imiric 5mo ago> Creativity is connecting ideas from different domains and see if something from one field applies to another. That's true. The question is whether the produced pattern has any value. LLMs are incapable of determining this, and will still often hallucinate, and make random baseless claims that can convince anyone except human domain experts. And that's still a difficult challenge: a domain expert is still needed to verify the output, which in some fields is very labor intensive, especially if the subject is at the edge of human knowledge. The second related issue is the lack of reproducibility. The same LLM given the same prompt and context can produce different results. This probability increases with more input and output tokens, and with more obscure subjects. The tools are certainly improving, but these two issues are still a major hurdle that don't get nearly as much attention as "agents", "skills", and whatever adjacent trend influencers are pushing today. And can we please stop calling pattern matching and generation "intelligence"? This farce has gone on long enough.
- agiipullor 5mo ago> And can we please stop calling pattern matching and generation "intelligence" thats literally what an IQ test tests - abstract pattern matching. but I guess you dont like IQ tests either
- hirvi74 5mo agoSome IQ tests like the WAIS test on retained, common facts. They are not all just pattern matching. Also, I do not like IQ tests either (having taken one myself). They are unbelievably boring, pointless, and measure more than just "intelligence."
- lysecret 5mo agoThere is a great recent episode of latent space about a similar topic it’s worth a watch even with the click baiti thumbnail and title https://youtu.be/9d899Ram9Bs?is=pQMoVmlWVsTNKfRK https://youtu.be/9d899Ram9Bs?is=pQMoVmlWVsTNKfRK
- locknitpicker 5mo agoFrom the article: > Conversely, for problems where one’s initial reaction is to be impressed that an LLM has come up with a clever argument, it often turns out on closer inspection that there are precedents for those arguments, so it is still just about possible to comfort oneself that LLMs are merely putting together existing knowledge rather than having truly original ideas. How much of a comfort that is I will not discuss here, other than to note that quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques. This is exactly what leads me to believe that the real impact of LLMs in human history is yet to come. My work as a researcher was mostly spent on two classes of workloads: reading papers that were recently published to gather ideas and keep up with the state of the art, and work on a selection of ideas gathered from said papers to build my research upon. It turns out that LLMs excel at the most critical component of both workloads: parsing existing content and use it when prompting the model to generate additional content based on specific goals and constraints. I mean, papers are already a way to store and distribute context.
- ionwake 5mo agoone thing I was wondering, is, if LLMs are word completions seemingly coming up with new solutions could this just be because stuff that was kept secret and now - is no longer is due to ingestion? I dont know enough about it tho
- dist-epoch 5mo agowhy would you keep secret this particular mathematical idea? it's not extraordinarily important, it's not on the path to some other major result, doesn't seem useful in financial trading. even author calls it good reasonable problem for a PhD thesis.
- arjie 5mo agoThe question of where the creative input is was a big thing around Experiments in Musical Intelligence and co-composing. But it seems perhaps that it’s a transient state we needn’t spend too much effort it. The machine has failed to disappoint repeatedly. Perhaps this is as far as it gets or perhaps we will be like people in Catching Crumbs by the Table by Ted Chiang where almost all science is interpretation of papers by vastly greater intellects.
- zingar 5mo agoThe post talks about LLM+human contributions being recognized in some different category from human-only. But is it possible to spot the difference between the two?
- amelius 5mo agoMakes sense as a mathematician basically has two powers (1) using their intuition and (2) an enormous amount of mental stamina. A mathematician builds their intuition by reading maths books. It is thus not surprising that an LLM is well equipped to take over the tasks of the mathematician.
- MrDrDr 5mo ago> "Even though I can motivate it in retrospect, ChatGPT’s idea to use h^2-dissociated sets to control relations of order at most h feels quite ingenious. As far as I can tell, this idea is completely original." The question that keep bothering me is can an LLM generate an idea that is truly novel? How would/could that actually happen? But then that leads to the question - what are we actually doing when we think? Perhaps it's as simple as the ability to just make mistakes that matters, the same things that powers evolution. As long as the LLM can make mistakes, it's capable of generating something genuinely novel. And it can make more mistakes much faster than we can.
- ikari_pl 5mo agoHow do you define a new idea? To me, it's rearranging the information you had in a way that hasn't been applied or published before. That's literally what LLMs are built for.
- LiamPowell 5mo agoTrivially the answer is yes by the infinite monkey theorem. If we allow the sampler to pick any token then any stream of arbitrary tokens can be generated. Therefore if an original idea can be represented with written words then a LLM can generate it. That is perhaps not the most satisfying answer, but if you want a better one you'll need to provide a function that determines if an idea is original.
- eterm 5mo agoMy own take, and it's veering into the Philosophy of Mathematics, but there's a debate about whether Mathematics is "Invented" or "Discovered". If it's "invented", then it requires ingenuity. If it's "discovered", then it was always already there, just waiting for the right connections to be made for it to be uncovered and represented in a way we can understand. Invention requires ingenuity, but discovery does not. So if LLMs can generate truly novel mathematics, for me that settles it that mathematics is indeed discovered, as LLMs are quite capable of discovery yet I don't consider them possible of invention.
- MrDrDr 5mo ago
- zkmon 5mo ago>>
- zkmon 5mo ago>> but it was definitely a non-trivial extension of those ideas, and for a PhD student to find that extension it would be necessary to invest quite a bit of time digesting Isaac’s paper The "non-trivial" is for human abilities. The weights lifted by a crane are also "non-trivial". People keep getting amazed at machine's abilities. Just like a radio telescope can see things humans can't, microscope can see the detail humans can't, we need not be amazed. The sensory perception of patterns is at different level for AI. It's a machine.
- svnt 5mo agoToo many people are wrapped around the ego axle thinking (assuming) their ideas are both them and somehow unique and special. It usually takes dissolving that, often through difficult experiences, before they can see it as a machine, something that could be separated from them.
- dag100 5mo agoI think the more pressing issue is that there isn't really much space left for humans in the economy if thinking can also be automated.
- iandanforth 5mo agoI found the section on publishing very interesting. Even if the quality of the output is up to snuff, where should it go? Arxiv doesn't allow AI written work. The author proposes that only work that has been certified by human should be published. However, now the field is in the same boat as software engineering where we are facing a glut of pull requests and not enough time and people to review them.
- kang 5mo ago> The lower bound for contributing to mathematics will now be to prove something that LLMs can’t prove, rather than simply to prove something that nobody has proved up to now and that at least somebody finds interesting. 5.5pro is amazing but this implication might not be true & is the core argument of this piece. AI will prove all sort of things - interesting, boring & incorrect. To sort it will be the task of the PhD.
- layer8 5mo agoThe task of a proof verifier is much simpler than the task of a proof finder (it’s basically equivalent to P vs. NP), and hence the bar for the required skills is lower. Merely verifying proofs isn’t research, and doesn’t impart research skills.
- kang 5mo agoVerification on its own is not research, but judgement is research. "Hey, Prove something a machine can't", sure I can't, "Hey, Say something worth proving & judge it well", ah, now I might have a few unique observation/ideas/curiosities/problems from my having being a human. Imo, the feeling of intelligence or the process of originality(originativity) test for ai is subjective & is coming down to 4 paths: novel relative to a reference class, valuable within a domain, counterfactually sensitive to internal state and environment, and revisable through learning.
- ianm218 5mo agoVerification is generally a much lower bar than solution generation. I don’t think it’s likely sorting out the right from wrong will end up being this huge PhD level effort.
- kang 5mo agoVerification & solution generation are both part of problem generation & defining the passing test - judgement.
- casey2 5mo agoI think mathematicians like LLMs because this is the first time we have something like a computer for the kinds of math most people do, high level, hand wavy abstractions that are (relatively) easy for people to grok but hard to explain to traditional computers.
- sexylinux 5mo agoUnfortunately it still does create errors. This is of enormous importance but still is being actively ignored by many professionals or dismissed as as a minor issue. Our emotional human brains are very enthusiastic about these new kind of "intelligent" products ("partners") and we want to believe so hard that they are finally "there" that we tend to ignore how big of a problem it is that LLMs carry a fundamental design problem with them that will make them produce errors even when we use a grotesque amount of resources to build "bigger" versions of them. The potential for errors will never go away with the current AI architecture. This is a fundamental paradigm shift in computing. Instead of putting a lot of energy into building an architecture that will produce reliable results, we are now maximizing on a system / idea that will never give us 100% reliable results. Basically it is just a marketing stunt. Probably the computer science guy building it knew very well that he would still need some fundamental break troughs to get to a real product, but the marketing guy saw that there is still potential to make a lot of money by selling a product that will produce correct results only 80% of the time. The marketing guy was right and marketing is now dominating science, but humanity will pay a big price for that. Putting enormous amounts of money into a fundamentally flawed system that we can not optimize to produce reliably error free results is just stupid. The big achievement of "classical" computing is that the results are reliably error free. We have still some known issues eg. with floating point math and bad blocks on disk / bit flipping etc. but these are observable and we can handle / avoid them. Generally "non-ai-computing" was made so reliable, that we can depend on it for many very important things. This came not by accident but was created by a lot of people who put a lot of resources into research to achieve that result. LLMs introduce a level of uncertainty and unreliability into computing that makes them practically useless. Because if you have enough knowledge to verify the result and AI is only quicker in producing the result, what is the point then putting so much resources in it (besides making money by re-centralizing computing, of course). Verifying a lot of results that have been produced quicker is still slow, so the people who are now just AI verifiers should just produce the results themselves, makes the whole process quicker. AI is only of value if it can produce results about things that you or your organization does not know anything about. But these results you can not verify and therefore potentially wrong results can be fatal for you, your organization and all the people that are affected by actions generated based on these wrong results. Many people have already been killed because decision makers are not able to follow that very simple logic. So we can still create "interesting and enjoyable results", but finally it is a gigantic miss-allocation of resources of historic idiocy. It fits, of course, very well in a timeline where grifters are on top of societies around the world. It is a fundamentally wrong path that should not be followed and scientists around the world should articulate exactly that instead of producing marketing blog posts for a system with such fatal inherent issues.
- quinndupont 5mo ago[flagged]
- vessenes 5mo agoI don’t love the tone here, but I do think you get at a key question in mathematical philosophy. Mathematicians have engaged, vigorously, on this very philosophical question for centuries - is math discovered truth, or is it more akin to building an edifice where you first define the materials, then the structure, and see where it leads? There are lots of strong feelings on both sides. For instance: “God created the integers, the rest is the creation of man” — Kronecker, 19th century sums up one particular perspective. To me, it’s probably a mix of both - some fantastic results in imaginary numbers show up as describing key electromagnetic effects many decades after they were first ‘discovered’ by theoretical mathematicians. NB: My original comment led with a pejorative, which was rightly flagged.
- dang 5mo ago> Who hurt you bro? Please don't respond to a bad comment by breaking the site guidelines yourself. That only makes things worse. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- dares2573 5mo agoI think the biggest advantage of ChatGPT compared to Claude is that there are fewer things outside the model itself, such as KYC, account bans, etc.
- solenoid0937 5mo agoThis is just grossly misinformed. OAI and Anthropic both require KYC for models of similar intelligence. They both do account bans if the classifiers fire wrong. You simply hear about it less with OAI because Codex has fewer prosumers.
- tmp10423288442 5mo agoCan you name any instance of OpenAI being as trigger-happy with bans as Anthropic has been in the past few months? Codex may have fewer prosumers, but they've added a lot in that time.
- solenoid0937 5mo agoThere will obviously be more reports about Anthropic because hating on Anthropic has been trendy lately, it has more users, and gullible people (including most HN'ers) fall victim to OAI's guerrilla marketing. In terms of bans and KYC they are not meaningfully different.
- MagicMoonlight 5mo agoChatGPT pro is garbage. It’ll spend 20 minutes on an answer, doing all kinds of ridiculous things like writing scripts… instead of just outputting plaintext. And then the answer isn’t even right.
- OsamaJaber 5mo agothe bottleneck isn't generation, it's verification
- alpama 5mo agoThis is scary. Ai is growing faster than our knowledge. We are not prepared
- deleted 5mo ago[deleted]
- cindyllm 5mo ago[dead]
- TrackerFF 5mo agoThe vast, vast majority of students going into higher education this fall will not contribute much to science until 4-5 years down the road (should they do research). Realistically 6-7 when they're in full swing with their Ph.D. If we look where these models were 5-7 years ago...the existential threat of the Ph.D. was not even on the radar back then. The people finishing up their doctorate now are the first that can truly leverage these tools. Now, if these to-be researcher students feel defeated (enough to quit), or completely lean on AI models the work for them, we're going to have a problem. Same with the funding of those Ph.D. positions. If we move away from "funding to produce researchers" to "funding to achieve results", will money that was usually spent to fund Ph.D. students start to flow towards compute? If we look at it a bit cynically: Some researcher will be able to pump out a lot more papers by spending money on compute, than a couple of years of training students. Interesting times. But also so much uncertainty. I feel terrible for the students that will have to decide now what they want to do, with all this knowledge.
- odyssey7 5mo agoIt’s not as if institutions have been lavishing PhD students with money up until now. As an afterthought budget item, those funds aren’t exactly attractive targets to raid for pursuing an expensive, different process.
- robot-wrangler 5mo ago> Now, if these to-be researcher students feel defeated (enough to quit), or completely lean on AI models the work for them, we're going to have a problem. [..] If we look at it a bit cynically: Some researcher will be able to pump out a lot more papers by spending money on compute, than a couple of years of training students. Obviously this is already happening and will accelerate. Outside of grad work, you could already just buy a degree. Certainly in the softer disciplines, you can currently just buy a phd thesis and a good publication history. If you're in industry instead of academics, you can even buy a promotion. If your employer gives an AI budget to all workers then you quietly double that budget out of your own pocket for as long as it takes to get a promotion, then stop and just enjoy a bigger paycheck.
- ndkap 5mo ago
- goopthink 5mo agoAn interesting takeaway is that heretofore most of that advances have been not from “invention” but from a breadth of visibility. LLMs have been able to be “creative” because of the volume of work that they cover and can draw lines and associations between, not in discovering things that did not exist previously (though an argument can be made that something like AlphaFold was “discovering” and “intuiting” associations that were not explicit anywhere previously, uniquely found by the AI… but I’d argue back something about the bitter lesson and we’d go on for more than a few threads). Somewhat ironic then, to not make this more explicit in an article about solving a combinatorial problem.
- rklampp 5mo agoGowers has always been a proponent of Lean (naturally). He receives funding from the "AI for Math" fund, which is sponsored by a fund that is a front organization for venture capitalists: https://www.renaissancephilanthropy.org/ https://www.renaissancephilanthropy.org/ The "brighter future" of course is that everyone is redundant and all capital is further concentrated. It is always Gowers, Tao and Lichtman (math.ínc startup) who are pushing these technologies.
- logicprog 5mo ago> It is always Gowers, Tao and Lichtman (math.ínc startup) who are pushing these technologies. In your mind does this mean that they are lying, or driven by motivated reasoning and cognitive bias, or whatever you'd like to say? Because I feel like people bring up these facts as a way to discount everything that these people are saying, but whether or not they've chosen to align themselves with AI aligned venture capital funding or not. The question is really, did what they say is happening happen or not? Are these capabilities real or not? To my mind, mathematics is pretty definitely, externally, objectively verifiable, so it would be easy to catch them in a lie. In the case of the Erdös problem that was recently solved in a novel and productive way, it wasn't even initiated by them and the chat GPT transcript is public for all to see. And the proof could easily be verified by other people, for instance. In addition, I think it's unlikely that they're not explaining things as they honestly see them and also doing their due diligence to make sure that they are seeing them as close to correctly as possible. Because their positions with these organizations not to mention their entire reputation and life's work and passion depends on their reputation in academic mathematics. If they were to give that up by falsifying these claims or not verifying them sufficiently, they would lose everything. I think it's also worth pointing out that it is totally possible for someone to align themselves with such organizations after the fact because they agree with them instead of being bought out by such organizations. Otherwise, it would be possible to dismiss the opinion of anyone working at any NGO dedicated to being against AI and denying AI's capabilities or whatever, as well by the same logic of their salary being paid by an organization dedicated to pushing those ideas.
- chalr 5mo agoThere have always been attempts at settling all mathematics by using mechanized approaches. Often by mathematicians who already had made an impact and then wanted an automated approach. The Bourbaki group was one of the first who attempted a mechanized approach (using pen and paper still of course) to set theory and were literally accused of wanting to end all mathematics. The approach was largely ignored in practice. Gowers and a handful of others who work on computerized approaches also seem to want to end human mathematics and have sharecropper mathematics for a monthly tithe. So far they are largely ignored in practice.
- robot-wrangler 5mo agoA very interesting comment from Baez, I'll just quote part of it. > Where does the value of thinking and having deep ideas come from? We need to think about this now. If it comes primarily from their scarcity – the fact that having certain ideas is hard – then indeed this value may drop precipitously when the manufacture of ideas can be automated. But if the value comes from the utility of the ideas – the benefit that the idea brings – then the story changes: perhaps creating more good ideas is actually better, not worse. Here I’m using “utility” in a broad sense, not just in the sense of what people often call applied mathematics. > In other words, mathematicians may need to adjust to a transformation from a scarcity economy to an abundance economy. https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/#comment-540064 https://gowers.wordpress.com/2026/05/08/a-recent-experience-...
- qwrahg 5mo agoI note that it is always the same online pundits (even if they are distinguished academics) who push anything new. Meanwhile Wiles and Perelman stayed offline and solved real problems.
- DoctorOetker 5mo agowould Wiles be willing to transcribe his proof for the metamath verifier? it can be done offline indeed...
- williamstein 5mo agohttps://github.com/ImperialCollegeLondon/FLT https://github.com/ImperialCollegeLondon/FLT
- DoctorOetker 5mo agoI asked about Wiles, because others frequently run into issues while formalizing.
- robot-wrangler 5mo ago
- eranation 5mo agoLike coding, if you get inspired by AI for a novel idea, and can reproduce the same result independently (could code the same thing by hand) or at least understand and check every single argument (self review your code, test on your machine) and get it peer reviewed (code review, but with a real human) then I don’t see why the industry accepts the latest iteration of ChatGPT being 99% written by codex, but rejects a valid math result inspired by it.
- tmp10423288442 5mo agoIt's interesting that ChatGPT Pro is the real deal that can write novel physics or math papers (for certain values of novel), while Claude Pro is crap that, depending on the A/B test, may not even provide Claude Code or at the very least doesn't provide Opus. Shows how LLM naming conventions are currently a mess.
- robot-wrangler 5mo agoDespite this coming from an independent expert and not from OpenAI, we need to be honest that this is more like a marketing campaign than open science. I assume the progress really is valid, the experts are indeed impressed, and we have the accurate time that it takes to produce the results. What we don't have is detail about true cost, or a CoT trace, or anything like that. The implication is: we're ready to let everyone go wild with this very soon. Ok, go wild with what exactly? How do we know that influential VIP users who might make very friendly blog posts aren't getting allocated exclusive access to a billion dollars worth of hardware when they ask questions? I mean literally giving a certain group of people temporary privileged access to like 90% of all available compute would be a completely reasonable business decision for OpenAI. Would a reveal like that change how we think about the result? What if half that amount of cash/compute could enable some completely non-AI approach of numerical brute forcing that settles the question even if it didn't write the paper? My other question is always whether the latest is purely using giant models or if we're now deeply into harnesses that use MCTS and such. Understandable to keep that a trade secret I guess. But IMHO we should at least get the CoT trace as a proxy for true cost, or else maybe we're just getting played to do the hype for corporate.
- electriclove 5mo agoMarketing campaign???
- energy123 5mo agoWe know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro, not by familiar names that OpenAI is sneaking additional compute to behind the scenes (which is too conspiratorial of an explanation for my liking regardless).
- robot-wrangler 5mo ago> We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro Where's that? The stuff I've seen is from celebrities. Were those problems as hard as this one, or the ones that Tao posts about? Regardless.. what's the argument against more transparency here to just settle this kind of thing? > which is too conspiratorial of an explanation for my liking regardless OpenAI is not, in fact, open. Why do they deserve the benefit of the doubt? Regardless.. special treatment for special customers isn't conspiracy, it's SOP literally everywhere and especially if you're helping to beta test. Anyone who's ever interacted with any technical account manager has seen waived quotas, free resource allocations, etc. The quid-pro-quo is obviously that your cheap early access means you get to give talks at a conference (or make a blog post that a lot of people read and talk about).
- jacktu 5mo agoThat's a real shift. The value of an open problem used to be that it was unsolved. Now an open problem needs to be unsolvable by something that can read the entire literature and try a hundred approaches in an hour.
- highfrequency 5mo ago> LLMs have got to the point where if a problem has an easy argument that for one reason or another human mathematicians have missed (that reason sometimes, but not always, being that the problem has not received all that much attention), then there is a good chance that the LLMs will spot it.
- Jweb_Guru 5mo agoThis jives with what I've experienced in the brief time I had access to 5.5 Pro. It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. The downside (not noted in the article, but noted by others here) is cost. It uses tokens at an insane rate, the tokens cost a lot, and using it with subagent flows that you can use to have it tackle large problems with high accuracy costs even more. It is also much "slower" for large scale problems because of context limitations -- it has to constantly rediscover context for each part of the problem, and in order to make it accurate you need to wipe its context before progressing to the next small part, or launch even more agents. For mathematical proofs like these, where the required context to understand the problem and proof besides stuff that's already available in its training set is small and the problems are considered "important" enough, this might not be a problem, but for many of the tasks I would like to use it for (ensuring correctness of code that affects large codebases, or validating subtle assumptions) it definitely is one. So I think it will be a while before the impressive capabilities of these models really percolate into our lives as programmers, unless you're one of the lucky ones given unlimited access to 5.5 Pro.
- y1n0 5mo ago> This jives with what I've experienced Just as an fyi, the word you are looking for is jibes. Jive is something else entirely.
- refulgentis 5mo agoThat ship sailed looooong ago.
- ignoramous 5mo ago> looooong Just as an fyi, the words you are looking for are ages/eons/an eternity.
- theptip 5mo ago> what should we do with this kind of content? Had the result been produced by a human mathematician, it would definitely have been publishable, so I think it would be wrong to describe it as AI slop. On the other hand, it seems pointless even to think about putting it in a journal, since it can be made freely available, and nobody needs “credit” for it (except that Isaac deserves plenty of credit for creating the framework on which ChatGPT could build). I understand that arXiv has a policy against accepting AI-written content, which makes good sense to me. So maybe there should be a different repository where AI-produced results can live. But various decisions would need to be made about how it was organized. Interesting question, I guess a starting point is “moltbook”, but perhaps a better one is something like GitHub, where Lean proofs and preprints can go, and trending items can get boosted. I also think that posting this stuff on x or bluesky has merit, but again the existing paradigm doesn’t quite work; perhaps you can create a completely separate identity for your agent (à la Moltbook) but I think you want some sort of reputational association with the human piloting the agent, at least for now. (Maybe eventually there are enough agents critically engaging with content so that “interesting” results get agent likes, and so we’ll-piloted agents stand on their own merit.)
- readgrounded 5mo agoQuantitative finance went through a smaller version of this in the 2010s. The apprenticeship was building a Black-Scholes pricer from scratch, then a vol surface, then a calibration loop. Sweat problems that taught you what the math meant. Then libraries got good, platforms got good, and a junior could be productive without ever reeeeeally knowing how it worked. On some level yes the finer detailed knowledge is going to be lost because it gets locked in but in some ways we do get a "higher level api" to presumably solve more difficult problems.
- ortusdux 5mo agoAn important lesson in web/blog design - I cannot for the life of me figure out who this author is (using only the website).
- qzgrid37 5mo ago[dead]
- robeym 5mo agoI think progressing as humans is something to be proud of. I care less about who gets credit and more about what we can now do. I also do not think this makes people less capable of solving hard problems. The bar just moves up. More people can now work on harder problems with better tools. If the goal is credit or proving real skill, then focus on harder problems, like ones AI can't reach
- screenstop 5mo ago[flagged]
- ljasdjasdfasdf 5mo ago[dead]
- YeGoblynQueenne 5mo agoAll this sounds to me like mathematicians spooking themselves with stories of how ChatGPT solved a problem, when it's mathematicians solving a problem using ChatGPT as a tool. E.g. from the twitter thread by Timothy Gowers: >> All I did was say things like, "Yes, it would be great if you could explore that idea and see whether you can get it to work," or "Could you rewrite that argument as a LaTeX file in the style of a standard mathematical preprint?" Yeah, so all he did was take the horse to the water and make the horse drink. The collaboration with the other two mathematicians wasn't a trivial part of the problem solving either: every time Timothy Gowers figured ChatGPT had goe somewhere with its problem-solving, he stopped, asked it to render the answer in LaTex, and sent the answer off to be verified by the other two. The reason for that is not to be underestimated: ChatGPT can produce answers to questions you ask it for as long as you ask it to do so but it has no capability to determine whether an answer is correct or not. That's why it needs a human with domain expertise to evaluate those answers. And of course to discard wrong answers in the process, because of course the process that's described here glosses over many false starts and back-and-forths and "you're absolutely rights, here's a new version of that"'s etc. that are common experience when using LLMs for problem-solving tasks. The existential questions that the article poses about mathematics then are easily answered by taking all of the above into account. If LLMs are a useful tool for mathematicians, then nothing changes. Mathematicians of all levels can still do their job and perhaps do it faster or better with the new tool. If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening.
- whimsicalism 5mo agoSorry, just so I fully understand your comment - your claim is that asking it to “explore that idea further” and “write the paper in latex” constitutes “taking the horse to the water and making the horse drink”? thank you for the morning laugh
- deleted 5mo ago[deleted]
- YeGoblynQueenne 5mo ago
- richard_chase 5mo agoDude needs to step down off his pedestal before he gets knocked down.
- dpweb 5mo agoDon’t quite agree with the implication that if the answer to a problem is readily available (from an LLM) there is no use in the struggle to find the solution. However I think it’s very important to approach such questions objectively, or at least self uninterested, and not as one who’s worried about one’s job or sense of self worth threatened by LLM technology. The value is in the development of one’s own mental faculties. In math classes they tell you you have to work the problems. Even if LLMs become capable of solving entire classes of problems that that set expands over time, the value in developing one’s ability never goes out of style.
- zhouquanxi 5mo ago[flagged]