12 ms·
Human mathematicians are being outcounterexampled
- wizzwizz4 3mo agoHuman mathematicians have been being out-counterexampled for at least two decades. The main difference, as I understand, is that (A) we now have a lot more compute to throw at such things, and (B) it is currently trendy to do so. But the sizes of counterexample we're seeing are around about what I'd expect pre-generative-AI counterexample search systems to be able to find. It's not easy to find a counterexample to the Jacobian conjecture, by any means – by which I mean to say that naïve brute-force search will take too long – but the scope of existing searches listed on Wikipedia[0] suggest that many tricks are already known, and that people just hadn't looked, systematically, for a counterexample in three variables before. Wikipedia writes: > Tzuong-Tsieng Moh checked the conjecture for polynomials of degree at most 100 in two variables.[17][18] where reference 17 is from 1983, and reference 18 is a preprint with no given date. Knowing very little about this problem, my impulse is to side with the unnamed faculty member cited in the article: > [who] said to me that the fact that the counterexample was so easy to find just indicated that humans had not spent enough time thinking about the problem, For context, the auto-generated counterexample is in three variables, has degree 7, and was discovered in 2026.
- NitpickLawyer 3mo ago> about what I'd expect pre-generative-AI counterexample search systems to be able to find. The difference today (and the reason why everyone is excited about it) is that the same system that does advanced math can write poetry, play an above average game of chess, code frontend/backend stuff and do cybersec. These are not "expert systems", nor are they trained for each task individually. That's the catch. > just indicated that humans had not spent enough time thinking about the problem Heh, this is a weak excuse. We've seen variations on this theme every time something cool gets solved by the models.
- am17an 3mo ago> nor are they trained for each task individually. They are explicitly trained for each task individually.
- NitpickLawyer 3mo agoThey are not. Pretraining just dumps every piece of content in the mix, and only has one objective - next token prediction. And you can get pretty good results even with base models, you just have to manage context differently. Later stages (mid, post training) involve RL that "surfaces" the right "traces" out of the pre-training. But they are not trained individually, as we used to do.
- am17an 3mo agoHave you looked at the what data companies (e.g. Scale, Mercor) hire for? Why do you think Meta records their employees every keystroke/mousestroke/eye-movement? EDIT: just re-read your comment. I don't think you have a good understanding here, no offence.
- NitpickLawyer 3mo agoNone taken, but it would be odd, since I've been training GOFAI models since 2010s and have had LMs in production since before chatgpt came out (before RLHF), so I think I have a pretty good understanding :) But I'm always open to learning. I think the misunderstanding comes from "individually". You are thinking about diverse datasets, but that's not what individually means in this context. In ML individually trained means that for each task you prepare an architecture, dataset and eval and train that model on that data. And each model has its own objective that you train for. In LMs the objective is singular, for every data point - next token prediction. And, importantly, you train on every datapoint, the more diverse the better, but not independently. The cool thing is that training on diverse datasets improves scores on other downstream tasks, while the training objective is the same. So you can have a training run on common crawl + programming that improves scores on logic puzzles, or common crawl + novels that improves scores on planning tasks. But the important thing is that it's all trained together, not independently.
- gowld 3mo agoThe word "just" is the mark of a coward. AI is only as good as a human mathemetician, which we don't have enough of? "just"? https://en.wikipedia.org/wiki/Jacobian_conjecture https://en.wikipedia.org/wiki/Jacobian_conjecture > The conjecture was first stated for two variables by Ludwig Kraus in 1884 [...] an example of a difficult question in algebraic geometry that can be understood using little beyond a knowledge of calculus. > The Jacobian conjecture is number 16 in Stephen Smale's 1998 list of Mathematical Problems for the Next Century. It was notorious for the large number of published and unpublished false proofs that turned out to contain subtle errors.
- wizzwizz4 2mo agoThis is a counterexample for the three-variable case, which hasn't been studied as much. The long-standing two-variable case remains open.
- paulpauper 3mo agomathematicians have been using computers for well over half a century, but this was after "bounding" the problem first and then running through the cases with a computer. Now AI is doing the first part. However, mathematicians are still needed at crafting prompts, and knowing where to look, still. The prompt for the Jacobian conjecture was obviously not random. the search space is too big to just try all the combinations of 3 variable polynomials.
- hobonation 3mo agoThis is the best take. Computers don't care about this stuff. A computer could make a movie, but only a human can appreciate it. We're a good team, and that's ok.
- nullsanity 3mo ago[dead]
- criddell 3mo agoFor now. I wonder if we will ever get to the point where the computer starts doing mathematics that we just can't understand. Surely there must be some limit to what we can understand (like how a gorilla will never understand prime numbers, there are probably limits to our intelligence as well).
- zeroonetwothree 3mo agoMathematics only really matters insofar as humans can understand it.
- stabbles 3mo agoNot really, a theorem with a hard proof can have simple but important corollaries. It's also not unthinkable that theorems exist with proofs that cannot reduce to something simple/short.
- 3mo ago
- satvikpendem 3mo agoThat's a good thing. It saves people wasting time trying to prove something they now know to be false, so that they can move on to other things to prove, it's a more fruitful use of humanity's time overall at least in the field of mathematics.
- parpfish 3mo agoproofs by counterexample are effective but ultimately unsatisfying. they get you to an answer but they don't help help you understand and bend you r mind into seeing how the math works and lead you on to the new set of questions. and for now as long humans are going to judge of what counts as an elegant or illuminating proof, there's going to be work for human mathematicians
- taneq 3mo agoMaybe I’m just not pure enough but I find the whole concept of proof by counterexample to be elegant, and I don’t see why proving that something must be true is superior to proving that it can’t be false.
- bananaflag 3mo agoYou mean proof by contradiction, which is something different.
- taneq 2mo agoHmm, I think I might have conflated the two. Thanks!
- chowells 3mo agoIt's elegant if all you're concerned with is whether a conjecture is true or false. Answered, move along! But mathematics is not a collection of facts. Mathematics is the study of abstraction. And what do you learn from a single data point? What can you abstract from that? That's why just being a counterexample isn't really interesting. There has to be more than "counterexample" for there to be something to abstract. Was it generated from an analysis of the problem? Can the counterexample be generalized to explore the problem further? Is the counterexample a surprise in a way that suggests something is missing from current understanding? Being a counterexample doesn't mean that something isn't interesting to a mathematician. But it's also not the interesting part.
- QuesnayJr 3mo agoIf the poster's (is it Kevin Buzzard?) suggestion works out and AI finds a counterexample to the Hodge conjecture, that would be a really big deal. It's one of the Millenium problems, for example. One thing that he mentions that already quite surprising is that AI was able to autoformalize the Golod-Shaferevich theorem and proof.
- mcshicks 3mo agoIt is Kevin Buzzard. It's kinda small font on my phone but if you look at the "about xena" link it says it's his site.
- williamstein 3mo agoAgreed, it's definitely Kevin. His writing style is unmistakable.
- OG_BME 3mo agoI think he was being provocative and maybe a bit tongue-in-cheek when he said that. A candidate object alone doesn't resolve the Hodge Conjecture. Any apparent counterexample would have to prove that no algebraic cycle exists, no invariant subspace exists, or that every element of an infinite ideal is nilpotent. Much harder, but not impossible.
- kevinbuzzard 2mo agoIndeed I was being slightly tongue-in-cheek -- but if you ask geometers whether they believe the Hodge conjecture then you certainly don't always get an unqualified "yes"! This is in contrast to e.g. asking number theorists whether they believe Birch--Swinnerton-Dyer, where they are almost always very confident.
- hintymad 3mo ago> The Jacobian Conjecture Interestingly, Yitang Zhang of the twin-prime-conjecture fame spent 7 years working on the Jacobian conjecture under the advisor Tzuong-Tsieng Moh at Purdue. A key step in his thesis used a corollary of Moh's. It turned out that the corollary was incorrect. As a result, Moh refused to write any recommendation letter for Zhang, and Zhang couldn't find any teaching or research job and ended up spending years working at a Subway[1]. Imagine Zhag had ChatGPT in 1986 when he started working on the Jacobian Conjecture. [1] Of course now this has become an inspiring story. That said, the story definitely invokes complex emotions. The best way to describe it is probably this Chinese poem, which I have no idea how to translate: 庾信平生最萧瑟,暮年诗赋动江关
- derbOac 3mo agoInspiring? Because of the twin prime conjecture success following his time in the wilderness? I suppose so. I'm tired of tales like this in academics though. That's not a criticism of you for telling the tale, I'm just so tired of this kind of thing in academics in general. So, so, so much politics and public reputation management. Zhang should have never had to suffer like that. As my own research has drifted more into math, I've been surprised at how many assertions in the literature turn out to be false. Not just false, but propagated into the applied literature extensively, and even when you point out the problems a lot of defensiveness and denial about it along the lines of Zhang's story. I agree about wondering what would have happened if LLMs had been around in 1986. My guess is the outcome would have been the same for the same reasons? My experience with LLMs in proofs is they can be very helpful, but also very wrong. It's like having another person with another set of hunches about what path to go down.
- hintymad 3mo ago> So, so, so much politics and public reputation management. Zhang should have never had to suffer like that. Very true. Unfortunately, when there are people, there will be politics. I remember when reading Yau's autobiography, I kept marvel how much calculation, or "politics" if you will, that Yau mentioned or implied in the book. > My guess is the outcome would have been the same for the same reasons? At least Zhang didn't have to spend 7 years working on the Jacobian conjecture. He said in an interview that he always wanted to work on number theory. Moh asked him to work on Jacobian, and he obliged.
- Dove 3mo agoWhen I was in grad school, I had the opportunity to take a course from my adviser in which he discussed his current research and some open questions. It was a relatively accessible subject area and the questions were sometimes easy enough that we could meaningfully contribute. On one particular Friday afternoon, he stated a conjecture that he hoped was true, and invited us to try to help him prove or disprove it. It was the sort of thing that he really wanted to be true; he liked things smooth and beautiful. I, on the other hand, hoped it was false as I like the weird and exceptional in mathematics. It was also the case that I had absolutely no command of the sort of machinery that one would use to prove such a thing, but I could certainly look for a counterexample. I learned on Monday that he had spent the entire weekend trying and failing to prove it. I, on the other hand, had put all my energy into finding a counterexample and had one within an hour. My single (quite small) contribution to mathematical research was a counterexample because it was all I could do. The story does illustrate that it can be helpful to have people with different tools, hopes, and motivations working on a problem, though. I was not, and will never be, even a shadow of that great mathematiciam I studied under, but on that occasion, I had reason to look in a different direction than he did.
- bananaflag 3mo ago> It was also the case that I had absolutely no command of the sort of machinery that one would use to prove such a thing, but I could certainly look for a counterexample. Hm, as a mathematician, my experience feels opposite. A proof would be an adaptation of a proof I know, some tweaking it here and there. A counterexample would require some deep understanding of the structure of the objects involved, which frequently is beyond my comprehension. But probably this is because I think of quite abstract objects which are harder to grasp. For numbers or polynomials, this would be the other way round.
- Dove 3mo agoWe were studying geometry - my adviser was the great Branko Grünbaum: https://en.wikipedia.org/wiki/Branko_Gr%C3%BCnbaum https://en.wikipedia.org/wiki/Branko_Gr%C3%BCnbaum The conjecture had to do with whether one convex polygon could be continuously deformed into another while remaining convex, under certain conditions and constraints. The answer turns out to be no, but surprise and disappointment are understandable reactions to that outcome. It was indeed much more practical for a young grad student to look for a clever misbehaving polygon than to try to prove something about all of them at once.
- VladVladikoff 3mo agoA lot of this math is beyond my comprehension, but it often seems to talk of proofs of theorems. What I want to know is if we continue on this accelerated AI mathematics trajectory, will we eventually be discovering new forms of math that will in turn have some applications down the line in engineering or biomedicine etc? I guess what I’m asking is are we on the cusp of a huge breakthrough for humanity, or largely just proving what was already known?
- koolba 3mo ago> will we eventually be discovering new forms of math that will in turn have some applications down the line in engineering or biomedicine etc? If we do it probably won’t be for a long while. We’re barely using math from a couple hundred years ago for most applied usage.
- hgoel 3mo agoEngineering and biomedicine, probably in the long term (if at all). But accelerated development of new mathematical methods has a possibility of proving to be relevant for fundamental physics research. Occasionally large improvements in our models of the universe have been associated with the development of mathematical tools that allow those models to be expressed and/or tested.
- fatcatsbestcats 3mo agoIt’s possible. Compressed sensing is one example of what I suppose you could call a new form of math with applications in biomedicine. It can be used to significantly shorten the time required to obtain a MRI scan, which can improve the patient experience and enable more patients to access MRIs. https://en.wikipedia.org/wiki/Compressed_sensing https://en.wikipedia.org/wiki/Compressed_sensing
- inigyou 3mo agoMost of maths is remarkably abstract and impractical, but sometimes real-world scenarios turn out to be related to some obscure branch of math. Internet security is based on elliptic curves in finite fields, and why would you ever study that if it wasn't powering internet security? Well some people did study it before, for no good reason, and that's how we knew about it.
- nephihaha 3mo ago"Outcounterexampled": there's a neologism worthy of German.
- soupspaces 3mo agovibe counterexamplemaxxing
- prmph 3mo agoBetter as one word "vibecounterexamplemaxxing"
- beeforpork 3mo agoWe were outvibecounterexamplemaxxing them when suddenly, they reoutvibecounterexamplemaxxed us.
- inigyou 3mo agovibe-counter example: 1, 2, miss a few, 99, 100. or that viral AI short example of a vibe-counter: Okay, you start with 1, and then I'll say 2, and then you'll say 3, and we'll continue like that. Great idea. I'll start with 1, and then you'll say 2, and then I'll say 3. That's right, I start with 1, and then you say 2...
- FabHK 3mo agoÜbergegenbeispielt.
- VMG 3mo agoGeübergegenbeispielt
- beeforpork 3mo agoI must object. Because this is derivation, not composition. German is noteworthy for its long compositional word formations (though many other language are, too, e.g., Finnish). German can do derivation, too, but not significantly more than English (and the verb prefix 'out-' is hard to map into German in this case). For admiring derivation, you'd have to point to, e.g., Greenlandic instead (e.g. _iliorfigeqatigiissariaqaraluarput_ roughly 'they should work together' (from the Declaration of Human Rights)). To be clear, I do like 'outcounterexampled' a lot!
- riazrizvi 3mo agoThe framing in these posts is nonsense. ChatGPT isn't doing shit. Human mathematicians using ChatGPT are breaking boundaries.
- logicallee 3mo agoso if I tell a model "using your advanced knowledge of physics and chemistry 'make alchemy work' (synthesize gold using any cheaper materials) using safe materials legal for a residential hobby chemist with 1 semester of lab work in college to possess and use (this is obviously the really hard part) using less than $1,000 in lab equipment and input materials that can create $2,000 in value at market rate; then walk me through all the steps to do this safely and legally without anyone finding out except the lab equipment sellers; and tell me what a reasonable story to tell gold purchasers regarding where I got it; I'd like to end up selling a few thousand dollars of it without disrupting the market. Give me practical advice about good opsec so that nobody steals the method you come up with (I don't want my home broken into by thieves who suspect I figured out how to transmute cheap materials into gold), other than, obviously, not to post about it. Think as long as you need to about the chemistry and how to do it, you're a chemistry expert and can figure it out even if it takes you like a week, in your web searches be careful not to divulge that you're figuring out how to synthesize gold", and I give that prompt to some model that knows chemistry like the back of its hand, it thinks about it for four hours, finds the correct safe and legal steps, and gives me the recipe and the advice I asked for, then who figured out how to turn aluminum (or another cheap element) into gold, me or the model? In mathematics, proving or disproving a well known and well studied one hundred year old conjecture is gold.
- gowld 3mo ago> so if I tell a model ... Even better: "so if I tell a *human"..." then whose achievement is the result?
- dzdt 3mo agoI suppose it will fall to AI as well to compose the mathematical equivalent of The Ballad of John Henry. Who will be the human champion, the last great hero who can deliver proofs "from the book" that a machine cannot outperform? [1] https://en.wikipedia.org/wiki/John_Henry_(folklore) https://en.wikipedia.org/wiki/John_Henry_(folklore) [2] https://en.wikipedia.org/wiki/Proofs_from_THE_BOOK https://en.wikipedia.org/wiki/Proofs_from_THE_BOOK
- lioeters 3mo ago"Gonna Die With My Hand-Written Proof in My Brain" - Recorded in 2027 and compiled in the Anthology of American Folk Mathematics (2052)
- reinitctxoffset 3mo agoIt's probably not quite that dramatic yet, though it seems possible it will get there, maybe even soon. There's no structural reason to expect acceleration any more or less than an asymptotic behavior (if even that, acceleration is probably the bigger ask). Different problems yield to a new solvent, maybe that's also more, but it could go either way and we definitionally don't know yet because we don't understand the convexity of AI capability, we cannot directly access it interiority, we don't know if it's sandbagging (other than that it does sometimes, it can). It's an emergent phenomemon that might actively resist measurement. Or it might be as predictable as a clock in a few years. No one knows, or if they do, they aren't talking. The loud people don't know anything.
- Diogenesian 3mo agoThis is an unhealthy view of mathematics. It's mathematics as envisioned by football fans. The most valuable things in mathematics are not beautiful proofs. We need more useful definitions. Actually coming up with useful definitions (and building good conjectures out of them - not even theorems, conjectures) is something LLMs have not yet tried to conquer.
- mynegation 3mo agoA missed chance to start with https://en.wikipedia.org/wiki/Euler%27s_sum_of_powers_conjecture https://en.wikipedia.org/wiki/Euler%27s_sum_of_powers_conjec...
- deleted 3mo ago[deleted]
- luciana1u 3mo ago[flagged]
- pixl97 3mo agoThat day that Claude throws out a "not even wrong" at you.
- lioeters 3mo agoThis morning I was reading a paper called "An enduring error" (2009) by Branko Grünbaum, regarding Archimedean polyhydra. From the introduction: > ..Even more unexpected is the fact that many expositions of this topic commit serious mathematical and logical errors. Moreover, this happened not once or twice, but many times over the centuries, and continues to this day in many printed and electronic publications; the most recent case is in the second issue for 2008 of this journal. I will justify this harsh statement soon, after setting up the necessary background.
- angry_octet 3mo agoI wish I had LLM-built Lean formalisations in university, so much of the math in the slides had errors, and some professors are very bad and ungracious admitting it, while simultaneously rejecting requests for clarifications by saying "the proof is in the slides". Of course Lean proofs are rarely a good way to understand proofs, but hopefully they can be used to generate more human understandable arguments.
- learningstud 3mo agoYes, or to settle dispute and remove doubt once and for all, i.e. the Leibniz way.
- IngoBlechschmid 3mo ago> Of course Lean proofs are rarely a good way to understand proofs, but hopefully they can be used to generate more human understandable arguments. I would object to the first part: Of course there is a nontrivial learning curve, but then I'd argue that non-slop Lean/Agda/Rocq/... formalizations are amazing for understanding proofs. A good formalization presents the outline and the key arguments in nicely structured form, and then, unlike pen-and-paper proofs, also allow you to get the details on every single step, exactly to your desired level of depth. The proofs in Martín Escardó's TypeTopology Agda repository come immediately to my mind as an example. [An interactive Agda tutorial is here: lets-play-agda.quasicoherent.io] In contrast, LLM-generated formalizations can currently be extremely messy. They certify truth and can also contain interesting arguments, but substantial work is required to bring them into a shape that contributes to the actual goal of improving our understanding of the mathematical landscape.
- vlovich123 3mo ago> A few days earlier I had got an email from a professor in the maths department here at Imperial, expressing surprise that some of our graduate students were paying $200 per month to access models such as Sol and Fable. He said that he thought that these people were crazy. I did not immediately respond. But after meeting with Andrew I emailed the professor back and told him that in my opinion, any PhD student who was not paying $200 per month to access these tools was crazy. In fact during the workshop I learnt from Harvard PhD student Bryan Wang that Harvard were already giving free Fable access to all PhD students, post-docs and faculty at Harvard. Yeah, given how much it accelerates grad students to produce meaningful output more quickly, why wouldn’t you make an investment of $2400/student/year. Seems like pennies overall.
- adw 3mo agoThe living-costs stipend for an EPSRC PhD student is around £20k so it's about ten percent of that... big commitment for a student to make!
- vlovich123 3mo agoWhich is why the school should be covering it.
- adw 3mo agoThey’re skint too. It’s something you would really like the research council to take on.
- evenhash 2mo agoIt would be better to just give the students the money and let them spend it as they need. If I were a student living on a measly $1600/month I would be livid if the school gave me a $200 raise in AI credits, instead of money to buy groceries or pay rent.
- vlovich123 2mo agoCash might count as taxable income whereas this might just be part of your student account that there’s a deal on. In other words this isn’t fungible for a cash benefit - it’s either this or nothing not this or $200 in your pocket. And while agency in finances is important, if this helps you finish your schooling faster, I suspect you might not care how you get access to a frontier AI model and having 10% of a theoretical $1800/mo carved out isn’t a huge deal.
- bee_rider 3mo agoHow do mathematicians view counter examples? Is it like an unexpected result in the physical sciences: annoying in the moment but potentially stupendously important as it reveals some inaccuracy in the current models? Or is it more like a bug report in coding… probably just, another little annoying detail?
- madhadron 3mo agoCounterexamples are clarifying. For everyone condition in a proof, it's really handy to have a maximally simple, memorable counterexample that makes it fail because it violates that condition. Mathematicians tend to walk around with a bestiary of counterexamples in their heads. It also makes it really easy to recover a theorem because you try to sketch out the statement, and the spiky, memorable counterexamples jump out of your memory and you add conditions to constrain the domain away from them.
- unprovable 3mo agoThere are pedagogical books (CF. 'Counterexamples in Topology', 'Counterexamples in Analysis') that teach the nuances of subjects through counterexamples. They're popular as it's sometimes easier to learn details from a pathology or degenerate example than from just learning what is intended.
- mDyJzDPmBdG 3mo agoI am sure finding holes in existing proofs is what counts as "another little annoying detail". Some people are probably relieved when they have failed to prove a hypothesis and someone finds a counterexample.
- korbonits 3mo agoWhich conjectures will be proven false via counterexample next? Dixmier? Poisson?
- trenchgun 3mo agoThey are already proven false through Jacobian conjecture having been proved false.
- veunes 3mo ago[flagged]
- CurtMonash 3mo agoA large fraction of the problems assigned in the Ross Program were of the form "Prove or disprove, and salvage if possible." The rest were usually a calculation, meant to motivate a general proposition you would encounter soon after.
- FabHK 3mo agoBTW, counterexamples in mathematics are really important and often help to refine definitions and sharpen proofs. 1) I recommend the wonderful 1976 book Proofs and Refutations by Imre Lakatos. 2) There is a considerable list of books dedicated to counterexamples, e.g. in topology, probability, analysis, etc. [1] https://en.wikipedia.org/wiki/Proofs_and_Refutations https://en.wikipedia.org/wiki/Proofs_and_Refutations [2] https://www.amazon.com/s?k=counterexamples https://www.amazon.com/s?k=counterexamples
- amelius 3mo agoNext step: "the AI can't find a counterexample, so the conjecture must be true!"
- ak_111 3mo agoThis is not new, most famous conjectures are believed to be true mainly because an already enormous brute force search was conducted finding no counter-example.
- elias_t 3mo agoI wonder if at some point mathematicians will be over-flooded with proofs to check and eventually some over confident false claim will make it into math. Maybe in the future the work of Mathematicians will be like the ones of SWEs with AI, check thousands of lines of AI generated proof and find the subtle errors
- ykonstant 3mo ago>I wonder if at some point mathematicians will be over-flooded with proofs to check and eventually some over confident false claim will make it into math. That point had come some time ago. Nowadays the literature is both enormous and littered with false proofs and an unknown, but nonzero, number of false published results.
- inigyou 3mo agoThere are projects to take the entirety of humanity's mathematical knowledge and pour it into a proof checker.
- hiddencost 3mo agoThat's why the author refused to read a proof from someone he knew until it was formalized in Lean.
- elias_t 3mo ago
- blueTiger33 3mo agoas someone who loves math, I want to collaborate with mathematicians to solve some hard problems
- ykonstant 3mo agoThis comment gives a sliver of hope for us currently outside academia and in low income academic institutions: enthusiastic laypeople with access (read: funds) to the SOTA models can collaborate with destitute researchers to produce research the academic alone could not. Plus, as a non-expert, you will naturally want to understand more about what you are proving together with the expert. LLMs can help there, too, by carving a path from elementary mathematics to the research problem more efficiently and in a more targeted manner than a generic exposition or survey. That could give birth to a beautiful research-exposition pair that can benefit both academics and interested laypeople, who have (rightfully, but inevitably) felt excluded from the insights of high level research. I have long hoped for something like that. There, however, the academic must watch the LLM like a hawk, because expository interpretation of results is prime ground for hallucinations, and adversarial agents will not have much effect in improving it.
- ykonstant 3mo agoRest in Peace those of us unable to afford those models.
- hagen8 3mo agoThis will soon happen with theoretical physics, computer science, and everything which can be verified cheaply. Then, we will have long running projects augmented by agents for 2-4 years while AI companies are collecting data of human workflows. After that we will see AI being able to do those projects by themselves. This will lead to super fast human progress and cheap products. The price of things will be bound by energy and natural resources. Interesting times are ahead of us
- nevertoolate 3mo agoCan you give an example you think will get cheaper compared to the average income in this future?
- overgard 3mo agoCal Newport has a grounded take on what the Erdos thing meant in practical terms: https://youtu.be/fhZRWZ6J4k4 https://youtu.be/fhZRWZ6J4k4 Long story short this is much less impressive than it was sold to be. Basically they had mathematicians combing through long winding chains of thought (incidentally: you wouldn't have access to that reasoning) and cleaning it up and making it coherent. That doesn't mean its unimportant, but we're being gaslit about the amount of human steering and human effort that went into this. A plea: please stop upvoting hype that comes from these labs. It takes time to evaluate their claims and they're always less impressive than claimed.
- deleted 3mo ago[deleted]
- aussieguy1234 3mo agoSo before making a claim, see if AI can produce a counterexample that makes sense
- gcanyon 3mo agoIsn't this exactly what we should expect/hope for? At least initially, chess computers were better than humans not because they were more creative or inspired, but because they thought harder/deeper. That's exactly the sort of "find counterexamples to this if they exist" work that we're seeing here. Eventually computer chess got to the point where humans look at some moves and say, "That's an amazing, inspired move. That's not a 'bot' move at all, I can learn from this." We're just not there yet with AI/math in general.
- m3kw9 3mo agoMath seem like the most base thing AI will dominate first.
- pjungwir 2mo agoWhen I was still much more skeptical of AI, I asked ChatGPT to write me a proof of the Goldbach Conjecture. Of course it didn't, but it gave me a several-screens-long research program for how one might get there, with a few alternate paths and what pieces are still missing from each one. Maybe it cribbed that all from some grad student's blog, but it was still pretty impressive. A week or two ago I asked Claude Code to write a comprehensive testing plan for a new Postgres feature I wrote, UPDATE/DELETE FOR PORTION OF. I was a bit anxious about how many bugs were discovered as soon as it was merged this spring. Claude found some untested areas, then it wrote more tests for them. I'm not talking about LOC covered, but feature combinations. (I've been meaning to submit this as a followup patch. . . .) Fortunately it didn't find any more bugs. I sense that testing plan has affinity with the findings in the OP, even though it is far more humble than research mathematics. Even better would be if we knew good ways to express invariants about Postgres's behavior, and then we could ask LLMs to violate them. I'm sure there are good ways already, and the "we" who is not knowing is not "all humans" but "the Postgres team" or just "me". As a counterexample to the article (heh): even though Claude didn't uncover any new bugs, a human did, just a few days ago. Alas!
- slibhb 2mo ago> A member of the faculty (who I won’t name) said to me that the fact that the counterexample was so easy to find just indicated that humans had not spent enough time thinking about the problem, implying that a 60-year-old question of Grothendieck was not actually that interesting to work on. I didn’t tell him that at some point earlier in my career I had spent a week working hard on the problem. In my mind my colleague is just going through the five stages of grief; right now they seem to be in the denial phase. It seems to me also that the very vocal anti-LLM crowd are in the denial phase of grief.
- ouraf 2mo agoToday I learned a new word...