6 ms·
Caltech Mathathon – first hackathon ever devoted to research level mathematics
- deleted 28d ago[deleted]
- charlieyu1 28d agoInteresting, if only I still have energy to work on something 40 hours non-stop
- deleted 28d ago[deleted]
- Semkas 28d agoWon't deny that this is an interesting idea, but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing / educative. More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.
- Aboutplants 28d agoIf the goal is to accomplish something then why limit yourself with available tools? I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
- loloquwowndueo 28d agoIf the goal is to run 42km why limit yourself? Use a car and win.
- falcor84 28d agoBut the goal here is not to run 42km; to stay with the outdoors metaphor, it's more like deciding where and how to set up a bivouac - use whatever tools you have at your disposal to analyze the area you're in, and find the best site to stay in overnight.
- a2ff6eeb0 28d agoIt seems like making progress on math is letting the AI run fully autonomously for a few days, occasionally asking it to keep going. I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.
- charlieyu1 28d agoAre they actually autonomous? I’d say subject knowledge at the prompt stage plays a large part towards getting proper results
- a2ff6eeb0 28d agoWhen Claude made progress on the Riemann conjecture, here are the kind of prompts used: > Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress. And left it for a long time. Jarred isn't a mathematician, he's the maintainer of a janky JavaScript environment. Here's the transcript: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e9548c3f44f.pdf https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...
- hgoel 28d agoPrompting for some of the results was almost the "Computer, do a breakthrough. Make no mistakes." meme. Just someone telling the model to keep trying a couple of times. Unfortunately we don't actually know what kind of prompting was done for the more prominent results.
- charlieyu1 28d agoIt’s definitely not how I work, I’d need to read the model responses and set a direction for the model to go
- deleted 28d ago[deleted]
- Fraterkes 28d agoDid you read the article? This is not a hackathon where you build software, it’s one where you’re trying to get a model to make progress on a frontier math problem. The point is that that activity may not map well onto the shape of a hackathon
- thatseasy 28d ago> why limit yourself with available tools Because the companies that run frontier models are malevolent by every metric. They are destroying the environment, especially those in neighborhoods of low income people. They are empowering their owners who are some of the most deplorable and duplicitous people living. They are destroying personal compute to avoid competition with local models by buying all computer components with “promised money” and forcing their P into AI. They stole the entire creative output of humanity and are trying to sell it back to us. They are only good for giving wealth access to skill while removing from the skilled the ability to access wealth. They are being used to kill in war and for surveillance. Seriously why would you use them? Your use only emboldens them; making you complicit in their nefarious success. I for one, am one who walks away from Omelas.
- Donald 28d agoHave you done any math hacking with sol/astra or fable? It’s more fun than using them for coding. The models are great at the monotony, like constructing a Gröbner-basis, etc. But they’re all still absolutely awful at coming up with new ideas, new proof methods, or new constructive forms. So you spend all your time on coming up with novel hypotheses yourself and handing off the rote work to an agent. It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.
- rookienumbers30 28d ago"A mathematician is a person who can find analogies between theorems; a better mathematician is one who can see analogies between proofs and the best mathematician can notice analogies between theories. One can imagine that the ultimate mathematician is one who can see analogies between analogies." I wonder how models perform on finding analogies between analogies
- falcor84 28d agoBut it's not "waiting on the output of an LLM for 40 hours" any more than a regular hackathon is "waiting for my damn teammates to finish their part for 40 hours". From my experience using agentic coding for hackathons, the best teams are those that coordinate with the AI agents in relatively quick cadence, generally giving it small tasks and steering it often. Teams may want to run some long-running sessions too, especially closer to the deadline, but even then, they'd probably want to run and follow several sessions in parallel, and continuously inspect their work so that they have reasonable confidence that their main efforts will wrap up before the deadline. There is an art to it.
- Ey7NFZ3P0nzAe 28d agoI can't wait to see if the team that fares best is the one that steers often or the one that interferes the least. So far humans failed at those problems. Also IIRC there was a guy that proved a substantial problem 2-3 months ago by basically pasting over and over "keep looking for a solution" or something like that for 2 days with little formal math background.
- groundzeros2015 28d agoAre you sure it’s letting it run and not going back and forth interactively?
- brian-bfz 28d ago1. IMO the hard part isn't prompting. It's selecting the problem and understanding the solution. It'd be especially exciting if a participant formulates their own conjecture, proves it with AI, then generalizes it to a new theory. 2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
- youoy 28d agoShould I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...
- sb10128 28d ago[dead]
- blondie9x 28d agoYeah it's a bit of a tricky situation. It's almost like a bribe in a sense.
- tzs 28d agoProbably many mathematicians want answers to the questions from the page: > This AI advancement raises the following questions: (a) How much can AI speed up the process from ideation to peer-reviewed publication? (b) What is the role of a mathematician when AI can solve conjectures faster? and the big AI companies agreed to sponsor them to find out because it is good publicity for the companies.
- isotypic 28d agoObviously the latter - this fact is betrayed by how the page lists the S^6 complex structure result, which was released as a 100 page barely readable mess (in fact even this might be too charitable), as still "unverified". Clearly a situation labs would like to avoid for future claimed results.
- bwfan123 28d ago> Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go. AI harnessing human labor with the incentive of 2M in free tokens is the way my skeptic eye sees it, or humans being duped as reverse-centaurs.
- chaoxu 28d agoI've applied as a team, hope I get in. Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost. Most generic harness we have now are not made for maximizing reasoning. I've tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.
- amelius 28d agoWill BigAI support this with free access to lots of hardware loaded with frontier models?
- boothby 28d ago> It will be the first hackathon ever devoted to research level mathematics. Well, that's pretty damned ignorant; I was attending William Stein's hackathons on the BSD conjecture and the Sage Math project nearly 2 decades ago.
- viccis 28d agoNo it's not, there are tons of programs in mathematics where you go there, form teams, work a problem as a team that's likely to get a result, and then publish the results from all the groups in the conference proceedings. They are called Research Collaboration Workshops. Very consistent pattern from these tech companies in their mathematics press releases that shows a conspicuous lack of experience in the research math world.
- corinthia 28d agorecent caltech grad here! and know some of the organizers well caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!
- deleted 28d ago[deleted]
- xqcgrek2 28d agoNo self respecting mathematician is going to willingly become a marketing tool for these companies solving neglected and irrelevant puzzles.
- milkshakes 28d agoi predict this perspective will not age well at all
- brian-bfz 28d agoHey guys, I'm one of the organizers. AMA. - We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors. - We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants. - Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html https://mathathonchallenge.com/faq.html
- fred123123 28d agoHey! Any indication on what area of mathematics theses questions are from?
- brian-bfz 28d agoYou pick your own problem! You can even formulate your own conjecture and then prove it. Picking an impactful problem is part of our judging criteria.
- fred123123 28d agoThat sound cool!, do you have hints on the cash prizes the website said something like 2 M ???
- jegutman 28d agoThat’s tokens available I assume for all competitors during the competition.
- maypop 28d agoIs this hackathon only for those with formal math backgrounds? I've seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.
- 28d ago
- ondrejdvorak 28d agoHi check zeta.pukapasoft.xyz am I eligible?
- deeznuttynutz 28d agoThis is so freaking cool. I'm super jealous as an old man.
- nill0 28d agoIf anybody is competing in the contest, this may be helpful as a starting point! Open Problems in Computational Geometry Listed by Erik Demaine, Joseph Mitchell, Joseph O'Rourke in 2024 https://topp.openproblem.net/ https://topp.openproblem.net/ And also the popular list below, which contains some of the frontier problems and undefeated beasts that have remained unsolved for decades, some even for centuries. https://en.wikipedia.org/wiki/List_of_unsolved_problems_in_mathematics https://en.wikipedia.org/wiki/List_of_unsolved_problems_in_m... Caution: Solvability is not guaranteed!
- 6z3uiomgd 27d agobro accept my application already
- jereggy 26d agoWhen will results be shared for who got in?
- jereggy 26d agoWhen will results for who got in be released?
- lamkka 26d agoWhy is the Department of Mathematics of Caltech not involved in the organization and sponsorship of this event?
- math_prof 25d agoI am against this Hackathon because it allows unscrupulous companies like OpenAI and Anthropic to use our data for their publicity. If you almost succeed to solve the problem during the hackathon, but perhaps miss out towards the end, they might steal the ideas from your attempts and eventually claim only they solved it. Also, they do not help one bit in making the proofs easier to understand by revealing intermediate inner workings of their products. Math research is not just about stating a Theorem and giving a proof. It is about understanding the theory so that we can continue exploring various aspects. Such a Hackathon will only promote the solving part. Please stop this nonsense.