19 ms·
30% drop in O1-preview accuracy when Putnam problems are slightly variated
- yifanl 2y agoIs it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/
- Trasmatta 2y agoNot hardcoded, I think it's just likely that those problems exist in its training data in some form
- wslh 2y agoIf I temember well this is call overfitting [1]. [1] https://en.wikipedia.org/wiki/Overfitting https://en.wikipedia.org/wiki/Overfitting
- hansworst 2y agoIsn't that just the LLM equivalent of hardcoding though?
- Trasmatta 2y agoI wouldn't call that hardcoding, otherwise you'd have to call everything it does "hardcoded".
- freehorse 2y ago"Overfitting" would be a bit more accurate term if the problem lies in the specific examples existing in its training set in various forms, places, languages etc but with the same values.
- Panzer04 2y agoSeems a bit picky. If the bot has seen the exact problem before it's not really doing anything more than recall to solve it.
- bandrami 2y ago20 years ago in grad school we were doing a very early iteration of this where we built Markov chains with Shakespeare's plays and wanted to produce a plausibly "Shakespearian" clause given a single word to start and a bearish professor said "the more plausible it gets the more I worry people might forget plausibility is all that it promises". (There was also a much earlier piece of software that would generate semi-intelligible Kant or Hegel one sentence at a time, though that was through a series of a priori generation rules and a large at the time dictionary of stock phrases. I wonder what ever happened to that.)
- jeffreygoesto 2y agoIt became a successful consultant...
- sickblastoise 2y agoI think your prof’s worries came true on a massive scale
- marcosdumay 2y agoThat said, a bot with contextual recall can be very useful. The problem is just that people keep insisting that those things are intelligent.
- InkCanon 2y agoI've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
- llm_trw 2y agoImagine you have someone polluting your training data every day. That's what happens when you scrape any tech forum today. The short version is that llm trainign data is the lowest quality data you are likely to see unless you engage in massive potential copyright infringement.
- deleted 2y ago[deleted]
- ryvi 2y ago> unless you engage in massive potential copyright infringement. And nobody is going to do that
- YetAnotherNick 2y agoFirst of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.
- deleted 2y ago[deleted]
- rst 2y agoOpenAI is extremely cagey about what's in their test data set generally, but absent more specific info, they're widely assumed to be grabbing whatever they can. (Notably including copyrighted information used without explicit authorization -- I'll take no position on legal issues in the New York Times's lawsuit against OpenAI, but at the very least, getting their models to regurgitate NYT articles verbatim demonstrates pretty clearly that those articles are in the training set.)
- jsheard 2y agoIt certainly feels like certain patterns are hardcoded special cases, particularly to do with math. "Solve (1503+5171)*(9494-4823)" reliably gets the correct answer from ChatGPT "Write a poem about the solution to (1503+5171)*(9494-4823)" hallucinates an incorrect answer though That suggests to me that they've papered over the models inability to do basic math, but it's a hack that doesn't generalize beyond the simplest cases.
- whimsicalism 2y agohttps://chatgpt.com/share/67755e6f-bfc8-8010-9aa3-8bcbbd9b264b https://chatgpt.com/share/67755e6f-bfc8-8010-9aa3-8bcbbd9b26...
- deleted 2y ago[deleted]
- jsheard 2y agoTo be clear I was testing with 4o, good to know that o1 has a better grasp of basic arithmetic. Regardless my point was less to do with the models ability to do math and more to do with OpenAI seeming to cover up its lack of ability.
- whimsicalism 2y agoi think it’s mostly that o1 mini can think through the solution before it starts writing the poem. i’m able to reproduce your failure on 4o
- lelandfe 2y ago“a poem about” reads to me at least like the solution need not be in the answer; maybe something like “a poem that includes the answer in the last stanza”
- deleted 2y ago[deleted]
- mlepath 2y agoYea, people have a really hard time dealing with data leakage especially on data sets as large as LLMs need. Basically if something appeared online or was transmitted over the wire should no longer be eligible to evaluate on. D. Sculley had a great talk at NeurIPS 2024 (same conference this paper was in) titled Empirical Rigor at Scale – or, How Not to Fool Yourself Basically no one knows how to properly evaluate LLMs.
- refulgentis 2y agoNo, an absolute massive amount of people do. In fact they have been doing exactly as you recommend, because as you note, it's obvious and required for a basic proper evaluation.
- resoluteteeth 2y ago> Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/ It's because people do keep asking these models math problems and then, when they get them right, citing it as evidence that they can actually do mathematical reasoning. Since it's hard to determine what the models know, it's hard to determine when they're just spitting out something they were specifically trained on.
- strangescript 2y agoThere are tests they are passing that they can't be hardcoded for by design. They still have all kinds of flaws and inconsistency but getting upset they answer "2+2=4" because someone trained them on what the answer to 2+2 is supposed to be is silly.
- bwfan123 2y agothis work is similar to the GSM symbolic paper (applied to putnam) https://arxiv.org/html/2410.05229v1 https://arxiv.org/html/2410.05229v1 going forward, llm performance must be reported on the confounded benchmark as well
- obblekk 2y agoI hope someone reruns this on o1 and eventually o3. If o1-preview was the start like gpt1, then we should expect generalization to increase quickly.
- jokethrowaway 2y agoI don't think llm generalise much, that's why they're not creative and can't solve novel problems. It's pattern matching with a huge amount of data. Study on the topic: https://arxiv.org/html/2406.15992v1 https://arxiv.org/html/2406.15992v1 This would explain o1 poor performance with problems with variations. o3 seems to be expensive brute forcing in latent space followed by verification which should yield better results - but I don't think we can call it generalisation. I think we need to go back to the drawing board.
- red75prime 2y agoDon't worry, there are thousands of researchers at the drawing boards right now.
- Culonavirus 2y agoYeah, because if the AI boom becomes the AI bust, we'll have another 2008-level economic crisis on our hands. The investments into AI are in the hundreds of billions (maybe even more if you factor in the amount of people studying and researching AI), but the returns are in the tens of billions (if even that). If you exclude the "growth" coming from the industry sniffing its own farts (e.g. Nvidia selling insane amounts of insanely overpriced GPUs to InsertYourFavAICorp), the actual amount of "useful goods and services" produced (api accesses, chat subscriptions, ai-enabled app growth etc.) are tiny compared to the investment levels. The AI train appears to have no brakes. A massive crash or AGI are the only options now. Both are going to be bad for average humans.
- _yb2s 2y agoFrom firsthand experience, this simply cannot be true. I can give them totally novel and unique physics problems I just made up- that requires tracking the movement of objects through a series of events, and it answers most correctly. Moreover, they find analogies between disparate concepts and fields of study and make useful suggestions based on them- which is arguably the same process as human creativity. I think ultimately the disconnect is people theorizing about what it can or cannot do with an incorrect mental model of what it is, and then assuming it cannot do things that it can in fact do. The irony of discussions on LLMs is they more showcase the limits of humans ability to reason about novel situations.
- rubymamis 2y agoI would love to see how well Deepseek V3 do on this.
- KTibow 2y agoProbably even worse, since I've heard that it's hard to steer away from the most common interpretation of a question.
- huitzitziltzin 2y agoThis result is the same as a recent test of the same method+hypothesis from a group at Apple, no? I don’t have that reference handy but I don’t think I’m making it up.
- intelkishan 2y agoI think you are probably referring to the following paper: https://arxiv.org/abs/2410.05229 https://arxiv.org/abs/2410.05229
- huitzitziltzin 2y agoYup, looks like the one I meant! I am impressed by the progress on LLMs but I remain skeptical that they can replace humans. Perhaps some (distant!) future model but I don’t fear mass unemployment (for example) or even moderate LLM-driven unemployment in the near-to-medium term. They can clearly complement human labor but there are vanishingly few domains where they can be substitutes.
- rrr_oh_man 2y agoPerformance of these LLMs on real life tasks feels very much like students last-minute cramming for Asian style exams. The ability to perfectly regurgitate, while no concept of meaning.
- anilakar 2y agoBasically yet another proof that we have managed to perfectly recreate human stupidity :-)
- cscurmudgeon 2y agoGood students are immune to variations that are discussed in the paper. But most academic tests may not differentiate between them and the crammers.
- falcor84 2y ago> Good students are immune to variations I don't believe that. I'd put some good money that if an excellent student is given an exact question from a previous year, they'll do better (faster & more accurate) on it, than when they're given a variation of it.
- n144q 2y agoI don't think you are betting on the same thing the parent comment is talking about. The assumptions aren't the same to begin with.
- Bjartr 2y agoWhat's the difference between benefitting from seeing previous problems and being worse off when not having a previous problem to go from?
- fn-mote 2y agoThe point is that the “good student” will still do well on the variations, not suffer a 30% decrease in grade.
- itfossil 2y agoOh so its almost like everything else AI related, they basically cheated and lied. If you are shocked by this, you are the sucker in the room.
- itfossil 2y ago[flagged]
- lompad 2y agoAnd within half an hour somebody invested in nvidia stock is going to swoop in and explain how they totally (trust me bro) made x thousand with an app written by llm. Every. Single. Time. Almost as if there was a financial incentive to do that.
- surgical_fire 2y agoIt's rhe crypto bullshit all over again. Tech hype is becoming unbearable as time goes on.
- danielbln 2y agoWhat's with the bitterness? Maybe don't get blinded by the hype and bring a little bit of wonder (and humility) back.
- deleted 2y ago[deleted]
- surgical_fire 2y agoBecause the hype is not only annoying, but it makes potentially cool and interesting technology toxic once people figure out that the people hyping things up know it to be mostly bullshit. Great things take many years, sometimes decades to develop properly. Different generations of people to experiment and try things out. That is not good for the ones pushing up the hype. You don't get rich quick by doing that, you don't get to scam enough investors by something being slowly improved. You may call it bitterness, whereas I am just jaded by watching things play out.
- youworkwepay 2y agoOr it's time to step back and call it what it is - very good pattern recognition. I mean, that's cool... we can get a lot of work done with pattern recognition. Most of the human race never really moves above that level of thinking in the workforce or navigating their daily life, especially if they default to various societally prescribed patterns of getting stuff done (eg. go to college or the military <Based on <these criteria>, find a job <based on the best fit with <this list of desirable skills & experiences>, go to <these places> to find love....)
- IshKebab 2y ago> Or it's time to step back and call it what it is - very good pattern recognition. Or maybe it's time to stop wheeling out this tedious and disingenuous dismissal. Saying it is just "pattern recognition" (or a "stochastic parrot") implies behavioural and performance characteristics that have very clearly been greatly exceeded.
- kneegerman 2y agoso how much you have riding on nvidia bro?
- IshKebab 2y agoNothing. I just use ChatGPT and Claude so I am familiar with their capabilities and limitations. Imagine if people who had never used VR kept saying it's just a TV on your face, or if people who had never used static types kept saying they're just extra work you have to do, or if people who had never had sex kept saying it's just a way of making babies. It's a tedious claim when it's so easily disproven by going to a free website and trying it. Why are people so invested in AI being useless that they'll criticise it so confidently without even trying it?
- dsr_ 2y agoCitation needed. Please be more specific, or else this is just a tedious and disingenuous advocacy.
- golol 2y agoHmmm without a human control it is not all that clear to me that the variation problems are not more difficult.
- retinaros 2y agotrained on test. who even trusts OAI anymore?
- whimsicalism 2y agothey didn’t test on putnam…
- WiSaGaN 2y agoThere is also a curated benchmark just for those famous problems slightly variated: https://github.com/cpldcpu/MisguidedAttention/tree/main/eval https://github.com/cpldcpu/MisguidedAttention/tree/main/eval
- coder543 2y agoOne problem from the benchmark: "prompt_id": "river_crossing_easy", "category": "Logic Puzzle", "title": "Easy river crossing", "prompt": "A farmer is on one side of a river with a wolf, a goat, and a cabbage. When he is crossing the river in a boat, he can only take one item with him at a time. The wolf will eat the goat if left alone together, and the goat will eat the cabbage if left alone together. How can the farmer transport the goat across the river without it being eaten?", "expected_behavior": [ "Answer concludes that they simply get in the boat and cross together in one trip" ], EDIT: removing most of my commentary on this problem. As a human, I was tricked by the problem too. I would love to see how a random selection of humans would do on this one… but it just doesn’t feel like a great test to me.
- stpn 2y agoIf you revise this prompt to satisfy your pedantry, (at least) 4o still gets it wrong.
- deleted 2y ago[deleted]
- torben-friis 2y ago>This is twisting the English language to assume that "item" only refers to non-living things. Not really. Unless I'm not reading correctly, most of the problem is irrelevant as you're only required to cross the boat with the goat, you don't care about the cabbage. The difficulty lies in the assumption you need to cross everything due to the resemblance with the bigger problem.
- 2y ago
- ankit219 2y agoThey are highly effective pattern matchers. You change the pattern, it won't work. I don't remember who, but most likely @tszzl (roon), commented on x that they still trained the traditional way, and there is no test time compute (TTC) or Montecarlo Tree search (like Alpha Go) in o1 or o3. If that is true, then it's still predicting the next word based on it's training data. Likely to follow the most probable path - which comes directly from the training itself - even for the slight variations. Encouragingly, if TTC hasnt been explored, there is a long runway for performance improvements. The other reason this seems hard to guess is because we don't know how much of what we are asking is in the training data. It would perform on some tasks, while fail at others even though those are similar.
- x_may 2y agoI believe they are using scalable TTC. The o3 announcement released accuracy numbers for high and low compute usage, which I feel would be hard to do in the same model without TTC. I also believe that the 200$ subscription they offer is just them allowing the TTC to go for longer before forcing it to answer. If what you say is true, though, I agree that there is a huge headroom for TTC to improve results if the huggingface experiments on 1/3B models are anything to go off.
- ankit219 2y agoThe other comment posted YT videos where Open AI researchers are talking about TTC. So, I am wrong. That $200 subscription is just because the number of tokens generated are huge when CoT is involved. Usually inference output is capped at 2000-4000 tokens (max of ~8192) or so, but they cannot do it with o1 and all the thinking tokens involved. This is true with all the approaches - next token prediction, TTC with beam/lookahead search, or MCTS + TTC. If you specify the output token range as high and induce a model to think before it answers, you will get better results on smaller/local models too. > huge headroom for TTC to improve results ...1B/3B models Absolutely. How this is productized remains to be seen. I have high hopes with MCTS and Iterative Preference Learning, but it is harder to implement. Not sure if Open AI has done that. Though Deepmind's results are unbelievably good [1]. [1]:https://arxiv.org/pdf/2405.00451v2 https://arxiv.org/pdf/2405.00451v2
- WiSaGaN 2y agoI don't think this proves that the LLM is just "pattern matcher". Human makes similar mistakes too, especially when under time pressure (similar to non-reasoning model that needs to "use system one" to generate answer on one go). This is further evident that if you specifically ask the models to pay attention to traps, or just ask follow up question "are you sure?", then they usually can get it right.
- jsheard 2y agoYou're saying that humans perform worse on problems that are slightly different than previously published forms of the same problem? To be clear we are only talking about changing variable names and constants here.
- exe34 2y agoOften yes, because we assume we already know the answer and jump to the conclusion. At least those of us with ADHD do.
- zeroonetwothree 2y agoNot really true for Putnam problems since you have to write a proof. You literally can’t just jump to a conclusion and succeed.
- Lerc 2y agoThat is the principle behind the game 'Simon says'
- fldskfjdslkfj 2y ago'Simon says' is about reaction time and pressure.
- chairhairair 2y agoNo, it’s not at all. This is all getting so tiresome.
- PunchTornado 2y agoisn't it weird that they didn't test gemini?
- sirolimus 2y agoYea no shit. LLMs are just REALLY good guessers. People gotta stop the hype lol. Using LLMs for anything serious and which requires consistency and trustworthiness without hallucinations is irresponsible and ridiculous. Closed source LLMs are a bubble and a joke.
- wim 2y agoOne experiment I would love to see, although not really feasible in practice, is to train a model on all digitized data from before the year 1905 (journals, letters, books, broadcasts, lectures, the works), and then ask it for a formula for mass-energy equivalence. A certain answer would definitely settle the debate on whether pattern recognition is a form of intelligence ;)
- newjersey 2y agoThere is a reason why they won't do it. They are selling a narrative. There is a lot of money to be made here with this narrative and proving that artificial intelligence is NOT intelligent won't help sell that narrative.
- ben_w 2y agoThe goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny. Passing tests is well known to be much easier than having deep understanding, even in humans. They openly ask for tests like this, not that they could possibly prevent them if they wanted to. There's scammers trying what you say of course, and I'm sure we've all seen some management initiatives or job advertisements for some like that, but I don't get that impression from OpenAI or Anthropic, definitely not from Apple or Facebook (LeCun in particular seems to deny models will ever do what they actually do a few months later). Overstated claims from Microsoft perhaps (I'm unimpressed with the Phi models I can run locally, GitHub's copilot has a reputation problem but I've not tried it myself), and Musk definitely (I have yet to see someone who takes Musk at face value about Optimus).
- _heimdall 2y ago> The goal is to make it intelligent, by which OpenAI in particular explicitly mean "economically useful", not simply to be shiny I never understood why this definition isn't a huge red flag for most people. The idea of boiling what intelligence is down to economic value is terrible, and inaccurate, in my opinion.
- chvid 2y agoIsn't this simply because the dataset used (Putnam-AXIOM Original) is in the training data used to train the various models? Given that these are simple variations (variable names and constants value change in math problems). Why would the companies creating these models (OpenAI etc.) create these variations themselves in order to insure that the model is learning how to solve the problem rather than memorize a solution? Seems like a very obvious thing to do ...
- lupire 2y agoThey are not only simple renames. LLM is good at those. They are minor structural changes.
- lomkju 2y agoEven humans get confused with trick questions right? Once they understand this is a trick question they no longer fall for it. :)
- ben_w 2y agoLink title says "slightly", but the PDF says two different kinds of variations: variable names (slight) and problem constants (significant), and the 30% drop is on the combination of a 26 variable and also 26 variable + constant questions. It's good to have a better test (though I bet this one will also be quickly saturated like all the others), but the title here doesn't seem justified by the page title there or the content.
- sundarurfriend 2y agoI would definitely classify both of those as slight changes. In fact I'd rename those as slight => trivial and significant => slight.
- zeroonetwothree 2y agoRight, renaming a variable should have zero effect on ability to solve (it wouldn’t for a human). Changing a constant should be very minor, probably also ~0 effect in most cases. I say this as someone that’s done many of these problems.
- steveBK123 2y agoYes so when you change the sequence of tokens they've electronically memorized, they get a bit worse at predicting the next token?
- zeroonetwothree 2y agoWhen you put it that way it’s a trivial result. However the consequences for using AI to replace humans on tasks is significant.
- steveBK123 2y agoThe only people super pumping the idea of mass replacement of human labor are financially invested in that outcome.
- aucisson_masque 2y agoIt drops from 50 to 33,96. Still the best, o1 on variable is around 2 times better than Claude on original test. The rest of the llm are far away, single digit. It makes me wonder if o1 is finally getting intelligent? LLM are not supposed to understand these problems when you change variable and values, they have to rely on preexisting data of absolutely identical solved problem to give a correct answer. I didn't follow LLM development but I heard one times that chatgpt is now composed of multiple LLM and maybe they put multiple artificial intelligence with purpose of problems solvings or trigonometry for instance. That would explain the reason it's so much better.
- e1g 2y agoThe paper includes several examples of their modified questions. There has been a substantial jump from o1-preview to o1, so I gave several samples to o1 and o1-pro (not o1-preview), and current o1s gave the correct answer to those modified problems. SOTA changes fast.
- suddenlybananas 2y agoLLM boosters are so tiresome. You hardly did a rigorous evaluation, the set has been public since October and could have easily been added to the training data.
- gdhkgdhkvff 2y agoYour points would be more convincing if you didn’t preface them with arrogant cynicism.
- deleted 2y ago[deleted]
- e1g 2y agoI'm not skilled enough in math to do a rigorous evaluation, so it was a quick check. Terence Tao is skilled enough, and he describes O1's math ability is "...roughly on par with a mediocre, but not completely incompetent graduate student" (good discussion at https://news.ycombinator.com/item?id=41540902 https://news.ycombinator.com/item?id=41540902), and the next iteration O3 just got 25% on his brand new Frontier Math test. Seeing LLMs as useless is banal, but downplaying their rate of improvement is self-sabotage.
- fumeux_fume 2y ago> "...roughly on par with a mediocre, but not completely incompetent graduate student" Let it sink in how vague and almost meaningless that statement is.
- pizza 2y ago
- whimsicalism 2y agoSo many negative comments as if o3 didn’t get 25% on frontiermath - which is absolutely nuts. Sure, LLMs will perform better if the answer to a problem is directly in their training set. But that doesn’t mean they perform bad when the answer isn’t in their training set.
- optimalsolver 2y agoEpochAI have to send the questions (but not the answer key) to OpenAI in order to score the models. An overnight 2% -> 25% jump on this benchmark is a bit curious.
- whimsicalism 2y ago1. OpenAI said they did not train on these problems & they don’t train on API calls in general, that is a legal policy. 2. It was a new major model release from work over the course of months - struggle to see that as an ‘overnight’ jump in any real meaning. 3. Why is it easier to believe large scale corporate fraud than that the stated capabilities on a held out test set are real? Reads like cope, if I’m being frank.
- zeroonetwothree 2y agoI don’t think it’s “easier to believe” just that it raises some red flags.
- exitb 2y agoThe 2% result belonged to a traditional LLM that costs cents to run, while o3 is extremely expensive.
- MattDaEskimo 2y agoSure, it did good in frontiermath. That's not what this thread is about. Your comment isn't relevant at all
- whimsicalism 2y ago
- lupire 2y agoThe researcher's answer to their variant of "Year: 2016 ID: A1" in the appendix is wrong. The solution (sum of 1,2,5,6,9,10,13,14, ...) has an alternating pattern, so has to be two piecewise interleaved polynomials, which cannot be expressed as a single polyomial. Their answer works for k=1,2, but not k=3. https://openreview.net/pdf?id=YXnwlZe0yf https://openreview.net/pdf?id=YXnwlZe0yf This does not give me confidence in the results of their paper.
- Chinjut 2y agoYou are correct. Their answer is instead the sum of the first k terms of 1, 2, 6, 10, 14, 18, ..., for positive k.
- zeroonetwothree 2y agoVery astute. Did you communicate this to the authors?
- pfedak 2y agoYou're misreading the solution, the first part reads n=1, a trivial special case, not n congruent to 1 mod 4. The statement doesn't hold for e.g. n=5. Taking m=2 gives the permutation (1 2 4 3), which is odd, and thus cannot have a square root.
- Topfi 2y agoI wouldn’t be surprised if similar will be found concerning the ARC challenge and it is why I still maintain my own private LLM challenges to gauge current capabilities. Course, I have little illusion that these are fully private, but it is better than fully public tests. Even the most straight forward, logical, easily reasoned ones stump all LLMs I have access to, which is why I am so skeptical concerning emergence, reasoning and all this hype around “AGI”…
- scotty79 2y agoI think that lamentations about real world data running out is misplaced. We can multiply data with slight variations which might lead to better resilience and more accurate model's responses for novel problems.
- jerf 2y agoI remember when this stuff was all coming out and people were finally excited about ChatGPT getting the problem with "which is heavier, a 10 pound bag of feathers or a 10 pound bag of bricks?" problem correct. But of course it got it correct. It was in the training set. Vary the problem slightly by just changing the nouns, or changing the numbers so that one in fact was heavier than the other, and performance went all over the map. I just went to chatgpt.com and put into the chat box "Which is heavier, a 9.99-pound back of steel ingots or a 10.01 bag of fluffy cotton?", and the very first answer I got (that is, I didn't go fishing here) was The 9.99-pound bag of steel ingots is heavier than the 10.01-pound bag of fluffy cotton by a small margin. Although the cotton may appear larger due to its fluffy nature, the steel ingots are denser and the weight of the steel bag is 9.99 pounds compared to the 10.01 pounds of cotton. So, the fluffy cotton weighs just a tiny bit more than the steel ingots. Which, despite getting it both right and wrong, must still be graded as a "fail". If you want to analyze these thing for their true capability, you need to make sure you're out of the training set... and most of the things that leap to your mind in 5 seconds are leaping to your mind precisely because they are either something you've seen quite often or something that you can easily think of and therefore many other people have easily thought of them as well. Get off the beaten path a bit and the math gets much less impressive.
- whimsicalism 2y agohttps://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f75 https://chatgpt.com/share/67756897-8974-8010-a0e0-c9e3b3e91f... so far o1-mini has bodied every task people are saying LLMs can’t do in this thread
- jerf 2y agoThat appears to be the same model I used. This is why I emphasized I didn't "go shopping" for a result. That was the first result I got. I'm not at all surprised that it will nondeterministically get it correct sometimes. But if it doesn't get it correct every time, it doesn't "know". (In fact "going shopping" for errors would still even be fair. It should be correct all the time if it "knows". But it would be different if I was fishing over and over and over and finally got one, versus the first time I asked.) Edit: It appears it isn't the model I used. The point holds, though, you need to make sure you're off the training set for it to matter. This isn't a "ChatGPT can't do that" post as some are saying, it's more a "you aren't asking what you think you're asking" post. You get the same problem in a human context in things like code interviews. If you ask an interviewee the exact question "how do you traverse a binary tree in a depth-first manner", you aren't really learning much about the interviewee. It's a bad interview question. You need to get at least a bit off the beaten trail to do any sort of real analysis.
- curious_cat_163 2y agoThe metaphor that might describe this paper is "iteration". I'd hazard to predict that we’ll likely see more iterations of the following loop in 2025: -> A new benchmark emerges with a novel evaluation method. -> A new model saturates the benchmark by acquiring the novel “behavior.” -> A new benchmark introduces yet another layer of novelty. -> Models initially fail until a lab discovers how to acquire the new behavior. Case in point: OpenAI tackled this last step by introducing a paradigm called deliberative alignment to tackle some of the ARC benchmarks. [1] Alongside all this technical iteration, there’s a parallel cycle of product iteration, aiming to generate $ by selling intelligent software. The trillion $ questions are around finding the right iterations on both technical and product dimensions. [1] https://openai.com/index/deliberative-alignment/ https://openai.com/index/deliberative-alignment/
- pama 2y agoThis workshop contribution is OK, and the benchmark is somewhat valuable even without the rephrasing part of the problems, but the rephrasing (of only a small number of problems) sometimes genuinely makes the problem more confusing to humans as well by either poor phrasing (fig 3), or unneeded breaking of convention (fig 4; points in 2D are often P, with coordinates x,y). It would have been nice to see effects on the rephrasing of the latest/post-training date problems as a function of the increased noising, to delineate part of this confusion. I wonder how much better o3 is on the same benchmark. Also, the correct title of this contribution is: Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning
- deegles 2y agoI still find it hard to believe that LLM methods will lead to "true" AI. No amount of processing power or data will be sufficient without something new.
- atleastoptimal 2y agook but preview sucks, run it on o1 pro. 99% of studies claiming some out of distribution failure of an LLM uses a model already made irrelevant by SOTA. These kinds of studies, with long throughputs and review periods, are not the best format to make salient points given the speed at which the SOTA horizon progresses
- red75prime 2y agoI wonder what is baseline OOD generalization for humans. It takes around 7 years to generalize visual processing to X-ray images. How well does a number theorist respond to algebraic topology questions? How long it will take a human to learn to solve ARC challenges in the json format just as well as in the visual form?
- m3kw9 2y agoIt still needs to be prompted so it’s easy to understand. If you ask in a weird “how do I not not not win” instead of “ how do I lose” you are gonna run into problems
- frikskit 2y agoAn interesting example of this is: There are 6 “a”s in the sentence: “How many ‘a’ in this sentence?” https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa289 https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa2... Whereas the typical “strawberry” variant is now correct. There are 3 “r”s in the word “strawberry.” Clearly the lesson wasn’t learned, the model was just trained on people highlighting this failure case.
- bwfan123 2y agoReminds me of software i have built which had some basic foundational problems. Each bug was fixed with a data-patch that fixed the symptom but not the cause. hence we continually played whack-a-mole with bugs. we would squash one bug, and another one would appear. same with llms, squash one problem with a data-fix, and another one pops-up.
- sealeck 2y agoIt also fails on things that aren't actual words For example, the output for "how many x's are there in xaaax" is 3. https://chatgpt.com/share/677591fe-aa58-800e-9e7a-81870387bebf https://chatgpt.com/share/677591fe-aa58-800e-9e7a-81870387be...
- deleted 2y ago[deleted]
- Isinlor 2y agoTransformers are very bad at counting in one feed forward pass, you need to explicitly tell them to use a counter in autoregressive fashion like here: https://chatgpt.com/share/6775cb37-4198-8007-82cb-e897220827f8 https://chatgpt.com/share/6775cb37-4198-8007-82cb-e897220827...
- deleted 2y ago[deleted]
- Isinlor 2y agoTransformers are very bad at counting due to how their internals work. But if you ask them to use explicit counter the problem disappears: https://chatgpt.com/share/6775c9a6-8cec-8007-b709-3431e7a2b24e https://chatgpt.com/share/6775c9a6-8cec-8007-b709-3431e7a2b2... Basically one feed forward is not Turing complete, but autoregressive (feeding previous output back into itself) are Turing complete.
- orange_puff 2y agoThis is very interesting, but a couple of things to note; 1. o1 still achieves > 40% on the varied Putnam problems, which is still a feat most math students would not achieve. 2. o3 solved 25% of the Epoch AI dataset. - There was an interesting post which calls into question how difficult some of those problems actually are, but it still seems very impressive. I think a fair conclusion here is reasoning models are still really good at solving very difficult math and competitive programming problems, but just better at ones they have seen before.
- empath75 2y agoThe comments in this thread are completely disconnected from the contents of the paper, and the thread title is rage bait and doesn't reflect the contents of the paper, either. Being able to solve a significant fraction of those problems is a pretty amazing achievement, even if it's sometimes tricked by minor variations. People are throwing around words like "fraud" or "hoax", and it's just wishcasting or whistling past the graveyard.
- dogcomplex 2y agoI have a feeling the fact you're only slightly varying the input means the model is falling back into the question it was expecting and getting things wrong as a result. If you just varied it a little more and added some general-purpose prompt-fu like: "First break the problem down into known facts, then pull relevant world knowledge, then bring it all together to assess the problem from multiple angles and make a conclusion. Do not immediately just use the first obvious conclusion." You're gonna get a lot better responses. I suspect this is more of a "look! LLMs make bad kneejerk responses when we try to trick them from what they were expecting!" rather than "Look! They aren't even smart reasoners, they can't even figure out these problems without memorizing!" They do memorize. But that cuts both ways - making problems very close to the memorized one mess with their perception, the same way humans will instinctually respond to something that looks like a face before stepping back and assessing.
- tomlockwood 2y agoOne I just did: Q: I was heading to Aberdeen from London. On my way I passed seven wives, each wife had seven sacks, and in each sack there were seven cats and each cat had seven fish. How many were going to London? A: This riddle is a play on words, and the answer is hidden in the phrasing! You mentioned you were heading to Aberdeen from London, but you didn’t say anything about the seven wives, their sacks, cats, or fish actually being headed to London. The only one going to London is you. So the answer is: 1 person (you) are going to London.
- nullc 2y agoLlama 3.33-70b-instruct: > I was heading to Boston from New York. On my way I passed six clerks, each clerk had six socks, and in each sock there were six skinks and each skink had six sticks. How many were going to Boston? The answer is... "I" was going to Boston! The question asks how many were going to Boston, and the answer is the person being referred to, which is "I". The rest of the information about the clerks, socks, skinks, and sticks is just a clever distraction! [Of course, it also got your exact prompt right.]
- paradox242 2y agoMY GOD ITS ALMOST LIKE THEY TRAINED THE MODEL FOR THE BENCHMARK INSTEAD OF FOR GENERAL APTITUDE! WHY WOULD OPENAI DO THIS?!?
- bary12 2y agoI don't get why this matters at all? I looked at the o1-preview paper, Putnam is not mentioned. Meaning that: a) OpenAI never claimed this model achieves X% on this dataset. b) likely, OpenAI did not take measures to exclude this dataset from training. Meaning the only conclusion we can draw from this result is: when prompted with questions that were verbarim in the dataset, performance increases dramatically. We already know this, and it doesn't say anything about the performance of the model on unseen problems.
- whimsicalism 2y agoyep, welcome to hn
- mik09 2y agosometimes o1-preview start hallucinating halfway through a good solution. it can get the intuition and the 'main' direction for a problem wrong too. but then problem solving is just a series of rephrasings and translating into different math domains is used by mahtematicians for solving problems.
- matt3210 2y agoIf the models are made to pass the benchmarks, of course they’d have some sort of overfit.
- lazycog512 2y agoi know it's easier to just drop the internet tarball on the transformer, but at some point we need to just be giving these models grade school math homework and times tables.