7 ms·
I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never s
by InkCanon 2y ago
I've always assumed they removed it, because it's such a basic and fundamental part of ML training that you separate your test and train data. And yet I never see any papers even mention if/how they do this. And I wonder if they do, how do they guarantee with high reliability that their massive terabytes of data don't contain the answer.
- llm_trw 2y agoImagine you have someone polluting your training data every day. That's what happens when you scrape any tech forum today. The short version is that llm trainign data is the lowest quality data you are likely to see unless you engage in massive potential copyright infringement.
- deleted 2y ago[deleted]
- ryvi 2y ago> unless you engage in massive potential copyright infringement. And nobody is going to do that
- YetAnotherNick 2y agoFirst of all, Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly. Secondly, removing it from internet data is not 100% accurate. There are translations of the problems and solutions or references and direct match is not enough. MMLU and test set benchmarks show more resilience though in some previous research.
- deleted 2y ago[deleted]
- rst 2y agoOpenAI is extremely cagey about what's in their test data set generally, but absent more specific info, they're widely assumed to be grabbing whatever they can. (Notably including copyrighted information used without explicit authorization -- I'll take no position on legal issues in the New York Times's lawsuit against OpenAI, but at the very least, getting their models to regurgitate NYT articles verbatim demonstrates pretty clearly that those articles are in the training set.)
- fn-mote 2y agoLet’s think about this. > Putnam is not in the test data, at least I haven't seen OpenAI claiming that publicly What exactly is the source of your belief that the Putnam would not be in the test data? Didn’t they train on everything they could get their hands on?
- whimsicalism 2y agodo you understand the difference between test data and train data? just reread this thread of comments
- YetAnotherNick 2y agoI don't know why I and you are getting downvoted. Sometimes, HN crowd is just unhinged against AI.
- boroboro4 2y agoThese models are trained in two steps: training base model and then uptraining it. First step includes as much data as possible, everything company can find. For Llama models it's 15T tokens, which is ~40 TB of data. No-one really puts an effort on splitting this data into train/test/eval (and it's not very achievable either). It's just as much data as possible. So it's like 99.9999999% wrong to assume something public isn't on the train set, such as Putnam problems in this case. This is about it.
- whimsicalism 2y agoright, but where did someone assume it wasn’t in the train set? they just said it wasn’t in the test set
- boroboro4 2y agoWhat test set is being talked about here? Why does it matter what’s on this set?
- whimsicalism 2y agofunny that nobody replying to you seems to even know what a test set is. i always overestimate the depth of ML conversation you can have on HN
- chvid 2y agoIt is on the open internet - questions and suggested solutions: https://kskedlaya.org/putnam-archive/ https://kskedlaya.org/putnam-archive/ I would expect all llms to be trained on it.
- jprete 2y agoI don't see any reason to assume they removed it unless they're very explicit about it. Model publishers have an extremely strong vested interest in beating benchmarks and I expect them to teach to the test if they can get away with it.
- stingraycharles 2y agoAs usual, once a metric becomes a target, it stops being useful.
- franktankbank 2y agoWell, they are doing BigCorpStuff not Science
- whimsicalism 2y agoputnam isn’t an llm benchmark ahhhh none of these companies are reporting putnam scores there’s nothing nefarious about training on putnam problems
- jprete 2y agoAny problem set that can make news is implicitly an LLM benchmark.
- deleted 2y ago[deleted]
- captainbland 2y agoI think it's reasonable to assume that openAI is optimising for maximum hype at this point which may include wilfully overfitting for impactful benchmarks to generate positive reports.
- lupire 2y agoWhen 4 came out they released a document that did BOTH inflate scores by changing the exam conditions, and also bragged about scoring worse than guessing on a multiple choice test.
- whimsicalism 2y agoBut putnam isn’t an official test? I find llm discourse on hn so frustrating
- marcosdumay 2y agoHow could they remove it? Those are well known problems, that people talk about on different contexts. They would have to review their entire training set.
- woopwoop 2y agoI agree that openai is somewhat sketchy about this, but they're sketchy about everything. In the past though they have admitted up front to data contamination (e.g. the original gpt-4 press release did not use big-bench as a benchmark due to data contamination). For the Putnam in particular: this is not a benchmark that they use. There is no reason to exclude it since it is not part of the "test set" in any meaningful sense.