3 ms·
Grok 4 sets a new high score on my Extended NYT Connections benchmark (92.4), beating o3-pro (87.3): https://github.com/lechmazur/nyt-connections/ https://githu
by zone411 1y ago
Grok 4 sets a new high score on my Extended NYT Connections benchmark (92.4), beating o3-pro (87.3): https://github.com/lechmazur/nyt-connections/ https://github.com/lechmazur/nyt-connections/.
Grok 4 Heavy is not in the API.
- sebzim4500 1y agoVery impressive, but what do you think the chances are that this was in the training data?
- diggan 1y ago> but what do you think the chances are that this was in the training data? Pulled out of my ass, I'd say a 95% chance. NYT Connections is a fairly popular puzzle, it's been out for more than 2 years, and even if this particular GitHub repository with the prompts and methodology wasn't in the training data, it's almost guaranteed that other information, problems and solutions from NYT Connections is in any of the other datasets.
- simondotau 1y agoIf your definition of cheating is "it was fed the answers during training" then every LLM is surely cheating and the real question is why other LLMs didn't do as well in this benchmark.
- pornel 1y agoYou could get 100% on the benchmark with an SQL query that pulls the answers from the dataset, but it wouldn't mean your SQL query is more capable than LLMs that didn't do as well in this benchmark. We want benchmarks to be representative of performance in general (in novel problems with novel data we don't have answers for), not merely of memorization of this specific dataset.
- simondotau 1y agoMy question, perhaps asked in too oblique of a fashion, was why the other LLMs — surely trained on the answers to Connections puzzles too — didn't do as well on this benchmark. Did the data harvesting vacuums at Google and OpenAI really manage to exclude every reference to Connections solutions posted across the internet? LLM weights are, in a very real sense, lossy compression of the training data. If Grok is scoring better, it speaks to the fidelity of their lossy compression as compared to others.
- pornel 1y agoThere's a difficult balance between letting the model simply memorize inputs, and forcing it to figure out a generalisations. When a model is "lossy" and can't reproduce the data by copying, it's forced to come up with rules to synthesise the answers instead, and this is usually the "intelligent" behavior we want. It should be forced to learn how multiplication works instead of storing every combination of numbers as a fact. Compression is related to intelligence: https://en.wikipedia.org/wiki/Kolmogorov_complexity https://en.wikipedia.org/wiki/Kolmogorov_complexity
- frozenseven 1y agoYou're not answering the question. Grok 4 also performs better on the semi-private evaluation sets for ARC-AGI-1 and ARC-AGI-2. It's across-the-board better.
- emp17344 1y agoIf these things are truly exhibiting general reasoning, why do the same models do significantly worse on ARC-AGI-2, which is practically identical to ARC-AGI-1?
- frozenseven 1y agoIt's not identical. ARC-AGI-2 is more difficult - both for AI and humans. In ARC-AGI-1 you kept track of one (or maybe two) kinds of transformations or patterns. In ARC-AGI-2 you are dealing with at least three, and the transformation interact with one another in more complex ways. Reasoning isn't an on-off switch. It's a hill that needs climbing. The models are getting better at complex and novel tasks.
- Workaccount2 1y agoPeople have this misguided belief that LLMs just do look-ups of data present in their "model corpus", fed in during "training". Which isn't even training at that point its just copying + compressing. Like putting books into a .zip file. This belief leads to the thinking that LLMs can only give correct output if they can match it to data in their "model corpus".
- riku_iki 1y ago> the real question is why other LLMs didn't do as well in this benchmark. they do. There is a cycle for each major model: - release new model(Gemini/ChatGPT/Grock N) which beats all current benchmarks - some new benchmarks created - release new model(Gemini/ChatGPT/Grock N+1) which beats benchmarks from previous step
- frozenseven 1y ago"It also leads when considering only the newest 100 puzzles."
- bigyabai 1y agoBe that as it may, that's not a zero-shot solution.
- bilsbie 1y agoYou raise a good point. It seems like would be trivial to pick out some of the puzzles and remove all the answers from the training data. I wish Ai companies would do this.
- zone411 1y agoThe exact questions are almost certainly not in the training data, since extra words are added to each puzzle, and I don't publish these along with the original words (though there's a slight chance they used my previous API requests for training). To guard against potential training data contamination, I separately calculate the score using only the newest 100 puzzles. Grok 4 still leads.
- dangoodmanUT 1y agoGrok 4 Heavy is not a model, it's just managing multiple instances of grok-4 from what I can tell