4 ms·
Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specifi
by johnnyApplePRNG 1mo ago
Not surprised to see Claude significantly higher in scientific intelligence than Sol.
You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.
That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length.
Don't get me wrong, codex is great at finding bugs and building games. It's great.
- saithound 1mo ago> You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. If you have time, can you elaborate or give some examples of mathematical nuances? I am evaluating Sol and Fable on a fairly large dataset of subtly flawed informal mathematical arguments (task is to identify and name propositions with substantially incorrect proofs in a larger body of text), and Sol is saturating the benchmark, while Fable is below 50% even with the most generous grading. I don't work in the natural sciences, so I suspect you mean something different by "mathematical nuance".
- overdrive110 1mo agoIs this benchmark public? Anecdotally I have had decent results asking Sol to nitpick my proofs (mostly probability theory but nothing super dense). I have never tried Claude seriously, so I am very curious about what the failures look like with Fable.
- anmolkabra 1mo agoYes! All tasks are on github: https://github.com/harbor-framework/terminal-bench-science https://github.com/harbor-framework/terminal-bench-science. They were contributed through PRs so the discussion and reviewing (before tasks were accepted) is also fully public.
- deleted 1mo ago[deleted]
- WithinReason 1mo agoSo why is Sol the one solving Erdős problems?
- jvwww 1mo agoYeah exactly - found the parent comment quite funny. For instance, one of the other top comments berates Claude for terrible instruction following with regards to scientific papers, whereas this one is full of praise. Everyone is just making up their thoughts on these models based on vibes.