4 ms·
> You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding a
by saithound 1mo ago
> You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.
If you have time, can you elaborate or give some examples of mathematical nuances?
I am evaluating Sol and Fable on a fairly large dataset of subtly flawed informal mathematical arguments (task is to identify and name propositions with substantially incorrect proofs in a larger body of text), and Sol is saturating the benchmark, while Fable is below 50% even with the most generous grading.
I don't work in the natural sciences, so I suspect you mean something different by "mathematical nuance".
- overdrive110 1mo agoIs this benchmark public? Anecdotally I have had decent results asking Sol to nitpick my proofs (mostly probability theory but nothing super dense). I have never tried Claude seriously, so I am very curious about what the failures look like with Fable.
- anmolkabra 1mo agoYes! All tasks are on github: https://github.com/harbor-framework/terminal-bench-science https://github.com/harbor-framework/terminal-bench-science. They were contributed through PRs so the discussion and reviewing (before tasks were accepted) is also fully public.