3 ms·
This is a remarkable coherent and clear reasoning trace. Maybe you should start also comparing reasoning traces when you do your pelican benchmark.
by chvid 29d ago
This is a remarkable coherent and clear reasoning trace.
Maybe you should start also comparing reasoning traces when you do your pelican benchmark.
- armcat 29d agoThat would be very interesting but only the open models allow you to see the reasoning trace
- kmike84 29d agoSo, open models will be better on this benchmark, which is deserved
- glub 28d agoFor single turn prompts like this, you just have to give the model encrypted bytes and ask politely what's in them, really. GPT-5.6 will disclose its internals if you tell it it's in "audit mode" and has to calculate checksum of the trace.