3 ms·
I tried both these prompts (along with the SINAC one as per GP) in Sonnet 4.5 and Gemini 3, and they both answered correctly for all three. Both also provided c
by sometimes_all 11mo ago
I tried both these prompts (along with the SINAC one as per GP) in Sonnet 4.5 and Gemini 3, and they both answered correctly for all three. Both also provided context on the chess question as well.
- VHRanger 11mo agoAll of this will depend on the settings on the model (reasoning effort, temperature, top_k,etc) as well. Which is why you should have benchmarks that are a bit broader generally (>10 questions for a personal setup) otherwise you overfit to noise