3 ms·
Keep deluding yourself, unless you work for an LLM provider You're clearly operating in bad faith, but just for the record: the General365 problems are very di
by orangecat 2mo ago
Keep deluding yourself, unless you work for an LLM provider
You're clearly operating in bad faith, but just for the record: the General365 problems are very difficult as you can see from the examples at https://arxiv.org/html/2604.11778v1#A1 https://arxiv.org/html/2604.11778v1#A1. It's actually impressive that Gemini 3 Pro got 62%, and the strongest OpenAI and Anthropic models they tried were GPT-5.1 and Sonnet 4.5.