4 ms·
Based on the rest of your writing I’m going to assume that the prompt was the problem.
by someguyiguess 2mo ago
Based on the rest of your writing I’m going to assume that the prompt was the problem.
- coldtea 2mo agoHe was perfectly clear in both cases. If a human misunderstood this, they'd be a dumb human.
- fn-mote 2mo agoFor perspective, I agree with the GP. The writing is not perfectly clear. We don’t have enough evidence to know if that was part of the problem.
- tcp_handshaker 2mo agoKeep deluding yourself, unless you work for an LLM provider... "Frontier LLMs Still Struggle with Simple Reasoning Tasks" https://arxiv.org/abs/2507.07313 https://arxiv.org/abs/2507.07313 "General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks" https://arxiv.org/abs/2604.11778 https://arxiv.org/abs/2604.11778 "...General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks..."
- orangecat 2mo agoKeep deluding yourself, unless you work for an LLM provider You're clearly operating in bad faith, but just for the record: the General365 problems are very difficult as you can see from the examples at https://arxiv.org/html/2604.11778v1#A1 https://arxiv.org/html/2604.11778v1#A1. It's actually impressive that Gemini 3 Pro got 62%, and the strongest OpenAI and Anthropic models they tried were GPT-5.1 and Sonnet 4.5.