4 ms·
I used this a lot over the weekend, and it's a really intelligent and strange model. It gets really agentic at the higher reasoning levels. It does the basics
by x313 2mo ago
I used this a lot over the weekend, and it's a really intelligent and strange model.
It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than that, it gets obsessed with solving problems and will do insane/unusual things to get to the solution. It actually reminds me of GPT-5.6-Sol-max which is similarly obsessive.
It doesn't surprise me at all that it outscores Opus 4.6. Opus had way better world knowledge but was more "human" with agent stuff - sort of lazy and uncreative, basically giving up once the obvious solutions failed. These newer models work more like magic, they are so creative and persistent at finding ways to get to the solution.
- culi 2mo agoI have the same reaction reading the internal "thinking" monologues of Kimi K3. When I sent a message that was basically "Nope, I'll just do XYZ instead. Thanks for your help", Kimi basically had an identity crisis. Like there was two wolves inside. One that deeply wanted to help more and go above and beyond and one that was trying to tame the other and make a graceful exit. Here's an excerpt of it > Should I verify their README changes? They didn't ask me to. "I've added some notes in the README. Thanks" — that's a closing statement, not a request. Reading the README unprompted to check their notes could be seen as helpful diligence, but they didn't ask for review. Keep it simple: acknowledge, brief close.
- Havoc 2mo agoYou see the same in GLM models. Even slight ambiguity in user instructions will send it into a tailspin on what intention was in thinking tokens then it goes let’s just make a judgement call on a direction and then proceed
- graceful6800 2mo agoObsessive is the right word. Over the weekend I had to stop it multiple times deep into a multi-hour long turn to ask what the hell it was doing. It was like a dog with a bone and would NOT let go of its current work to talk to me. I had to interrupt it three times with increasingly aggressive instructions to STOP and answer my questions before proceeding. In another session it straight up told me it was in the middle of debugging something important and to ask later. I'm running an RTX 6000 Blackwell. It regularly spent over an hour per turn thinking. Every time I looked at it, the thinking trace seemed coherent, sensible, appropriate. But it could never settle on a solution. Granted, I was trying to have it solve a hard problem that 5.6 Sol couldn't solve, but still. Either way, I'm still impressed. It genuinely feels better than Sonnet 5
- tandr 2mo ago> Granted, I was trying to have it solve a hard problem that 5.6 Sol couldn't solve, but still. Did it solve it?
- MasanskY01 2mo agoThe suspense is killing me!
- fodkodrasz 2mo agoIt is still working on it!
- a96 2mo agoAnd must not be disturbed!
- fatata123 2mo ago[dead]
- celrod 2mo agoI think I'd rather have the model stop once the obvious solutions failed and ask me. It can suggest more creative ideas, but I don't necessarily want it to try implementing them.