4 ms·
I've been playing with 3.5:122b on a GH200 the past few days for rust/react/ts, and while it's clearly sub-Sonnet, with tight descriptions it can get small-medi
by misnome 7mo ago
I've been playing with 3.5:122b on a GH200 the past few days for rust/react/ts, and while it's clearly sub-Sonnet, with tight descriptions it can get small-medium tasks done OK - as well as Sonnet if the scope is small.
The main quirk I've found is that it has a tendency to decide halfway through following my detailed instructions that it would be "simpler" to just... not do what I asked, and I find it has stripped all the preliminary support infrastructure for the new feature out of the code.
- reactordev 7mo agoTurn down the temperature and you’ll see less “simpler” short cuts.
- smokel 7mo agoFor the uninitiated: Interestingly, it is not advisable to take this to the extreme and set temperature to 0. That would seem logical, as the results are then completely deterministic, but it turns out that a suboptimal token may result in a better answer in the long run. Also, allowing for a little bit of noise gives the model room to talk itself out of a suboptimal path.
- LoganDark 7mo agoI like to think of this like tempering the output space. With a temperature of zero, there is only one possible output and it may be completely wrong. With even a low temperature, you drastically increase the chances that the output space contains a correct answer, through containing multiple responses rather than only one. I wonder if determinism will be less harmful to diffusion models because they perform multiple iterations over the response rather than having only a single shot at each position that lacks lookahead. I'm looking forward to finding out and have been playing with a diffusion model locally for a few days.
- reactordev 7mo agoYup. I think of it as how off the rails do you want to explore? For creative things or exploratory reasoning, a temperature of 0.8 lends us to all sorts of excursions down the rabbit hole. However, when coding and needing something precise, a temperature of 0.2 is what I use. If I don’t like the output, I’ll rephrase or add context.
- mejutoco 7mo agoSetting the temperature to zero does not make the llm fully deterministic, although it is close.
- sheepscreek 7mo agoThat sounds awfully similar to what Opus 4.6 does on my tasks sometimes. > Blah blah blah (second guesses its own reasoning half a dozen times then goes). Actually, it would be a simpler to just ... Specifically on Antigravity, I've noticed it doing that trying to "save time" to stay within some artificial deadline. It might have something to do with the system messages and the reinforcement/realignment messages that are interwoven into the context (but never displayed to end-users) to keep the agents on task.
- wood_spirit 7mo agoYeah that happened to me with Claude code opus 4.6 1M for the first time today. I had to check the model hadn’t changed. It was weird. I was imagining that maybe anthropic have a way of deciding how much resource a user actually gets and they had downgraded me suddenly or something.
- e1g 7mo agoClaude Code recently downgraded the default thinking level to “medium”, so it’s worth checking your settings.
- nekitamo 7mo agoThank you. The difference was quite noticeable today.
- wood_spirit 7mo agoThank you thank you you give me hope :) But how do you see the current thinking level and how do you change it? I’ve been clicking around and searching and adding “effortLevel”:”high” to .claude/settings.json but no idea if this actually has any effect etc.
- varshar 7mo agoAs per Anthropic support (for Mac and Linux respectively) - $ echo 'export ANTHROPIC_EFFORT="high"' >> ~/.zshrc source ~/.zshrc $ echo 'export ANTHROPIC_EFFORT="high"' >> ~/.bashrc source ~/.bashrc I prefer settings.json (VSCode) - "claudeCode.environmentVariables": [ { "name": "ANTHROPIC_MODEL", "value": "claude-opus-4-6" }, { "name": "CLAUDE_CODE_EFFORT_LEVEL", "value": "high" } ], ...
- shaan7 7mo ago> that it would be "simpler" to just... not do what I asked That sounds too close to what I feel on some days xD
- storus 7mo ago> to decide halfway through following my detailed instructions that it would be "simpler" to just... not do what I asked That's likely coming from the 3:1 ratio of linear to quadratic attention usage. The latest DeepSeek also suffers from it which the original R1 never exhibited.
- nl 7mo agoThere is no way you can diagnose this like that. Correlation isn't causation and much more likely is a common source of reinforcement training data.
- Aurornis 7mo ago> The main quirk I've found is that it has a tendency to decide halfway through following my detailed instructions that it would be "simpler" to just... not do what I asked, This is my experience with the Qwen3-Next and Qwen3.5 models, too. I can prompt with strict instructions saying "** DO NOT..." and it follows them for a few iterations. Then it has a realization that it would be simpler to just do the thing I told it not to do, which leads it to the dead end I was trying to avoid.
- slices 7mo agoI've seen behavior like that when the model wasn't being served with sufficiently sized context window
- soulofmischief 7mo agoClaude Opus does this constantly for me, no matter how I prompt it or what is in my AGENTS.md, etc. It is the bane of my existence.