6 ms·
Its especially concerning / frustrating because boris’s reply to my bug report on opus being dumber was “we think adaptive thinking isnt working” and then thats
by JamesSwift 6mo ago
Its especially concerning / frustrating because boris’s reply to my bug report on opus being dumber was “we think adaptive thinking isnt working” and then thats the last I heard of it: https://news.ycombinator.com/item?id=47668520 https://news.ycombinator.com/item?id=47668520
Now disabling adaptive thinking plus increasing effort seem to be what has gotten me back to baseline performance but “our internal evals look good“ is not good enough right now for what many others have corroborated seeing
- whateveracct 6mo agoyou're using a proprietary blackbox
- JamesSwift 6mo agoSure, but that blackbox was giving me a lot of value last month.
- whateveracct 6mo agoso it's also a skinner box
- retinaros 6mo agoits a drug. that is how it works. they ration it before the new stuff. seeing legends of programming shilling it pains me the most. so far there are a few decent non insane public people talking about it :Mitchel Hashimoto, Jeremy Howard, Casei Muratori. hell even DHH drank the coolaid while most of his interviews in the past years was how he went away from AWS and reduced the bill from 3 million to 1millions by basically loosing 9s, resiliency and availability. but it seems he is fine with loosing what makes his business work(programming) to a company that sells Overpowered stack overflow slot machines.
- throwaway9980 6mo ago[flagged]
- bloppe 6mo agoI think you're loosing your ability to spell
- retinaros 6mo agonever said he was a looser. just that his take on genAi coding doesnt align with his previous battles for freedom away from Cloud. OAI and Anthropic have a stronger lock in than any cloud infra company. you got everything to loose by giving your knowledge and job to closedAI and anthropic. just look at markets like office suite to understand how the end plays.
- throwaway9980 6mo agoThose jobs are as good as loost already. There's no endgame where knowledge workers keep knowledge working they way they have been knowledge working. Adapt or be a loosing looser forever.
- bloppe 6mo agoIs office suite supposed to be an example of lock-in? I haven't used it since middle school. I've worked at 3 companies and, to the best of my knowledge, not a single person at any of them used office suite. That's not to say we use pen and paper. We just use google docs, or notion, or (my personal favorite) just markdown and possibly LaTeX. I think it's somewhat analogous with models. Sure, you could bind yourself to a bunch of bespoke features, but that's probably a bad idea. Try to make it as easy as possible for yourself to swap out models and even use open-weight models if you ever need to. You will get locked into the technology in general, though, just not a particular vendor's product.
- jibal 6mo agoloser (Didn't you notice being mocked for the spelling error?)
- heurist 6mo agoI work with some 'legends of programming' and they're all excited about it. I am too, though I am not a legend. It really is changing the game as a valid new technology, and it's not just a 'slot machine'. Anthropic is burning their goodwill though with their lack of QA or intentional silent degradation.
- retinaros 6mo agoit is a slot machine. you win a lot if what you do is in the dataset. and yes most of enterprise software is likely in it as it is quite basic CRUD API/WebUI. the winning doesnt change the fact that it is a slot machine and you just need one big loss to end your work. as long as you introduce plans you introduce a push to optimize for cost vs quality. that is what burnt cursor before CC and Codex. They now will be too. Then one day everything will be remote in OAI and Anthropic server. and there won't be a way to tell what is happening behind. Claude Code is already at this level. Showing stuff like "Improvising..." while hiding COT and adding a bunch of features as quick as they can.
- dyauspitr 6mo agoThe fact that they might gimp it in the future doesn’t mean it does offer very real world value right now. If you’re not using an LLM to code, you’re basically a dinosaur now. You’re forcing yourself to walk while everyone else is in a vehicle, and a good vehicle at that that gets you to your destination in one piece.
- retinaros 6mo agoas an overpowered stack overflow machine this is quite good and a huge jump. As a prompt to code generator with yolo mode (the one advertised by those companies) it is alternating between good to trash and every single person that works away from the distribution of the SFT dataset can know this. I understand that this dataset is huge tho and I can see the value in it. I just think in the long term it brings more negatives. If you vibecode CRUD APIs and react/shadcn UIs then I understand it might look amazing.
- dyauspitr 6mo agoYes, definitely CRUDs but also iPhone applications, highly performant financial software (its kdb queries are better than 95% of humans), database structure and querying and embedded systems are other things it’s surprisingly good at. When you take all of those into account there’s very little else left.
- NobleLie 6mo agoThe question is, are you getting value from your setups or not?
- butlike 6mo agoAnd now it isn't. Pray they don't alter the deal any further.
- slopinthebag 6mo agoWhoops haha. Surely that can't be how black boxes normally work right?
- mrandish 6mo agoMe too, but it was obviously wildly unsustainable. I was telling friends at xmas to enjoy all the subsidized and free compute funded by VC dollars while they can because it'll be gone soon. With the fully-loaded cost of even an entry-level 1st year developer over $100k, coding agents are still a good value if they increase that entry-level dev's net usable output by 10%. Even at >$500/mo it's still cheaper than the health care contribution for that employee. And, as of today, even coding-AI-skeptics agree SoTA coding agents can deliver at least 10% greater productivity on average for an entry-level developer (after some adaptation). If we're talking about Jeff Dean/Sanjay Ghemawat-level coders, then opinions vary wildly. Even if coding agents didn't burn astronomical amounts of scarce compute, it was always clear the leading companies would stop incinerating capital buying market share and start pushing costs up to capture the majority of the value being delivered. As a recently retired guy, vibe-coding was a fun casual hobby for a few months but now that the VC-funded party is winding down, I'll just move on to the next hobby on the stack. As the costs-to-actual-value double and then double again, it'll be interesting to see how many of the $25/mo and free-tier usage converts to >$2500/yr long-term customers. I suspect some CFO's spreadsheets are over-optimistic regarding conversion/retention ARPU as price-to-value escalates.
- iterateoften 6mo agoIt’s the official communication that sucks. It’s one thing for the product to be a black box if you can trust the company. But time and time again Boris lies and gaslights about what’s broken, a bug or intentional.
- CodingJeebus 6mo ago> It’s the official communication that sucks. It’s one thing for the product to be a black box if you can trust the company. A company providing a black box offering is telling you very clearly not to place too much trust in them because it's harder to nail them down when they shift the implementation from under one's feet. It's one of my biggest gripes about frontier models: you have no verifiable way to know how the models you're using change from day to day because they very intentionally do not want you to know that. The black box is a feature for them.
- bomewish 6mo agoIf you cared so bad you could make your own evals.
- whateveracct 6mo agoso pay anthropic money to maybe detect when the model is on a down week? lol
- chinathrow 6mo agopaying for - so some form of return is expected.
- whateveracct 6mo agothe issue is the return is amorphous and unstructured there's no contract. you send a bunch of text in (context etc) and it gives you some freeform text out.
- chinathrow 6mo agoSure, but I pay real money both to Antrophic and to JetBrains. I get a shitty in line completion full of random garbage or I get correct predictions. I ask Junie (the JetBrains agent) to do a task and it wanders off in a direction I have no idea why I pay for that.
- gowld 6mo ago> I have no idea why I pay for that. And Claude have no idea why it did that.
- chinathrow 6mo agoExactly, and we feel vindicated when it works but sold when it fails. Something will have to change.
- SyneRyder 6mo ago> Sure, but I pay real money both to Antrophic... I misread that as Atrophic. I hope that doesn't catch on...
- ai_slop_hater 6mo agoThis matches my experience as well, "adaptive thinking" chooses to not think when it should.
- andai 6mo agoI think this might be an unsolved problem. When GPT-5 came out, they had a "router" (classifier?) decide whether to use the thinking model or not. It was terrible. You could upload 30 pages of financial documents and it would decide "yeah this doesn't require reasoning." They improved it a lot but it still makes mistakes constantly. I assume something similar is happening in this case.
- rrvsh 6mo ago[dead]
- nomel 6mo agoIs knowing how hard a problem is, before doing it, solved in humans?
- biglost 6mo agoYes, everyweek when assigning fking points to tasks on jira/s
- arthurcolle 6mo agoAs a unit this is funny, Jira points assigned per second (now possible with parallel tool calling AIs)
- WobblyDev 6mo ago[dead]
- Gareth321 6mo agoI don't think so. If the model used to analyse the complexity is dumb, it won't route correctly. They clearly don't want to start every query using the highest level of intelligence as this could undermine their obvious attempt at resource optimisation. I faced the same issue using Open Router's intelligent routing mechanism. It was terrible, but it had a tendency to prefer the most expensive model. So 98% of all queries ended up being the most expensive model, even for simple queries.
- pkilgore 6mo agoSeconded. After disabling adaptive thinking and using a default higher thinking, I finally got the quality I'm looking for out of Opus 4.6, and I'm pleased with what I see so far in Opus 4.7. Whatever their internal evals say about adaptive thinking, they're measuring the wrong thing.
- hbbio 6mo agoUnless they're measuring capex
- echelon 6mo agoThat's why they put the cute animal in your terminal.
- SV_BubbleTime 6mo agoOk, side topic… but that little bastard cheerfully told me out of no where that I have a mall of without a null check AND a free inside a conditional that might not get called. It didn’t give me a line number or file. I had to go investigate. Finally found what it was talking about. It was wrong. It took me about 20 minutes start to finish. Turned it off and will not be turning it back on.
- darkwater 6mo agoI thought it just emitted tongue-in-cheek comments, not serious analysis. And I use the past tense because I had it enable explicitly and a few days ago it disappeared by itself, didn't touch anything.
- c0wb0yc0d3r 6mo agoThe buddies were Anthropics April fools day stunt. Buddies were removed from a newer version of Claude code. By default Claude code updates automatically.
- 6mo ago
- azrollin 6mo ago[dead]
- Moonye666 6mo ago[flagged]
- rkuska 6mo agoFor 4.7 it is no longer possible to disable adaptive thinking. Which is weird given the comment from Boris followed with silence (and closed github issue). So much for the transparency. > Claude Opus 4.7 (claude-opus-4-7), adaptive thinking is the only supported thinking mode. Thinking is off unless you explicitly set thinking: {type: "adaptive"} in your request; manual thinking: {type: "enabled"} is rejected with a 400 error. https://platform.claude.com/docs/en/build-with-claude/adaptive-thinking https://platform.claude.com/docs/en/build-with-claude/adapti... For my claude code I went with following config: * /effort xhigh (in the terminal cli) - To avoid lazying * "env": {"CLAUDE_CODE_DISABLE_1M_CONTEXT": "1"} (settings.json) - It seems like opus is just worse with larger context * "display": "summarized" (settings.json) - To bring back summaries. * "showThinkingSummaries": true (settings.json) - Should show extended thinking summaries in interactive sessions Freaking wizardry.
- arcanemachiner 6mo agoIt's early days for Opus 4.7, but I will say this: Today, I had a conversation go well into the 200K token range (I think I got up to 275K before ending the session), and the model seemed surprisingly capable, all things beings considered. Particularly when compared to Opus 4.6, which seems to veer into the dumb zone heavily around the 200k mark. It could have just been a one-off, but I was overall pleased with the result.
- captainregex 6mo agoI’m super envious. I can’t seem to do anything without a half a million tokens. I had to create a slash command that I run at the start of every session so the darn thing actually reads its own memory- whatever default is just doesn’t seem to do it. It’ll do things like start to spin up scripts it’s already written and stored in the code base unless I start every conversation with instructions to go read persistence and memory files. I also seem to have to actively remind it to go update those things at various parts of the conversation even though it has instructions to self update. All these things add up to a ton of work every session. I think i’m doing it wrong
- beaker52 6mo agoIt doesn’t really come as a surprise to me that these companies are struggling to reliably fix issues with software which relies on a central component which is nondeterministic. But they made their own bed with that one.
- thaanpaa 6mo agoWell, the fun part is that the algorithms themselves are deterministic. They are just so afraid of model distillation that they force some randomness on top (and now hide thinking). Arguably for coding, you'd probably want temperature=0, and any variation would be dependent on token input alone.
- hexaga 6mo agoMeh. Temp 0 means throwing away huge swathes of the information painstakingly acquired through training for minimal benefit, if any. Nondeterminism is a red-herring, the model is still going to be an inscrutable black box with mostly unknowable nonlinear transition boundaries w.r.t. inputs, even if you make it perfectly repeatable. It doesn't protect you from tiny changes in inputs having large changes in outputs _with no explanation as to why_. And in the process you've made the model significantly stupider. As for distillation... sampling from the temp 1 distribution makes it easier.
- ljm 6mo agoI've noticed a lack of product cohesion in general and it does make me wonder if it's a result of dogfooding AI. For example, chat, cowork and code have no overlap - projects created in one of the modes are not available in another and can't be shared. As another example, using Claude with one of their hosted environments has a nice integration with GitHub on the desktop, but some of it also requires 'gh' to be installed and authenticated, and you don't have that available without configuring a workaround and sharing a PAT. It doesn't use the GH connector for everything. Switch to remote-control (ideal on Windows/WSL) or local and that deep integration is gone and you're back to prompting the model to commit and push and the UI isn't integrated the same. Cowork will absolutely blow through your quota for one task but chat and code will give you much more breathing room. Projects in Code are based on repos whereas in Chat and Cowork they are stateful entities. You can't attach a repo to a cowork project or attach external knowledge to a code project (and maybe you want that because creating a design doc or doing research isn't a programming task or whatever) Use Claude Code on the CLI and you can't provide inline comments on a plan. There is a technical limitation there I suppose. The desktop app is very nice and evolving but it's not a single coherent offering even within the same mode of operation. And I think that's something that is easy to do if you're getting AI to build shit in a silo.