5 ms·
There's definitely a way to use Claude code that is token conscious. I've tried throwing unsupervised agentic software factory workflows against the wall, and
by tra3 4mo ago
There's definitely a way to use Claude code that is token conscious.
I've tried throwing unsupervised agentic software factory workflows against the wall, and they burned through my tokens like nobody's business but didn't produce much.
Supervised, human-in-the-loop process on the other hand is much more productive but doesn't consume nearly as much. Maybe that's why everyone's pushing agentic approaches so much.
- SubiculumCode 4mo agoAt the enterprise level though, its going to be hard to want to use a service in which costs are not predictable, and keeping those costs under control requires employee training.
- salawat 4mo agoThere's no fucking training to mitigate a slot machine.
- dgellow 4mo agoGames like Diablo are basically a whole bunch of slot machines, and there are strategies you can follow to optimize your run.
- gambiting 4mo agoYes, because in video games there is always a chance to win so you can optimize your strategy around that chance. If you have a 1% chance to drop a legendary weapon, the question becomes how do I manufacture 100 chances for a weapon drop in the shortest possible time. With agentic coding there is no such guaranteed chance - in a way it's worse than a slot machine that is guaranteed to pay out eventually. You could spend hundreds of millions of tokens and still not get what you asked for.
- echoangle 4mo ago> If you have a 1% chance to drop a legendary weapon, the question becomes how do I manufacture 100 chances for a weapon drop in the shortest possible time. Sidenote but I hope everyone realizes that 100 is kind of arbitrary here and does not mean the total chance to to get something is 100%.
- avadodin 4mo agoyou don't have to do the math unless it's on the exam, lol.
- dgellow 4mo agoYou’re right, the arpg analogy isnt great, it’s too simplistic. I was trying to come up with something heavily stochastic where people are coming up with strategies to get the odds in their favor. Maybe closer to speculating on the real estate market? But even that feels too simplistic compared to LLMs. Even the definition of a win isn’t well defined. Actually it’s really its own thing, I don’t think the slot machine analogy works too well, you also have fixed odds (and you know they aren’t in your favor), and a binary output
- skydhash 4mo agoThe analogy to slot machine is that you're spending your own resources in hope of a reward. So you're ultimately bound by your resources and your strategy doesn't count for much in the grand scheme of things. With employees, there's a lot of punishments in place for people to not want to mess up. Loss of wages and reputation, prison time,... Startup do not fail because they have a bug-ridden product, they fail because of the market. With AI, all bets are off. They're not aligned with your goals and it's very hard to discern when they go off unless you're an expert. And if you are one, at best it's just a slight boost in typing especially with all the works involved in software development.
- serf 4mo agothat analogy is so boring now with so many real world examples of actual LLM work. people still can't get over the unreasonable effectiveness of algorithms.
- greenchair 4mo agonondeterminism will always be anathema to the engineering mind
- arkadiytehgraet 4mo agoThere have also been winners of a slot machine gamba, so the analogy quite holds. I would even argue that there are considerably more slot machine gamba winners than the real world examples of actual LLM work.
- LPisGood 4mo agoThere’s actually been a ton of research on how to optimize “slot machines,” at least in a generalized sense. For more reading, check out the literature on multi armed bandits.
- subscribed 4mo agoLOL, that's a sophisticated and sometimes slightly unpredictable multitool. If this is the "analogy" you go for, you don't seem to be suited to make that comparison.
- __mharrison__ 4mo agoOdd, I train teams (at large companies) to use harnesses effectively. So some training does exist. I get the anti/skeptic sentiment. I've been called a lot of horrible things by a vocal contingent when they hear that I help train folks to learn software engineering best practices and then apply AI to that.
- layer8 4mo agoTo be fair, the cost of software development has always been fairly unpredictable. What may be different is that the cost used to be roughly proportional to man-hours spent, while now the number of agents running in parallel may be less predictable.
- xienze 4mo ago> To be fair, the cost of software development has always been fairly unpredictable. Yes, but in a "oops this is gonna take another two months to finish" kind of way, not the "oops this is the 12th time this month 8 developers have burned $2K in tokens in a single day and no one really knows how it happened" kind of way.
- kridsdale1 4mo agoWe’re all being given belt-loaded machine guns and tossed on to Planet K. We used to pay for the salaries of soldiers, now we have an Ammo Budget.
- dgellow 4mo agoA belt loaded spinwheel machine gun, where there are some chances the next bullet is a dummy round, or goes in the wrong direction. And everytime you reload a new soldier is in charge of the gun
- bluGill 4mo agoYou don't need that analogy as the normal use of a automatic gun in war is not to kill, it is to suppress - stop the enemy from moving. If you are hit by a gun in automatic mode it is your own stupid fault. When you want to kill someone you switch to one shot or maybe 3 round bursts.
- dgellow 4mo agoTIL. I know literally nothing about automatic guns
- sidewndr46 4mo agoAm I losing my mind, aren't there multiple headlines each day about companies penalizing employees for not using AI enough?
- iSnow 4mo agoThat was roughly 3 weeks ago, with the reprising of Claude 4.7 and GPT 5.5, things have become more spicy.
- sidewndr46 4mo agouse AI, don't use AI, this whole thing is getting really hard to follow
- andrekandre 4mo agoi've worked at so many places where the propaganda/marketing and reality on the ground is so disorienting/shocking i don't really expect this to be any different...
- lukan 4mo agoIt is allmost as if humans ain't of a single mind.
- foolserrandboy 4mo ago2 months ago: no limits. 1 month ago we had a leaderboard for whoever had the highest token spend not taking into account what was actually produced. This week: “everyone is using opus too much, just use it for planning.”
- basch 4mo agosince those headlines started ive felt it just encouraged inefficiency. "say as much as you can without saying anything." if you were accomplishing your task the need for more would end, thus there is incentive to never succeed.
- mrgoldenbrown 4mo ago>...use a service in which costs are not predictable, and keeping those costs under control requires employee training. Isn't this a (mildly exaggerated) description of AWS, which is a very successful service?
- noodletheworld 4mo agoMmm… but for AWS its pay for external use right? So your costs scale with the number of users you have. Thats an op ex that you can explain. For tokens for developers its maybe closer, cost/outcome wise, to hiring an external consulting company to write your code; money paid scales with work done, no promise of delivery, arbitrary unpredictable external price changes. Its not quite the same; though, similarly lucrative for consultants.
- logicchains 4mo ago>Mmm… but for AWS its pay for external use right? Not if you're using it for running builds, running research jobs, model training, etc.
- jochem9 4mo agoYou can put a limit on token spend and provide training (and even pre-configured workflows) on how to limit token spend. Like the other commenter said: cloud spend can also spin out of control if you don't pay attention, yet we've found ways to keep it under control (training, guardrails, limits, transparancy).
- darig 4mo ago[dead]
- harimau777 4mo agoThe problem that I see is what you do if someone runs out of tokens. It doesn't very well work to say "well I guess you just get fired because you can't work at full speed for the rest of the month". Personally, this feels like its just trying to push the work of managers in allocating resources onto developers so that they have more work to do and can be blamed if anything goes wrong.
- tracker1 4mo agoMy experience as well... I've only hit Antrhopic's 5hr threshold a few times, and two of them was within a half hour of the window. Also, all three times I'd already accomplished a LOT. I tend to work with the agent, and observe what's going on as well as review/test and work through results/changes. I spend a lot more time planning tasks/features than the execution, even using the agent as part of planning and pre-documentation. It works really well. I don't think people burning through the 5hr allotment in under an hour are actually reviewing/QC/QA the results of what they're doing in any meaningful way, and likely producing as much garbage as good (slop). I'm really curious as to HOW the MS employees were using the agents as much as what they were doing.
- kristjansson 4mo agoI suspect subscription limits are quite a bit higher than the equivalent tokens their dollar cost could purchase. I similarly feel like I can get a lot done with a $20/mo Claude Pro subscriptions, but also can easily spend $10-20/day at API pricing with similar usage.
- skeledrew 4mo agoCan verify that I've gotten about $400 worth of tokens from my $20 sub.
- kridsdale1 4mo agoNow that sounds like a business I’d like to invest in! When’s that Anthropic IPO anyway?
- lawn 4mo agoI don't understand why people are using the API pricing instead of the Pro/Max subscriptions? What am I missing?
- giancarlostoro 4mo agoEnterprise customers don't get that option. But also if you want a fully custom harness, you also don't get that option.
- CoolestBeans 4mo agoThe current thinking is automated agents is what turns this from an industry in the tens of billions to a multi trillion dollar one. So yes you are right on the money, agents stimulate demand for this thing they've built.
- joe_mamba 4mo ago[flagged]
- kridsdale1 4mo agoThere is always a quantity of lubricant that can get any machine moving. Just add so much that you create an all consuming river of lube and watch your thing sail away.
- ndesaulniers 4mo agoGood then that Amazon sells it by the 55 gal drum then. https://www.amazon.com/Passion-Lubes-Natural-Water-Based-Lubricant/dp/B005MR3IVO https://www.amazon.com/Passion-Lubes-Natural-Water-Based-Lub... > This product is out of stock Ah, shoot, there go my weekend plans. Bummer.
- deleted 4mo ago[deleted]
- dualvariable 4mo ago"The bureaucracy is expanding to meet the needs of the expanding bureaucracy"
- beardyw 4mo agoI didn't know that one. Loosely said to be Oscar Wilde.
- 4mo ago
- thegreatpeter 4mo agoyeah, by using codex
- brookst 4mo agoI get 98.6% cache hits on Claude code. Short of drastic arch changes it’s hard to imagine it getting much better.
- gobdovan 4mo ago98.6% cache hits doesn't distinguish an efficient workflow from an overly chatty linear agent repeatedly reusing the same context. Plus, it says nothing directly that the process has good useful progress per token.
- kridsdale1 4mo agoWe are all going to be graded by (tickets closed / tokens burned) soon enough.
- hedgehog 4mo agoYou pay for cache hits on every turn and even with the newest architectures longer context is slower/more energy intensive. Constructing concise turns that reuse prefix and stop when the new context is no longer useful help, as does pushing generation down into cheaper models while using stronger models for verification.
- KronisLV 4mo ago> There's definitely a way to use Claude code that is token conscious. Colleague used Sonnet 4.6 on some pretty normal agentic coding tasks through AWS Bedrock to keep the data in the EU, 100 EUR usage in a single day. In comparison, the Mistral subscription costs about 20 EUR per month and we tested that for similar tasks it was okay, the usage got to around 10% of that monthly limit in a single day. Or Anthropic's own Max (5x) plan where you get way, way more tokens to do with as you please. I feel like the sweet spot is having a monthly subscription with any of the providers (you're subsidized a bunch), but if you have to pay per tokens, now I'd just look in the direction of what tasks DeepSeek would be okay for, sadly probably not in the situation above. For a startup, though... On the other hand, this feels a bit hypocritical: > It was part of an effort to get project managers, designers, and other employees to experiment with coding for the first time, and sources tell me that Claude Code has proved very popular inside Microsoft over the past six months. They're gonna say that the future is all AI... until they get the bill.
- michaelbuckbee 4mo agoI was trying to get a better sense of the time cost quality matrix of these, so I threw together a quick eval of Sonnet 4.6, Mistral's dev model, and Opus 4.7 (figuring it's what you'd use if you were on Max). The results for a function implementation and test of levenshtein distance in js are pretty similar but Mistral is 30x cheaper than Opus 4.7 and 4x faster than Sonnet 4.6. https://5m6qnuhyde.evvl.io/ https://5m6qnuhyde.evvl.io/
- KronisLV 4mo agoThe one detail I did forget to mention is that if anyone goes with the Mistral subscription (instead of paying per-token), then the Mistral Vibe tool gives you their Medium 3.5 model by default, with a 200k token context. It will probably be enough for plenty of tasks, though there's also a noticeable difference between that and up to 1M.
- kaoD 4mo agoBut that's not very informative. Levenshtein distance is not only a well-understood problem, it's small, self-contained, and extremely well-represented in the training data. The kind of problem where even small/bad models can excel. The golden standard for those tasks is just "use a library" so no wonder the beefy models are expensive: you're chartering a commercial airplane to go grocery shopping. My personal benchmarks are software engineering tasks (ideally spanning multiple packages in a monorepo) composed of many small decisions that, compounded, make or break the implementation and long-term maintainability. There's where even frontier models struggle, which makes comparisons meaningful.
- nurettin 4mo ago---- Before it was: Me: We need to do this this that. Claude: <random stuff that approximates human outout> Me: Are you sure? Claude: Well actually there is a bug <more random stuff that looks right this time> ----- Now it is: Me: We need to do this this that. Claude: <random stuff that approximates human outout> Claude: Let me consult the advisor on that. Claude: advisor came up with some advice, adjusting according to that. <more random stuff that looks right this time>
- jstummbillig 4mo agoI think it's great. People at a broad scale are getting first hand experience with resource management. It's a fairly cheap way of doing it too (in contrast to: learning this by managing humans) and we can all benefit from the skill transfer.
- visarga 4mo agoI find myself observing how my lead manages meetings ... "ah, this is like when I do that with Claude", "this is where he wants to understand what happened, like when I ask Claude" ... it's funny.
- matheusmoreira 4mo agoYeah. Claude does good work but reviewing it all properly takes quite a bit of time. It got to the point I started having trouble maxing out my weekly allocation. Dealt with that by going all out and making an agentic parallel code review skill. Basically an infinite TODO list generator. Now I'm definitely getting 100% of the usage I paid for. It really burns tokens like nobody's business, and catches a lot of issues while at it. I've been looping this review/fix process every week. It's dramatically reduced the amount of stuff I need to pay attention to during my human review sessions.
- jdsnape 4mo agoI’m interested in how this works in practise - I guess you’ve written a skill to do code review, then your Claude.md file tells it to use it after every change as a bg task? So does this work as a background task while Claude is working on the next ‘feature’?
- matheusmoreira 4mo agoI just committed the skill to my dotfiles repository. https://github.com/matheusmoreira/.files/tree/master/~/.claude/skills/scrutinize https://github.com/matheusmoreira/.files/tree/master/~/.clau... There are many "critics", one for each quality I want reviewed. Correctness, consistency, maintainability, security, testing... Everything I could think of, and I keep adding more. https://github.com/matheusmoreira/.files/tree/master/~/.claude/skills/scrutinize/critics/programming https://github.com/matheusmoreira/.files/tree/master/~/.clau... The scrutinize skill is the entry point. The Opus I'm talking to becomes an agent coordinator. He explores and autodiscovers the project's structure, subdivides it into logical sections. Then he runs a truly absurd critic x section matrix against the entire project. Literally hundreds of these agents running in parallel, each focusing on one area. Ten minutes of this is enough to exhaust my Max 5x five hour window and put a serious dent in the weekly usage numbers. It literally takes days to run a full agent sweep. I designed it around the rate limiting. The agents do file system style journaling in order to resume cleanly. They commit all of their findings as they go into an orphan branch in the repository. Further review runs can build on it and avoid searching for known issues. The way it works in practice is I just run /scrutinize sweep and then go work on something else, or just go do my actual job, live my life, play video games, write an article for my blog or something. Come back five hours later to either resume the process or check the literally hundreds of issues that have been found by all the agents. Then Claude and myself will go in and evaluate and fix all of those issues one by one. Then review again. Then evaluate/fix again. I'm just gonna keep looping this over and over until zero issues are found. For all of my projects. Going from solo hobbyist programmer to this was pretty insane. I can only imagine what these corporations with infinite money must be doing.
- blitzar 4mo ago> There's definitely a way to use Claude code that is token conscious. By buying a subscription and dealing with the limits, using claude code and paying per token seems like the fast lane to the poor house.