3 ms·
Ask HN: For Enterprise coding agents, what's your company doing to control cost?
For those of you with enterprise pay-as-you-go plans, what's your company doing to keep costs from exploding? Do you have monthly limits? How do those limits spread across vendors? Does everyone have the same limit? What happens if you hit that limit halfway through the month?
My company is a medium size startup trying to support multiple vendors and harnesses. We're thinking of having tiers of limits (tech vs non tech) and allowing users to decide where they want to spend that.
Curious to hear what others are doing as prices and usage get higher and higher.
- steammaho 21d agoLooking to what's going on in our company I see that majority of employees doesn't understand efficient usage and how better to use models. Even programmers sometimes do crazy stupid things. I would think twice before allowing users to decide something that may directly influence spends
- robertcope 21d agoWe recently rolled out a dashboard so that you can see your token usage. Looking at my usage vs others on my team, I just cannot imagine what they're doing. I use Claude for literally everything and my usage is 10x smaller. I genuinely do not know how they can use that many more tokens. There's gotta be some kind of curve that shows token vs effectiveness and it is not linear.
- mio_ships 21d agoNot enterprise, but one thing that transferred from running a lot of agents on a small budget: the cost that got away from us was never the big deliberate runs, it was the default behaviours nobody had looked at. Two we measured and turned off: 1. Re-reading whole files after every edit "to verify". The edit result already tells you it applied. Forbidding the re-read unless a test fails cut a noticeable slice of tokens on long sessions with zero quality change. 2. Guessing loops. An agent that gets a fix wrong twice will happily try a third and fourth variant. We put a rule in the harness: after the second failed attempt it has to add instrumentation and report what it observed before it is allowed to change code again. That turned several hour-long loops into ten-minute fixes, and the token savings were incidental to the time savings. On limits: a hard monthly cap per person mostly moved the spend to the last week of the month. What worked better for us was a visible remaining-budget gauge in the tool people actually work in, so the number is in front of them while they decide whether to kick off another run. People self-regulate surprisingly well when the meter is on the dashboard rather than in a monthly report. If you do tier by role, I would tier by "how expensive is a wrong answer" rather than tech vs non-tech. A non-technical person running one careful summarisation a day is cheap; a developer with an agent in a retry loop is where the money goes.
- m0rde 21d agoThanks, great ideas to try out here!
- sshussain270 20d agoYes! Its the inefficient retries and loops that eat up the tokens. A good md file harness saves it.
- anthoniks 21d ago[flagged]
- deleted 21d ago[deleted]
- techblueberry 21d agoPraying that productivity explodes as well.
- Mohamed_Amineio 21d agoi built nexalione.com cause i was tired paying other website for things i could get for free but the most thing that reduced my cost is by hosting it alone with cloudflare it reduces cot of hosting by almost 100% and helps you run a smooth website
- kelvo_ran 21d ago[dead]
- FirstClassTree 21d ago[dead]
- baracop 21d ago[flagged]
- khantto 21d agoin my experience, i tried to use some Chinese LLM, the package plan is much affordable.
- Sattyamjjain 20d ago[flagged]
- softwaredoug 20d agoI personally think the root cause is losing hand coding skills You can be much more token efficient if you code 5-10% and let the LLM replicate those patterns. You can also maintain an understanding of the system. Avoid needless AI layered complexity. And ship less slop. I’m simultaneously bullish on human and AI coding.
- xcubic 17d agoCan you go into details on the hand coded patterns? Do you have an example?