7 ms·
I would be really curious to hear from devs at Databricks what the experience of development is like internally. I work at a small startup with essentially unli
by extr 2mo ago
I would be really curious to hear from devs at Databricks what the experience of development is like internally. I work at a small startup with essentially unlimited AI spend budget - the entire point is that I should be turning to it at every opportunity since our human labor is so expensive relative to tokens. So generally it's like:
- Spend most time prioritizing/discussing what to do.
- Once that's agreed, use Fable 5 High + 5.6 Sol XHigh come up with a design + plan. Agree on the high level plan. (Usually this just comes down to choosing where the change belongs on the spectrum between minimal patch <-> full redesign)
- Use Opus 5 or Sol Med to execute
- Auto-fix bugs and CI until green + thermonuclear review skill x3.
- Manual interrogation of change/nits
- Come up with QA plan and have Codex Computer Use execute on it
- Manually spot check the final result (usually a sizable diff, thousands of lines, complete feature E2E, etc)
I probably spend like $80 a day at least but I produce the output of 3 or 4 2022 engineers and probably at better quality. So it's easily worth it. Would I save money by switching to GLM 5.2 and such...perhaps? IDK. At our scale it's not worth the time spent building the eval harness to actually understand the performance tradeoff.
- biophysboy 2mo agoDo you have tips for generating clean productive output per dollar?
- the_sleaze_ 2mo agoin my humble experience it boils down to mastery. Are you at least conversational in the subject matter? You're gonna have a good time just by paying attention and adjusting your workflow. If you're getting a lot of back and forth with it, its asking a lot of planning type questions, stop, step back, rethink the whole feature, and start again from the beginning with everything more fleshed out. If you are in a brand new field, there's no way to bridge that divide. The issue is you don't know what is good or bad, or whether what you have learned is good or bad. You're in a sports car and you don't know how to drive much less what's track and what's field. You can spend a lot of effort getting good at prompting towards writing tests and E2E tests to at least verify your app does what you expect it to, regardless of experience.
- grigri907 2mo agoI appreciate this non-judgmental description of what it's like to approach a topic/technology from a newcomer's perspective. Thanks!
- extr 2mo agoThis is a great point and I agree. My own productivity varies based on what part of the codebase I'm working on. If it's "been in there before" and I know the right questions to ask, I can one-shot a good design/improvement. If I'm spending 20-30 minutes asking Fable to "draw a diagram so I can understand" - probably less so. But notably, I CAN get there in a fraction of the time it would have taken before. You can general personalized onboarding docs to ~anything.
- extr 2mo agoKeep the decision-making and execution separate. Use the high IQ models to chat about the design and make them drive subagents to do the actual work. "Chat" style threads are actually quite cheap. Where it gets expensive is having Fable 5 output thousands of lines of implementation where 95% of it was already overdetermined and there were only a few important judgement calls. I actually have no doubt that I could replace my Opus 5 Low/Medium subagent profiles with Grok 4.5/GLM 5.2/Deepseek v4 Flash and perf would probably be pretty similar. On top of that - highly recommend adding accurate cost counters to your statusline. You can't improve what you don't measure! (Or even have any intuition about).
- ai_fry_ur_brain 2mo ago[dead]
- RugnirViking 2mo agoDo you have issues with performance at the moment? Right now I tend to find that it produces absolutely terrible design patterns and especially performance. I mean maybe I don't know exactly what area you're looking at but yeah for us we tend to find it's terrible wrt dB/caching/scaling and often any performance improvements it proposes end up actually shooting itself in the foot and being worse than before but it's not very good at testing in an organized way to even notice it made it worse despite repeated prompts to do so I mean if I prompt it to test performance in a handheld structured way (it is very bad at finding out what performance to test and why) before making changes I can usually figure it out but it usually takes insistence on the specifics to really ensure a good solution that will actually fix the problem
- extr 2mo agoPerformance is better than ever. It's never been more practical to set up wildly complex synthetic test environments and measure perf wins. Plus the models will find every possible algorithmic/design improvement. It actually gives me quite an uncanny feeling, bulldozing over years of human optimization work with a newer, "perfect" design. Like bringing an AK-47 back to the middle ages.
- app13 2mo agoI needed to thoroughly test rerankers on my companies rather unique corpus. Opus and I wrote a parallelized test harness and labeled groundtruth in around 2 hours. In 2022 that would've likely been all I did for a couple sprints
- extr 2mo agoYes 100%. This morning I casually prompted Codex to drive the browser to complete extensive performance testing in-situ that would have literally been weeks of work before. Probably in reality it just wouldn't have been done, and performance guarantees would have been attempted up front via more careful design. In this case the design was also AI generated, and there were limited wins to be found because the design was already superb.
- reqo 2mo agoIME this works until it does not. This approach works well at the beginning of a greenfield project, but at the same time because it is so easy to add features, you will likely ship something that is way too over engineered. And that complexity will not amortize over next increments and will more likely lead to the entire project being a black box only fully understood by AI. However a more careful use of AI for targeted surgical changes is far more ”productive” in the long term IMO.
- extr 2mo agoDisagree. I operate this way inside a multi-million line legacy codebase.
- nujabe 2mo ago> I work at a small startup How does a “small startup” end up with a multi million line “legacy” codebase? Something not mathing
- extr 2mo agoHave you worked at many startups?
- nujabe 2mo agoNo, but not relevant. What is the point of working at a startup if you’re dealing with millions of lines of legacy code ? Isn’t the whole point of startups to create & innovate with a clean slate and modern tools?
- extr 2mo agoNo, actually. The point is to build a profitable business.
- sarchertech 2mo ago
- catlover76 2mo ago[dead]
- pizza234 2mo agoIn our team's experience, the product of agents is generally The Homer (1). It does work, but it's vastly overengineered. When I personally want tight code, I have to spend a considerable amount of time adjusting it manually: - It needs to be trimmed down. In my experience, at least one agent I use struggles to produce minimalist designs, and it's very frustrating - I need to consider whether there are solutions based on higher-level assumptions, that AIs typically miss - I need to check whether there are off-the-shelf solutions - AIs like to reinvent the wheel IMO, software production has become a mass-produced commodity in every sense - it's much more expensive to produce software manually, but the quality is not the same. (1) https://simpsons.fandom.com/wiki/The_Homer https://simpsons.fandom.com/wiki/The_Homer
- eitally 2mo agoAs a business user, the same thing is true for non-code documents. The biggest exertion is reducing the excessive slop down to concise, clear points.
- extr 2mo agoThis was more true a few months ago but Fable has improved the situation considerably. Also just remember - minimalist code looks and feels great but customers do not read your code. I have caught myself many times providing "corrections" to abstractions that were already ~fine, just not perfect. The average SWE costs $200/hr. Careful you don't burn $50 worrying about code that will likely be rewritten or can be better abstracted when that's actually needed.
- dieselgate 2mo ago> The average SWE costs $200/hr This is a pointless quibble but the hourly rate claim is not true--it's like ~$60 in the USA [0]. Maybe you meant at a specific Org but this is important context when comparing "pricing" between human and AI. [0] https://www.salaryexpert.com/salary/job/software-developer/united-states https://www.salaryexpert.com/salary/job/software-developer/u...
- 2mo ago
- jchook 2mo agoThis is very close to my workflow but you forgot one important step: - Suggest a better approach that makes the AI say, “That’s much simpler. And you’re right. My original plan was over-engineered.”
- nujabe 2mo ago> essentially unlimited AI spend budget > I probably spend like $80 a day This doesn’t sound like “unlimited”, I spend more than this out of pocket per day and I have a strict budget.
- extr 2mo agoIt's a fair point, it's not truly unlimited and I do wonder how that would change my workflow. I can definitely imagine if I was inside Anthropic or OAI with unlimited "fast" tokens, you would be more tempted to hand over even more of this process. I completely understand why they talk about "graph engineering" and such, my entire workflow above could be a graph and I could try to increase my leverage even further. Realistically though I am bounded by product decision making, not code output right now.
- deleted 2mo ago[deleted]
- gamblor956 2mo agobut I produce the output of 3 or 4 2022 engineers and probably at better quality. Possibly, but the output of a 2022 engineer is about 1/10th of the output of a 2010 engineer, so it's an extremely low bar.
- Krei-se 2mo agoalso - as always with these claims there's no actual product / repo / whatever one could check. I would love to see what these tools create but outside slop there's never: This works, is in production, here's the code. Any day now.
- dgellow 2mo agoIt’s crazy how we are like ~2y in this AI revolution and still do not have an answer to this question: can you show us the ROI? Where is the revolutionary software your team of agents created?
- bonoboTP 2mo agoWhy would it need to be revolutionary? It can be some ordinary thing. Software is mostly ordinary.
- what 2mo agoI found an interesting project recently. As I was looking through the source something felt off. Turned out to be entirely LLM written. There was duplicated code everywhere, same function defined in dozens of files (same name, same intended behavior) but none of them would produce the same output for an input. Dead code all over the place. Over architected. Useless comments. It was all generated in the last 4 months, so don’t come at me with the “but did they use a model from the last 6 months” nonsense.
- bonoboTP 2mo ago4 months is ancient. Fable and Sol are a different animal.
- K3UL 2mo agoThe output yes, but do you produce the impact and value of 3 engineers? I have seen this workflow being toyed with too, and I find it to produce massively overengineered stuff that actual people don't really wanna use
- what 2mo agoHe only spot checks 1000s loc diffs, so probably has no clue.
- catfood 2mo ago>Auto-fix bugs and CI until green + thermonuclear review skill x3. Gotta love this loop, I have it running while I'm asleep all the time.
- matsemann 2mo agoMy experience is that your description works for a certain time, since you're knowledgeable of the codebase and can guide it. But after too many iterations with not hand-holding the llm, it quickly gets unwieldy.
- stbenjam 2mo ago$80 sounds extremely low for what you're describing - are you on API token plans? I have had some $3,000 token days - even without Fable. I don't see how this is sustainable. My personal 20x plans get so much usage for so cheap. The consumer subsidies are crazy, but alas I can't use them for work.
- extr 2mo ago$80 is definitely low now that I look at my numbers. but not OOMs low, it's closer to like $200 on heavy days. i don't know how you're doing $3k/day, that's wild. i'm pretty aggressive about compaction and session restarts, and i reserve Fable 5/Sol XHigh for "main thread" orchestration
- Footprint0521 2mo agoDude $3k? Holy heck you should look into K3/Deepseek V4 Flash
- ajcp 2mo agoOnly spending $80 a day on Opus 5/Fable 5/GPT 5.6 Sol feels very low. I'll roll through a couple hundred dollars worth of credits a day with those models, the vast majority of which would be on non-coding tasks, and it's still a huge cost savings over me or my team having to do these things manually, if we'd even be able to do them at all. But that's also why it's now easy to justify the cost of an Nvidia or Intel inference server with Kimi K3 locked and loaded :)
- bryan0 2mo ago> Spend most time prioritizing/discussing what to do. you should probably be doing this discussion work along with Fable 5. It will give good feedback if you're working on the correct things. > Come up with QA plan and have Codex Computer Use execute on it QA plan should be part of the above "design + plan", not after it. The implementer needs to be able to fully test before publishing a PR. This is true whether humans or agents are writing the code. > Manually spot check the final result (usually a sizable diff, thousands of lines, complete feature E2E, etc) unfortunately this is not really scalable with amount of code agents can produce, so you need independent (fresh context) agent reviewers to help. Ideally they only escalate to a human when really stuck. > I probably spend like $80 a day at least at a small startup you should be on the $200/month plan(s).
- bdangubic 2mo agodo the same across 20 terminals (as you should) and now you are up to $1.6k/day. would that give you output or 60-80 engineers? not a chance, right? no one’s AI spent will be in question working a single terminal with carefully planned out and executed process you do
- drTobiasFunke 2mo agoOutput of 3 or 4 2022 engineers? Its that your self assessment? Output as in number of lines of code?
- swader999 2mo agoNot op, but we are at 2400 total points delivered over seven years. 1000 of those in the last six months. About 2-4 devs over that period, just two the last six months.
- drTobiasFunke 2mo agoNoone is questioning the volume of llm output. My question is whether all those points delivered improved your product and software in any meaningful way, or do you now have 20x more code that noone understands with the same quality of software product?
- swader999 2mo agoYeah it did. We've signed three new customers because of one. Another large feature was an integration that landed more. We rewrote the entire front end to modern stack. Used to be a fifteen year old react and backbone bFrankenstein app.
- aetherspawn 2mo agoI was just about to say, how could routing possibly be worth it at the risk that the work output is sub par?
- samesense 2mo agoYou have an unlimited budget, and you only spend $80/day? I’m up to $3k/week, and still expanding.
- 827a 2mo agoI'm probably between $50-$200/day depending on the day; we also have effectively unlimited budget, though a lot of that is because Azure gives startups $150,000 in credits for 2 years, which we've wired up to a LiteLLM gateway & OpenCode. Without that I think our appetite would be more around $400/month/employee. A lot of my high costs is because I just throw Sol at everything. If I were more selective and brought in Luna or v4 Flash every once in a while, I think I'd be more like ~$400/month. That's why I'm not aligned with the notion that "tokens are subsidized so that's why people are using so much": its not that I'll have to adjust to using less, its just that I'd need to think before I prompt a bit and be more judicious. I could easily see my raw token counts doubling or tripling in the coming months. I don't think that will change as subsidization subsides; though maybe lab revenue will; intelligence per dollar is getting cheaper every week. Its solely a function of adaptation to process, which takes time. The productivity gains per token are the single most asymmetrical thing I've ever seen in engineering. The engineers on our team are pretty effective with tokens; easily that 2x-4x output as you're seeing, spending $20-$200/day. Some of our security folks have also started contributing more-and-more code, and they're on the other side: they'll spend hundreds a day running in circles, eventually producing these +/-30k loc pull requests that take ages to get merged and are littered with issues. They weren't writing much code before, so arguably they're more productive by some multiplier greater than 1, but I think the drag on the rest of the team, and potential issues with what they produce, has overall created a net-negative situation. Inversely, some other company functions have produced a few one-off websites for things like sales processes, and those have been a huge win. The asymmetry is wild. There's almost a valley of incoming skill where if you know nothing about code, you'll leverage it well; if you know just a little bit, it makes you super dangerous; if you know a lot, you're the biggest winner. Really difficult situation to navigate.
- tfehring 2mo agoI'm also at a startup. My workflow is similar but I have Fable 5 xhigh drive the whole thing: it gets Codex CLI installed in its environment with an API key, and it's instructed to delegate ~everything to Codex and review its work, especially for code quality/conciseness. Fable delegates to Sol or Luna (fast mode) xhigh/max depending on the task - I think Luna xhigh on fast mode is basically a Pareto improvement over Sol medium.
- willsmith72 2mo ago> I probably spend like $80 a day Wait what? I don't understand these numbers. I spend $1k/day Your story about being told to use AI for everything I was expecting you to be well over that
- mrlongroots 2mo agoIn my experience, code is a small fraction of the work. I'm in an infra team and for the last 2 weeks or so I've been trying to understand whether a particular workload will catch fire if a switch is flicked. I'm also new to the team so partly it is me wearing training wheels, familiarizing myself with the telemetry etc, but I will state that I'm not completely lousy at this stuff. No model in my experience can do anything remotely comparable to the work "what happens to the workload if this switch is flicked" needs. They can't even design a reliable quick experiment to answer what cast should be applied to the binary trace_id in table A for the join to table B to work. They will happily do something idiotic and then conclude that the join does not work.
- ianmarcinkowski 2mo agoAn AI maximalist on my team put up 2 pull requests with ~120-140 changed files this week. If they spent $200 on tokens to achieve this, we spent $2500-3500 in human salary and opportunity cost reviewing it.
- z0mghii 2mo agoYou need to change the way you think about reviews
- eru 2mo agoInteresting. I would probably start with the QA plan first, or at least before implementation (and perhaps even before design.)
- ffsm8 2mo ago> I probably spend like $80 mate, if youre not using subscription then youre spending waaaaaaay more. the plan itself with fable/sol will most likely have already cost more then $80 -- ime thats more like 500-2k/day of usage. most harnesses let you see the usage in the status bar, i encourage you to enable it
- deadlast2 2mo agoDoes it not end. Like is there not a point with all this speedups and infinite intelligence that your software system is essentially done.
- lelanthran 2mo agoIs the revenue up by 2x or 4x?
- newsicanuse 2mo agoMakes me wonder the kind of startup this peron is working for where slop is encouraged
- imilev 2mo agoInterested to dive deeper on the upfront design discussion. Have you found these to more often then not translate into the real product. In my experience at the begining of the full agentic coding loop in our company we were more hands on with the codebase and had better judgement over the plans. Now it is quite often that the inital plan after executed needs more refinement and that made the plan review somewhat obsolete for us.
- tosh 2mo agofor complex open ended coding tasks better models are better (and mid-to-long-term, often also short-term end up cheaper than weaker models) this might change soon if we are reaching a certain capability threshold but right now that's still the case unless you are working on throw-away trivial stuff where iteration speed and trying many speculative things might give you an edge