8 ms·
I've recently been using AI a lot for performance optimisation during a particularly busy period at work. I would say it was almost completely useless at the hi
by VBprogrammer 2mo ago
I've recently been using AI a lot for performance optimisation during a particularly busy period at work. I would say it was almost completely useless at the high-level direction - it would point out suspicious parts of SQL queries for example but on back to back testing these almost never resulted in any performance change.
In fact, if it wasn't for the fact that it made making the actual changes I identified much easier (move these joins into a CTE etc) it would have been a detriment. Not only did I get sidetracked by a bunch of useless suggestions but I also had to put up with others dumping their raw AI output at me as if it was somehow a meaningful contribution.
- colechristensen 2mo agoThen you're not grounding it to reality properly. LLMs get bad pretty quickly when you just tell them to pull the answer out of thin air. Pin them to reality with actual performance tests to run to test theories and the results will be much better. Performance optimization involves simulating real world loads and measurable results. Give them the opportunity to actually have a closed loop if you want more than trivial improvements.
- maccard 2mo agoHave you any example materials that show a harness that an agent can be pointed at an app with some sort of telemetry tool, the results it gained and the cost of doing so? Because my experience is the same as the parents - the LLM goes on massive tangents, and the more tangents it goes on the worse the results get.
- colechristensen 2mo agoUltimately the harness is me and the experience is like managing an unruly toddler. You have to pay attention to what it does and issue corrections. Skills and prompts and AGENTS.md do some, nested sets of agents do some, but over it all is me keeping track of what it's doing and needs to do. You have to size the unit of work to its useful attention span, you have to have code architecture that is conducive to units of work, and you have to have tools like ticket managers, etc. that keep the big picture and smaller pictures in mind. Reverting from agile practices to more pre-planning whole project documentation and architecture decisions helps. But in the end it's you. LLMs have their limits and need humans to direct them.
- maccard 2mo agoGotcha, so no.
- bmurphy1976 2mo agoYou don't need a "blog post." You need experience managing projects, tracking progress, coordinating conversations, mentoring less experienced engineers, etc. That experience comes with practice and time. If you really need something to latch onto, there's plenty of educational material available about how to be an effective scrum master. Start there.
- colechristensen 2mo agoIndeed. I'm saying the human inside the engineering loop is essential. If a person needs training, particularly if the set of high level advice I set out in a reply is rejected, a blog post isn't going to cut it anyway. Some AI skeptics paradoxically insist humans are irreplaceable and then get almost mad(?) when you discuss how humans are necessary in the loop. At least that's my interpretation of what is going on, it's hard to tell. When building with LLMs I'm mostly product manager and QA, the whole task is having the judgement to evaluate and correct course. You can share advice but it's not a copy-paste situation where one set of prompts/harnesses/whatever solve all your problems forever. They have to be adjusted for a particular situation.
- maccard 2mo ago> If a person needs training, particularly if the set of high level advice I set out in a reply is rejected, a blog post isn't going to cut it anyway. My problem with the “high level advice” is that the results it gives are bad. Sure, sometimes it gets it right, and it helps but I’d wager about half the time the results are just bad. Hence why I’m asking people to show their homework here. I think that the anthropic rust rewrite is great example - it shows yes you can do a lift and shift assuming you’re ok with those constraints. > Some AI skeptics paradoxically insist humans are irreplaceable. But yet in this comment thread we have people saying “just point it at your code and let it go” like [0], you saying “just follow the vague instructions and if you can’t get the results that I’m telling you you’ll get them it’s your problem. But it’s too circumstantial for me to be able to tell you how to work” > where one set of prompts/harnesses/whatever solve all your problems forever I don’t want forever, I just want to know what is actually working today, or this week, or this month. [0] https://news.ycombinator.com/item?id=49122616 https://news.ycombinator.com/item?id=49122616
- ndriscoll 2mo agoNot sure what materials your hoping for, but I've pointed codex at pprof application endpoints and just told it to identify hotspots and propose lowest hanging fruit/highest ROI items to fix, which I then approved it to open PRs for. Reduced application CPU usage by ~30% on loaded servers. It's not magic. Fairly obvious oversights that I or someone else on the team could've found and fixed if we had looked, but it was never a business priority so we never did until I curiously spent 5 minutes asking it to do it for me.
- maccard 2mo agoIdeally a blog post on an open source project that shows harness, model, prompt and the before/after measurements plus the PR.
- redsocksfan45 2mo ago[dead]
- simonw 2mo agoTell the coding agent to install or build its own performance and profiling tools. A few lines of Python with a time.time() call is enough to start iterating on performance improvements, and the really good models know how to use much more advanced existing tools.
- WhyIsItAlwaysHN 2mo agoThe mistake there is to point it at code to figure out performance optimizations. The place to find them would be performance profiles, query plans, telemetry. The guidance for perf still applies, measure before and after change. The issue is that the code often does not contain the information to do a perf optimization. Eg. you can't tell your cache size, the volumes of data in your DB or the latency of your network through just the text. Should you provide this context, you can get better results.
- flohofwoe 2mo ago> The place to find them would be performance profiles, query plans, telemetry. Tbh, once this information is available (which is the tricky part), in 99% of cases there's really no AI needed to analyze the data, since the 'low hanging fruits' will usually stand out anyway. And once you get into the area where optimization hotspots are no longer obvious, you're already deep in the diminishing returns area and optimizations for one use case or hardware configuration may degrade performance on others. That's my experience anyway.
- dewey 2mo ago> there's really no AI needed once this information is available Most people are not arguing that problems are too tricky for a human to solve once presented with it, but AI can look at 1000 things at the same time across a whole code base and then just come back with the results a human can review.
- flohofwoe 2mo agoTrue, but at least IME, once you get beyond fixing the 'obvious bugs' (how many there are depends on the initial quality of the code base - but I agree that AI is really useful to find those), you'll get into a murky grey area of "maybe false positives" which require a lot of time to validate. As a result you sink a lot of time creating and tweaking code-base-specific rules to try filtering out false positives, and IMHO this is exactly the tipping point where the whole thing becomes pointless because it becomes a bureaucratic monster. E.g. a good code analysis tool needs to work predictably at button press on any code base. Still better than nothing of course, e.g. I actually think bug scanning / code analsysis is indeed the one area where LLMs are actually useful, but it suffers from the same 'diminishing returns' problem as traditional approaches, if not worse (e.g. still no silver bullet, but a mostly useful additional tool in the toolbox).
- sigmoid10 2mo agoWhich model/agent/harness tool did you use? I've found what you describe was my exact experience some ~6-8 months ago, but since about a month or so the game has completely changed. Using 5.6 Sol with highest reasoning setting in Codex or Fable in Code, the models come up with a list of possible improvements from static analysis (ranked by complexity/benefit), write and run their own custom profilers and deliver significant performance improvements with barely any input needed from my side. So this is no longer a model issue, it's a user toolchain issue.
- qsort 2mo agoVery hard to say anything definitive on this because it's a moving target, but last time I tried models still had a distinct sense of "consistently good, sometimes great at micro, bad at macro". Similar to how, even for relatively pedestrian CRUD, they'll do code that's objectively fine at the function/file/class level but can still make a mess if you don't supervise them at least at a high-level.
- sigmoid10 2mo agoWhich model and harness did you try?
- qsort 2mo ago5.6 Sol on Codex, Opus 5 on Claude Code. (See sibling message, I'm not saying they're useless)
- sigmoid10 2mo agoThen it's for sure a guidance or missing context issue. When given full access to everything they need, these models deliver results on par or above the best coders I've ever met and they are lightyears ahead of the lower 90% of developers. I feel more and more that when users report they can't solve their problems this way, it's like when a gorilla is mad at Einstein after talking to him and he didn't make bananas grow instantly.
- herrkanin 2mo agoThe thing that makes it work really well is to make sure it has all the tooling to verify its hypotheses. If you allow it to run the full lifecycle in loops you will be surprised how well it works.
- stogot 2mo agoWhat tooling makes this go?
- colordrops 2mo agoThe ability to run queries and get the metadata about the run, e.g. length of run, execution plan, engine, engine params, etc
- madeofpalk 2mo agoTests! Unit tests, integration tests, random adhoc scripts. You know - TDD! I’ve been working on UI component improvements and it was doing a lousy job until i specifically told it to test in a headless browser to validate it works. I think somewhere in an AGENTS.md i have an instruction to “don’t state your guesses as fact - validate findings and results”.
- Tade0 2mo agoIt bothers me that you have to explicitly state this to the agent. Makes me think what else is missing from that file which also needs to be explicitly stated, but I don't know what don't know. "Do a good job"?
- Filligree 2mo agoWill, it depends on the AI. Anthropic used to have a lecture-length system prompt for their models to explain this stuff—part of the secret sauce for Claude Code—and famously found that the 5 series models no longer need it. As usual, if you use anything but the best model available I’m going to state that the better ones do better. If you do use the best model available, then I’ll just mention that Fable still has limits and still needs some guidance. One thing it does not do is deliberately build tests which test nothing at all, or which restate the code under test. I mention this because certain other models absolutely would.
- kypro 2mo agoI've noticed the same thing. The trend I've noticed is that AI struggles to think outside the box when making optimisations, which exactly what's needed when you've made all of the practical DB and logic optimisations to the existing code. Often you need to take a step back and question how the system is working and if there would be better ways to design it so the bottlenecks you're hitting wouldn't exist in the first place. Caching things, adding indexes, tweaking logic – these can help, but you'll quickly hit diminishing returns once you've done all of the obvious stuff. I've seen people here say how AI is great at optimising code though, but I'm not sure if that's because they're giving it optimisation problems with a lot of low hanging fruit or if they're successfully getting AI to rework their systems to remove bottlenecks. This one area I find AI to still be particularly bad at.
- asutekku 2mo agoI've had a completely different experience. I usually know the problematic part and then ask the AI to optimize that exact part with benchmarks, telling it to try various different approaches, and it very often gains massive performance increases on various tasks, with me of course steering it and giving it pointers every now and then. I was able to optimize a physical water simulation that was almost unrunnable on browsers to a buttery smooth 60fps version.
- gieksosz 2mo agoI did a lot of SQL optimisation with claude code over the last year, what made it work for me was making sure I verify everything with explain analyze. The whole loop was not fully “agentic” because of this but it was still faster and better than when done entirely manually. Claude also suggested plenty of small optimisations I would otherwise not do.
- throwaway_20357 2mo agoAs others have said, these were uninformed guesses which LLMs seem still to be willing to hand out deliberately. If you provide more information (your schemas, the distribution of values in your database tables, the EXPLAIN ANALYZE outputs, DB configuration etc.) it will do a much better job.
- cubefox 2mo ago> I've recently been using AI That's uninteresting as long as you don't specify the model you used. For example, Mythos was far better at finding security bugs than previous models.
- mainmailman 2mo agoDo you really think they mean they’ve been using mythos
- cubefox 2mo agoObviously no. There is also a significant capability difference e.g. between Opus 4.8 and Fable.
- user43928 2mo agoIn my mind there are three tiers: The SOTA: Fable, GPT 5.6 Sol, Opus 5 The "enterprise admin did not turn on the new models": Opus 4.8, GPT 5.5 The "I love hallucinated garbage": Sonnet, Qwen 3.6, GPT 5.4 mini, GPT 5.3 Codex, etc. Results vary widely
- inigyou 2mo agoWhat did your tiers consist of when GPT 5.3 was the latest?
- user43928 2mo agoI believe I was still using Opus 4.6 with the Claude Code CLI. Before that, Gemini 3 Pro in Antigravity. I have no experience with GPT 5.3 beyond seeing the nightmares colleagues produced in their MRs with GPT 5.3 Codex. It could be that they had the distilled 5.3 Codex Spark selected, I am not sure.
- izacus 2mo agoYep, as a performance specialist I've had the same outcomes - it'll find "smoking guns" galore which will sound plausible and be completely wrong. But that's ok - I still get a lot of value of the tool elsewhere (e.g. writing data analysis scripts, exploring code, breaking down bug artifacts...), but it doesn't stop the tiring BS from other people going "why not just AI it?" (especially from managers). I usually let them try to "just AI it" and most of them learn quickly that it's not that simple. (The others stay deranged and are making my job miserable.)
- prash20026 2mo agoAsk it to create a small sample to reproduce the issue in isolation. If it can't reproduce the performance changes then it drops it. That might work better. I haven't tried it with SQL but it worked with a few complex UI issues I had. It identified the actual issue after a few false starts.
- Uptrenda 2mo agoPeople in the comments sound like they're literally asking claude: "hey claude, can you find areas for performance improvements." Nah, it doesn't work like that. You have to be more like: "can you profile [...] and identify hot spots." then looks at results that are suspicious. if you don't ground it to the real world reality of your project then don't be surprised when it makes up its own. that's the first mistake i made as a new claude user. i only knew what prompts people were sharing on forums. but since they had been written by vibe coders it was all complete non-sense. also think about this: software engineering is filled with highly specialised tooling that can easily improve software quality. the thing is, most developers dont know how to use them, and if they do -- using them is so slow that its almost not worth it. any tool you can name claude can use. obscure academic proof verifiers, memory corruption scanners, debuggers, profilers, coverage tools, backdoor supply chain pattern scanners... it can run these and have results in seconds. really is like living in the future.
- jongjong 2mo agoAI can uncover certain issues and resolve certain problems but it also has its blind spots. I've been using Claude to create a software architecture diagram. It came up with a lot of useful functionality that we had neglected to show but upon further examination a lot of steps didn't make any sense. It added a box to do validation on some data then it put another box downstream to do validation again on the exact same piece of data (the second box had a different name but upon questioning, Claude conceded that it was the exact same validation step). When doing inference on user input, it put a box to extract specific insights from the prompt and then it put another box to do inference again on the inferred data to categorize it... And this would add latency and unnecessarily strip out useful context when doing the second inference... It really only needed a single inference step... I re-uploaded the diagram to Claude and confronted it and it conceded to all of my points and it even noticed a theme between them and suggested that the diagram exhibited "a pattern of redundant steps." But it was incapable of suggesting viable solutions besides name changes. All of its solutions made the design more complicated. It was incapable of simplifying the design. I've noticed this with AI code; it can make things more complicated, add more features but struggles to simplify things. Also, AI is horrible at rating things... It's just too superficial. For example, if you write a huge amount of intentionally over-engineered, tightly coupled code with poor separation of concerns but you do proper linting, define all the types, add a lot of comments, it will give it a higher score than a shorter, more readable and maintainable snippet which exhibits loose coupling and high cohesion because it lacks the superficial aspects. It's horrible because now business people who understand nothing about coding may ask AI about code quality and it will consistently rank low quality over-engineered projects that are full of bugs, unmaintable and less secure, with a higher score. Frustratingly, if you point out problems in its judgement, it will concede to your points without reservation, even building on top of your argument... but it will never actually tell you this stuff up-front, unprompted if you don't already know it! It never seems to reveal new insights, at best it can only expand on existing insight which you've already had. It tells you what you want to know but not what you need to know.
- deleted 2mo ago[deleted]
- shadim80 2mo ago
- UqWBcuFx6NV4r 2mo agoYou have to learn to use your tools, not try whatever intuitively made sense to you at first (expecting the tool to do all the work) and then whining on Hacker News when it doesn’t work out for you.
- watwut 2mo agoAI can not fail. It can only be failed. AI can not have limitations. Only the user can be limited.
- whatevaa 2mo agoAre you still speaking about a tool or starting a religion?
- oblio 2mo agoThis part is obviously their fault and they need to use their tools better: > others dumping their raw AI output at me Sheesh.
- geraneum 2mo agoI’m terribly sorry on their behalf. I hope the expression of their experience has not hurt Claude’s feelings (IPO valuation). Won’t happen again.
- raincole 2mo agoWe're reliving the early day Google era. It's an objectively simple tool to use, but some people just refuse to put any effort to learn. I'm thinking the issue is probably that LLMs/harnesses are too easy to use? It crossed the thin line between magic and tooling and blur the mental model. If Claude Code were as hard to use as, say, ComfyUI, perhaps there would be less programmers having absurd expectation of it.
- IanCal 2mo agoThis is the research taste part - often they can be good but human experts are really good at this (also you know your codebase, without the prompt these models have no other background). Then being able to suggest several things to try and have them go off and build, measure and tweak is hugely useful in my experience. Also things like making custom visualisations for comparing changes.
- aurareturn 2mo agoYou're using it wrong. No really you are. Give Fabe 5 access to a test database with some data, or even restricted access to your live DB and tell it to optimize then. I've had stunning success optimizing for performance this way. In a single day, I made the core part of our app 2-3x faster.
- 2sk21 2mo agoThats fine but then Fable 5 should have requested this information instead of blundering along. So why didn't it? An expert human asked to do the same task would have surely asked for the additional data.
- aurareturn 2mo agoNo one said Fable 5 is human. That's like saying Fable 5 should have told me how to prompt it to solve the Riemann hypothesis like Terence Tao would. The skill of the user still matters.
- sideeye 2mo agoThat's what I add at the end of my prompts: "Ask me questions on anything that's not clear related to this task" or similar. Both the CLI and the VSCode plugin have a nice interface designed for this, asking the user questions.
- dwaltrip 2mo agoSure, but it didn't. And you get way better results if you do that, so do that.
- bmurphy1976 2mo agoLet me translate this for you: "I have not spent the time cultivating the soft skills necessary to leverage this tool successfully therefore it's the tool that sucks."
- sirsinsalot 2mo agoEven if true, this kind of response is unhelpful and fosters dismissal of your point. This kind of attitude more broadly paints AI advocates as cultists because they refuse to engage beyond "you're doing it wrong". How about give pointers on "doing it right?"
- deleted 2mo ago[deleted]
- hansmayer 2mo ago[dead]
- heaney-555 2mo ago[flagged]
- conjectures 2mo ago...and harnesses, routers &c
- taspeotis 2mo ago[flagged]
- timcobb 2mo agoDid you give it access to a planner/analyzer?
- moduspol 2mo agoMy favorite thing lately is that it's very good at resolving dependency security vulnerabilities. GitHub already does a good job at detecting and notifying which packages are vulnerable and at what versions they're fixed, and without AI, it's just a super tedious process to figure out what things all depend on that package, find a version that'll work for all of them (potentially including upgrading the things that depend on it), wire it up, test things, and make any minor changes necessary in our code. It's not "hard" or novel--it's just tedious grunt work. I'm happy to have AI solving that one.
- nonethewiser 2mo agoDid it ever tell you not to optimize what you should redesign?
- insanitybit 2mo agoI've consistently found that I see performance issues that the AI misses. It often says "that's not going to be what improves performance, it's noise" and then I get it to do it and it's like a global 30% throughput win lol. I think a lot of performance guidance it'll be trained on is shit - I see devs consistently misunderstand performance too and downplay the impact of anything other than "IO".
- CuriouslyC 2mo agoJust give them a profiler. They drill down just like a human would and test stuff and validate. It works great.
- bluGill 2mo agoAnd without a profiler they do about as bad as humans: spending a lot of time optimizing code that is rarely used (and then often with small n)
- insanitybit 2mo agoProfiling is awesome, I do that. But it's also way slower. Like, 1000000% I'm with you, that's the right way. But I've seen huge gains just from reducing allocations and changing hash algorithms.
- eviks 2mo agoWhy are you sidetracked instead of piping that to another AI?
- douglee650 2mo agoWere you using Fable 5?
- dhruvrrp 2mo agoA self validating loop goes a long way with getting AI to optimize stuff. We had some builds that used to take a long time (15-20~ minutes), that i just asked Claude to reduce the build time. It brought it down to 2 minutes and under 10 seconds for incremental builds. Turns out folks had been copy pasting some setup code which took 2 minutes to run, and only needed to be run once per environment. Claude refactored the code to get it to run once, and since we had tests which validated the build artifacts, it was able to roll back if its change had broke the build.
- wenc 2mo agoEchoing the other folks, I have a different experience. I profile sql performance and LLMs find more opportunities than I could. All it takes is real data, a sql repl and an agent. Just ask the agent to use the repl to EXPLAIN and profile the sql. It works amazingly most of the time.
- Version467 2mo agoYup, I do this regularly now. Whenever I think that something takes longer than it should I just tell it to set up a testing harness to measure each part of the process. Once I have the baseline I'll just set a goal to improve performance by 10x. Works better than it has any right to. Have to be careful to exclude changes that massively increase complexity for tiny gains after its done, but the big improvements are often sensible choices that I'd also make as an engineer (i.e. more efficient data representation, caching and memoization where it matters, parallel processing, etc.)
- mycall 2mo agoCTEs are often slower than using temp tables with indexes.
- epolanski 2mo agoWhat's the complain? That AI can't do your job yet?
- zzzeek 2mo agoI would never undertake a big performance optimization without first measuring that area's total impact. The LLM would then be great to work out the refactoring and you can give it your bench suite (which it probably wrote) to work against. I guess just another area where the LLM is useful only as long as you remain in charge using your own programming experience as a guide.
- PunchyHamster 2mo agoI had good results when I started AI with building the testing for the actual problem, then it can just run itself over and over without wasting much time on verifying its output. And at worst you still saved a bunch of time on writing the tooling to test the issue
- rib3ye 2mo agoWhich model/mode?
- vonneumannstan 2mo agoFree ChatGPT on Chat mode.
- tzone 2mo agoWhich model are you using? There is monumental difference between models. Even between "frontier models". When people tell these stories, it would be great to also add which model you were using. As an example Opus 5.0 is in completely different class compared to Cursor Grok 4.5 even if the benchmarks don't show such massive difference. Not even talking about regular stuff like Sonnet or Composer or stuff like that.
- jrockway 2mo agoYeah. If you need something to dig deep, you need to try Fable (optionally in /goal mode). For performance testing, I wrote isolated testbeds that try to impair the system in realistic ways (latency/jitter/bandwidth limit on logical WAN hops when load testing), and Fable is happy to send a bunch of agents at it and iterate until it gets the results it's looking for. I think that if you are used to Sonnet medium or something, this will surprise you, but models like Fable and Sol on high/xhigh will really dig deep until they meet your goal. (I mostly use this for bug hunting and not perf, but ... I think it can do perf if you set it up right.)
- vonneumannstan 2mo agoIt's almost always a version of "I used the free version of ChatGPT 10 months ago and it couldn't code well."
- dataviz1000 2mo agoIt is very cheap to: 1. Ask Claude Code or coding agent to research the internet, documentation, and Github for examples and learning working with similar problems. 2. Make the ten best examples of solving the problem based on the research. 3. Run each against the database enough times (100, 1000, or 10,000,000 depending) to empirically prove which is optimized. You can set other criteria based on your knowledge. 4. ??? 5. Profit The point is, it is cheap to research combing through 1,000s of examples and learnings and test the best and most relevant.
- Forgeties79 2mo ago[dead]
- wuliwong 2mo agoI am not denying what you are saying but I am interested in what models you were using for these tasks?
- stack_framer 2mo agoThe useless suggestions are exhausting. One of my co-workers uses Claude for all asynchronous communication, including Slack messages, Jira comments, code reviews, and emails. Every single message from him, literally every time he communicates, it's a massive wall of text, overflowing with scope-creep suggestions like nothing I've ever seen before. It's impossible to ask him a simple question and get a simple answer. And this seems to be a broader trend with AI in general. It's getting more and more difficult to retain some semblance of brevity with AI tools. I'm constantly asking them to "be brief," and "only answer the question I asked," and "don't provide more information than I requested," but they keep churning out more than I want in their responses. It looks like a subtle attempt to use up more tokens.
- chrisjj 2mo agoHave you ever seen him and Claude in the same room?
- rootusrootus 2mo agoAs much as I don't truly love the shop I'm at now, stories like this keep me from looking for anything new. At least until the AI psychosis [hopefully] passes. I can't even get my management to approve a $100 Claude Max subscription, which is kind of annoying but also means they have not changed their expectations on what the team produces. None of my coworkers is going to be shoveling AI slop my direction because then they'd have to explain to management why they are using an unapproved tool to look at internal company code.
- tracker1 2mo agoYeah, I can't even run AI on work hardware... I have been able to use it for writing small utils/libraries that I then pull into the work... but the divide is clear and I review all the code myself. For a couple examples, working through an animated loader for html/js/css with an svg for the org. Another was working through a library implementation to work against an interface that was designed for Mongo, but the org is using SQL Server. Latest was a quick util to extract a zip file of pdfs into a 1bit(b/w), zopfli compressed png file per page. Generally stuff I could do, but would take me a few days for research and experimentation vs an hour or two with AI.
- internet2000 2mo agoIs this with the good Anthropic models? With extended thinking? And a way for it to validate the changes? My experience is the complete opposite.
- stackskipton 2mo agoI hate these type of questions, they scream "YOUR HOLDING IT WRONG." My experience is if you have been racking up the tech debt and have some low hanging fruit, Claude is really good at cleaning it up. If you have been properly stewarding the project, it's much less effective.
- toader 2mo agoIf YOU'RE going to be snarky at least use proper grammar.
- sirsinsalot 2mo agoWay to prove their point.
- toader 2mo agoHow so?
- cyral 2mo agoI'm surprised to hear this. I've found AI incredibly useful for optimizing queries. Give it the EXPLAIN ANALYZE output of your query and it will have no problem optimizing it
- BowBun 2mo agoSame here, this seems like a setup issue. If you give the LLM the ability to run queries against prod-like data and review the paths, you have a fully autonomous system that can correctly optimize your ORM/queries. Been doing this for months. pgAnalyze is a great complementary tool for this and they offer MCP
- deleted 2mo ago[deleted]
- ei8ths 2mo agoFor sql, i found that it can recommend indexes that i didnt have in place. So for that part of it, its great. For optimizing code, it does an ok job, i guess.
- Bnjoroge 2mo agoWhat model is “AI”?
- nwatson 2mo agoOne of my tasks is to take various AI models the data science team has produced/fine-tuned and make them runnable with optimizations on GPUs (TensorRT, vLLM, Triton Inference Server, etc.), involving conversion, deploy, smoke-test, packaging, documentation, all within a uniform framework. The documentation includes model-provenance, simple deployment instructions (usually a Docker run), simple deployment test. The DevOps team takes this and tailors it to Kubernetes / helm-charts or whatever the target environment needs. I had worked on some earlier deployment environments for a few models where we focused on one version of Triton and one runtime technology. Even upgrading to a newer Triton version was brittle, involving a lot of command-line changes at various phases. This was written mostly pre-Claude-getting-real-good. I decided that probably Claude had matured and was way better at understanding the particulars of AI-model-GPU-deployment-and-technologies. I worked with Claude to make the framework a much more lightweight wrapper, ignorant for the most part of a lot of the deployment internals. After this refactoring and doing the first couple of models, it's quite amazing at how well Claude can figure things out. For any new model we now set up its "specialization" directory and its documentation subdirectory, point Claude at the proper AI-model files, point Claude to the sample non-optimized inference code and test data, and point Claude to a similar conversion we've done in the past. There a multi-layer class hierarchy dealing with various tiers. I ask Claude to explore the existing conversion/packaging, the model, any documentation that comes with it (a lot of times there are unexpected twists), the sample code, and the desired multiple use cases the model is meant to address, and the test data. Claude has been trained, I'm sure, on a lot of AI model conversion, so it's able to synthesize the full multi-stage conversion/deployment pipeline, come up with appropriate test cases for all the use cases involved. There are usually between 3 and 10 refinements after initial synthesis, fixing outright errors, refinemnts that the data-science team requests after playing with test deployments, etc. The options and pitfalls are vast, and without Claude each preparation likely would 10x or more longer. I just put most details in Claude's hands, and make sure the general framework is good enough to provide external uniformity. When all is working, it takes a couple of hours to make sure the documentation is good. All this to say that at least for this domain, Claude / AI has been a game-changer and has sped up the process amazingly.