6 ms·
As an Opus user, I genuinely don’t understand how someone can work for weeks or months without regularly opening an IDE. The output almost always fails. I repe
by superze 9mo ago
As an Opus user, I genuinely don’t understand how someone can work for weeks or months without regularly opening an IDE. The output almost always fails.
I repeatedly rewrite prompts, restate the same constraints, and write detailed acceptance criteria, yet still end up with broken or non-functional code.its very frustrating to say the least Yesterday alone I spent about $200 on generations that now require significant manual rewrites just to make them work.
At that point, the gains are questionable. My biggest success is having the model take over the first Design in my app and I take it from there, but those hundred lines if not thousand lines of code it generates are so Messi, it's insanely painful to refactor the mess afterwards
- christophilus 9mo agoI’ve had decent results from it. What programming language are you using?
- falcor84 9mo agoWhy would you spend $200 a day on Opus if you can pay that for a month via the highest tier Claude Max subscription? Are you using the API in some special way?
- jefffoster 9mo agoAt a guess an Enterprise API account. Pay per token but no limits. It’s very easy to spend $100s per dev per day.
- simonw 9mo agoThe $200/month plan doesn't have limits either - they have an overage fee you can pay now in Claude Code so once you've expended your rate limited token allowance you can keep on working and pay for the extra tokens out of an additional cash reserve you've set up.
- merlincorey 9mo ago> The $200/month plan doesn't have limits either... once you've expended your rate limited token allowance... pay for the extra tokens out of an additional cash reserve you've set up You're absolutely right! Limited token allowance for $200/month is actually unlimited tokens when paying for extra from a cash reserve which is also unlimited, of course.
- simonw 9mo agoI think you may have misunderstood something here. When paying for Claude Max even at $200/month there are limits - you have a limit to the number of tokens you can use per five hour period, and if you run out of that you may have to wait an hour for the reset. You COULD instead use an API key and avoid that limit and reset, but that would end up costing you significantly more since the $200/month plan represents such a big discount on API costs. As-of a few weeks ago there's a third option: pay for the $200/month plan but allow it to charge you extra for tokens when you reach those limits. That gives you the discount but means your work isn't interrupted. Extra Usage for Paid Claude Plans: https://support.claude.com/en/articles/12429409-extra-usage-for-paid-claude-plans https://support.claude.com/en/articles/12429409-extra-usage-...
- deleted 9mo ago[deleted]
- merlincorey 9mo agoThank you for the explanation, but I did fully understand that is what you were saying. What I don't fully understand is how you can characterize that as "not limited" with a straight face; then again, I can't see your face so maybe you weren't straight faced as you wrote it in the first place. Hopefully you could see my well meaning smile with the "absolutely right" opening, but apparently that's no longer common so I can understand your confusion as https://absolutelyright.lol/ https://absolutelyright.lol/ indicates Opus 4.5 has had it RLHF'd away.
- 9mo ago
- falcor84 9mo agoOh, I wasn't arguing that it isn't "easy to spend $100s per dev per day". I was just asking what the use-case for that is.
- miguel_martin 9mo agoThis is what an AGENTS.md - https://agents.md/ https://agents.md/ (or CLAUDE.md) file is for. Put common constraints to correct model mistakes/issues with respect to the codebase, e.g. in a “code style” section.
- cloudflare728 9mo agoSometimes I have a similar file or related files. I copy their names and say use them as reference. Code quality improves by 10 times if you do so. Even providing a a example from framework's getting started works great too for new project. Yeah the pain of cleaning up small mess is great too. I had some tests failing and type failing issues, I thought I will fix it later by only using AI prompt. As the size was growing, failing Typescript issues was growing too. At some point it was 5000+ type issues and countless number of failing unit tests. Then more and more. I tried to fix with AI, since it was not possible fixing old way. Then I discarded the whole project when it was around 500k lines of code.
- pca006132 9mo agoQuestion: How many LoC do you let the AI write for each iteration? And do you review that? It sounds like you are letting it run off leash.
- cloudflare728 9mo agoI had no idea how it would end up. It was first time using AI IDE. I had only used chatgpt.com and claude.ai for small changes before. I continued it for the experiment. I thought AI write too many tests, I will judge based on test passing. I agree, it was bad expectation + no experience with AI IDE + bad software engineering.
- shepherdjerred 9mo agoI hardly ever open an IDE anymore. I use Claude Code and Cursor. What I do: - use statically typed languages: TypeScript, Go, Rust, Python w/ types - Setup linters. For TS I have a bunch of custom lint rules (authored by AI) for common feedback that I've given. (https://github.com/shepherdjerred/monorepo/tree/main/packages/eslint-config/src/rules https://github.com/shepherdjerred/monorepo/tree/main/package...) - For Cursor, lots of feedback on my desired style. https://github.com/shepherdjerred/scout-for-lol/tree/main/.cursor/rules https://github.com/shepherdjerred/scout-for-lol/tree/main/.c... - Heavy usage of plan mode. Tell AI something like "make at least 20 searches to online documentation", support every claim with a reference, etc. Tell AI "make a task for every little thing you'll implement" - Have the AI write tests, particularly the more expensive ones like integration and end-to-end, so you have an easy way to verify functionality. - Setup Claude Code GHA to automatically review PRs. Give the review feedback to the agent that implemented it, either via copy-pasting or tell the agent "fetch review comments and fix them". Some examples of what I've made: - Many features for https://scout-for-lol.com/ https://scout-for-lol.com/, a League of Legends bot for Discord - A program to generate TypeScript types for Helm charts (https://github.com/shepherdjerred/homelab/tree/main/src/helm-types https://github.com/shepherdjerred/homelab/tree/main/src/helm...) - A program to summarize all of the dependency updates for my Homelab (https://github.com/shepherdjerred/homelab/tree/main/src/deps-email https://github.com/shepherdjerred/homelab/tree/main/src/deps...) - A program to manage multiple instances of CLI agents like Claude Code (https://github.com/shepherdjerred/monorepo/tree/main/packages/multiplexer https://github.com/shepherdjerred/monorepo/tree/main/package...) - A Discord AI bot in the style of my friends (https://github.com/shepherdjerred/monorepo/tree/main/packages/birmel https://github.com/shepherdjerred/monorepo/tree/main/package...)
- moffkalast 9mo ago> make at least 20 searches to online documentation Lol sometimes I have to spend two turns convincing Claude to use its goddamn search and look up the damn doc instead of trying to shoot from the hip for the fifth time. ChatGPT at least has forced search mode.
- shepherdjerred 9mo ago
- SkyPuncher 9mo agoMy trick is to explicitly roll play that we’re doing a spike. This gets all of the models to ignore all of the details they normally get hung up on. Once I have the basics in place, I can tell it to fix details. It’s _always_ easier to add more code than it is to fix broken code.
- tmaly 9mo agoWhat does your software creation workflow look like? Do you have a design phase?
- nowittyusername 9mo agoMost people have not fully grasped how LLM's work and how to properly utilize agentic coding solutions. That is the reason for issues when it comes to vibe coders having low quality code. But that is not the limitation of technology but the user (at this stage). Basically think of it this way everyone is the grandma that has been handed a palm pilot to use to get things done. Grandma needs an iPhone not a palm pilot but the problem is that we are not in that territory yet. So now consider the people who were able to use the palm pilot very successfully and well, they were few and they were the exception, but they existed. Same here. I have been using coding agent for over 7 months now and have written zero lines of code, in fact I don't know how to code at all. But i have been able to architect very complex software projects from scratch. Text to speech , automated llm benchmarking systems for testing all possible llama.cpp sampling parameters and more, and now im building my own agentic framework from scratch. All of these things are possible and more without writing one line of code yourself. But it does require understanding how to use the technology well to get this done.
- mirsadm 9mo agoIf you don't know how to code then you are not able to judge what your producing accurately.
- nowittyusername 9mo agohere you go I open sourced one of the projects https://youtu.be/EyE5BrUut2o https://youtu.be/EyE5BrUut2o
- krior 9mo agoAll of the applications you mention could be scoped as beginner projects. I don't think they represent good proofs of capability.
- nowittyusername 9mo agoWell why don't you look at it for yourself and tell me if this looks like a beginner project https://youtu.be/EyE5BrUut2o https://youtu.be/EyE5BrUut2o
- throwatdem12311 9mo agoI have a hell of a time just getting any LLM to write SQL queries that have things like window functions, aggregates and lateral left joins - even when shoving the entire database schema DDL into the context. It's so frustrating, it regularly makes me want to just quit the profession. Which is why I still just write most code by hand.
- deadbabe 9mo agoIf you really know SQL, writing an SQL query basically just feels like writing a prompt for a database client anyway, except it does exactly what you ask for.
- throwatdem12311 9mo agoI have a running joke at work. * LLMs are just matrix multiplication. * SQL is just algebra, which has matrix multiplication as part of it. * Therefore SQL is AI * Now who is ready to invest a billion dollars in our AI SaaS company? Or it’s just that astronaut with a gun meme: “Wait AI is just SQL?….Alway has been.”
- data-ottawa 9mo agoI write a lot of SQL and I haven't had these issues for months, even with smaller models. Opus can one shot most of my queries faster than I could type them. Instead of stuffing the context with DDL I suggest: 1. Reorganize your data warehouse. It needs to be easy to find the correct data. Make sure you use ELT clear layers, meaningful schemas, and have per-model documentation. This is a ton of work, but if done right the payoff is massive. 2. I built a tool for myself to pull our warehouse into a graph for fuzzy search+dependency chain analysis. In the spring I made an MCP server for it and Claude uses that tool incredibly well for almost all queries. I haven't actually used the GUI or scripts since I built the MCP. Claude and Devstral are the best models I've used for SQL. I cannot get Gemini to write decent modern sql -- even the Gemini data science/engineer agents in Google Cloud. I occasionally try the paid models through the API and still haven't been impressed.
- 9mo ago
- deleted 9mo ago[deleted]