12 ms·
Snorting the AGI with Claude Code
- dwohnitmok 1y agoOn the one hand very cool. On the other hand, every time people are just spinning off sub-agents I am reminded of this: https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality-is-the-tiger-and-agents-are-its-teeth https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality... It's simultaneously the obvious next step and portends a potentially very dangerous future.
- TeMPOraL 1y ago> It's simultaneously the obvious next step As it has been over three years ago, when that was originally published. I'm continuously surprised both by how fast the models themselves evolve, and how slow their use patterns are. We're still barely playing with the patterns that were obvious and thoroughly discussed back before GPT-4 was a thing. Right now, the whole industry is obsessed with "agents", aka. giving LLMs function calls and limited control over the loop they're running under. How many years before the industry will get to the point of giving LLMs proper control over the top-level loop and managing the context, plus an ability to "shell out" to "subagents" as a matter of course?
- qsort 1y ago> How many years before the industry will get to the point When/if the underlying model gets good enough to support that pattern. As an extreme example, you aren't ever going to make even a basic agent with GPT-3 as the base model, the juice isn't worth the squeeze. Models have gotten way better and I'm now convinced (new data -> new opinion) that they are a major win for coding, but they still need a lot, a lot of handholding, left to their own devices they just make a mess. The underlying capabilities of the model are the entire ballgame, the "use patterns" aren't exactly rocket science.
- benlivengood 1y agoWe haven't hit the RSI threshold yet and so evolution is so slow that it's usually terminated as not-useful or it solves a concrete problem and is terminated by itself or a human. Earlier model+frameworks merely petered out almost immediately. I'm guessing it's roughly correlated with the progress on METR.
- lubujackson 1y agoAm I the only one who saw in the prompt: > ${SUGESTION} And recognized it wouldn't do anything because of a typo? Alas, my kind is not long for this world...
- floren 1y agoI noticed it and then scrolled through looking for the place where they called it out... sadly disappointed but I don't know what I expected from lesswrong
- deleted 1y ago[deleted]
- SamPatt 1y ago>Claude code feels more powerful than cursor, but why? One of the reasons seems it's ability to be scripted. At the end of the day, cursor is an editor, while claude code is a swiss army knife (on steroids). Agreed, and I find that I use Claude Code on more than traditional code bases. I run it in my Obsidian vault for all kinds of things. I run it to build local custom keyboard bindings with scripts that publish screenshots to my CDN and give me a markdown link, or to build a program that talks to Ollama to summarize my terminal commands for the last day. I remember the old days of needing to figure out if the formatting changes I wanted to make to a file were sufficient to build a script or just do them manually - now I just run Claude in the directory and have it done for me. It's useful for so many things.
- jjice 1y agoI'm very interested to hear what your uses cases are when using it in your Obsidian Vault
- SamPatt 1y agoFormatting changes across lots of notes, creating custom plugins, diagnosing problems with community plugins, creating a syncing program that compares my vault (with publish:true frontmatter) to my blog repo and if see changes then automatically updates the repo (which is used to build my site), creating a tool that converts inline urls to markdown footnotes, etc. Obsidian is my source of truth and Claude is really good at managing text, formatting, markdown, JS, etc. I never let it make changes automatically, I don't trust it that much yet, but it has undoubtedly saved me hours of manual fiddling with plugins and formatting alone.
- Aeolun 1y agoThe thing is, Claude Code only works if you have the plan. It’s impossible to use it on the API, and it makes me wonder if $100/month is truly enough. I use it all day every day now, and I must be consuming a whole lot more than my $100 is worth.
- sorcerer-mar 1y ago
- tinyhouse 1y agoThis article is a bit all over the place. First, a slide deck to describe a codebase is not that useful. There's a reason why no one ever uses a slide deck for anything besides supporting an oral presentation. Most of these things in the post aren't new capabilities. The automation of workflows is indeed valuable and cool. Not sure what AGI has anything to do with it.
- bravesoul2 1y agoAlso I don't trust it. They touched on that I think (I only skimmed). Plus you shouldn't need an LLM to understand a codebase. Just make it more understandable! Of course capital likes shortcuts and hacks to get the next feature out in Q3.
- imiric 1y ago> Plus you shouldn't need an LLM to understand a codebase. Just make it more understandable! The kind of person who prefers this setup wants to read (and write) the least amount of code on their own. So their ideal workflow is one where they get to make programs through natural language. Making codebases understandable for this group is mostly a waste of effort. It's a wild twist of fate that programming languages were intended to make programming friendly to humans, and now humans don't want to read them at all. Code is becoming just an intermediary artifact useless to machines, which can instead write machine code directly. I wish someone could put this genie back in the bottle.
- DougMerritt 1y ago> It's a wild twist of fate that programming languages were intended to make programming friendly to humans, and now humans don't want to read them at all. Those are two different groups of humans, as you implied yourself.
- lelandbatey 1y agoThere is no amount of static material that will perfectly conform to the shape and contours of every mind that consumes that static material such that they can learn what they want to learn when they want to learn it. Having a thing that is interactive and which can answer questions is a very useful thing. A slide deck that sits around for the next person is probably not that great, I agree. But if you desperately want a slide deck, then an agent like Claude which can create it on demand is pretty good. If you want summaries of changes over time, or to know "what's the overall approach at a jargon-filled but still overview level explanation of how feature/behavior X is implemented?", an agent can generate a mediocre (but probably serviceable) answer to any of those by reading the repo. That's an amazing swiss-army knife to have in your pocket. I really used to be a hater, and I really did not trust it, but just using the thing has left me unable to deny its utility.
- abhisheksp1993 1y ago``` claude --dangerously-skip-permissions # science mode ``` This made me chuckle
- BoredPositron 1y agoIf people would be as patient and inventive to teach junior devs as they are with llms the whole industry would be better of.
- sorcerer-mar 1y agoYou pay junior devs way way way more money for the privilege of them being bad. And since they're human, the juniors themselves do not have the patience of an LLM. I really would not want to be a junior dev right now... Very unfair and undesirable situation they've landed in.
- mentos 1y agoAt least it’s easier to teach yourself anything now with an LLM? So maybe it balances out.
- sorcerer-mar 1y agoI think it's actually even worse: it's easier to trick yourself into thinking you're teaching yourself anything. Learning comes from grinding and LLMs are the ultimate anti-intellectual-grind machines. Which is great for when you're not trying to learn a skill!
- jyounker 1y agoYeah, you have to be really careful about how you use LLMs. I've been finding it very useful to use them as teachers, or to use them in the same way that I'd use a coworker. "What's the idiomatic ways to write this python comprehension in javascript?" Or, "Hey, do you remember what you call it when..." And when I request these things I'll try to ask in the most generic way possible so that I then get retype the relevant code, filling in the blanks with my own values. That's just one use though. The other is treating it like it's a jr developer, which has its own shift in thinking. Practice in writing details specs goes a long way here.
- 1y ago
- jasonthorsness 1y agoThe terminal really is sort of the perfect interface for an LLM; I wonder whether this approach will become favored over the custom IDE integrations.
- drcode 1y agosort of, except I think the future of llms will be to to have the llm try 5 separate attempts to create a fix in parallel, since llm time is cheaper than human time... and once you introduce this aspect into the workflow, you'll want to spin up multiple containers, and the benefits of the terminal aren't as strong anymore.
- jyounker 1y agoHaving command line tools to spin up multiple containers and then to collect their results seems like it would be a pretty natural fit.
- jtms 1y agoTmux?
- peab 1y agodagger does this: https://www.youtube.com/watch?v=C2g3vdbffOI https://www.youtube.com/watch?v=C2g3vdbffOI
- sally_glance 1y agoWho or what will review the 5 PRs (including their updates to automated tests)? If it's just yet another agent, do we need 5 of these reviews for each PR too? In the end, you either concede control over 'details' and just trust the output or you spend the effort and validate results manually. Not saying either is bad.
- smallnamespace 1y agoIf you can define your problem well then you can write tests up front. An ML person would call tests a "verifier". Verifiers let you pump compute into finding solutions.
- intralogic 1y ago[flagged]
- CGamesPlay 1y agoIn general, "reader mode". I don't use Chrome but Google suggests that it's in a menu <https://support.google.com/chrome/answer/14218344?hl=en https://support.google.com/chrome/answer/14218344?hl=en>. Many Chrome-alikes provide it built-in (Brave calls it Speedreader), and many extensions can add it for you (Readability was the OG one).
- deleted 1y ago[deleted]
- bionhoward 1y agoAssuming attention to detail is one of the best signs people give a fuck about craftsmanship, isn’t the fact the Anthropic legal terms are logically impossible to satisfy a bad sign for their ability to be trusted as careful stewards of ASI? Not exactly “three laws safe” if we can’t use the thing for work without violating their competitive use prohibition
- alwa 1y agoI can’t speak for their legal department, but their product, Claude Code, bears signs of lavish attention to detail. Right down to running Haiku on the context to come up with cute appropriate verbs for the “working…” indicators.
- deleted 1y ago[deleted]
- konexis007 1y ago.
- jilles 1y agoHow does this compare with Apples or Orange?
- brcmthrowaway 1y agoHow does this compare with Code::Blocks?
- blahgeek 1y agoAsking it to explain rust borrow checker is one of the worst examples to demonstrate its ability to read code. There are piles of that in its training data.
- dundarious 1y agoAgreed, ask it to explain how exceptions are handled in python asyncio tasks, even given all the code, and it will vacillate like the worst intern in the world. What's more, there's no way to "teach" it, and even if there was, it would not last beyond the current context. A complete waste of time for important but relatively simple tasks.
- gilbetron 1y ago"There are piles of that in its training data" Such a weird complaint. If you were to explain the rust borrow checker to me, should I complain that it doesn't count because you had read explanations of the borrow checker? That it was "in your training data"? I mean, do you think you just understand the borrow checker without being taught about it in some form? I mean, I get what you are kind of saying, that there isn't much evidence that they tools are able to generate new ideas, and that the sheer amount of knowledge it has obscures the detection of that phenomenon, but practically speaking I don't care because it is useful and helpful (within its hallucinatory framework).
- deleted 1y ago[deleted]
- dirtbag__dad 1y agoThis article is inspiring. I haven’t had the moment to get my head out of the Cursor + biz logic water until now. Very cool to think about LLMs automagically creating changelogs, testing packaging when dependencies are bumped, forcing unit tests on features. Is anyone aware of something like this? Maybe in the GitHub actions or pre-commit world?
- pjm331 1y agohttps://docs.anthropic.com/en/docs/claude-code/github-actions https://docs.anthropic.com/en/docs/claude-code/github-action...
- citizenpaul 1y ago>automagically creating changelogs, testing packaging when dependencies are bumped, forcing unit tests on features. Yeah now companies that paid lip service to those things can still not have them but pretend they do cause the AI did it....
- rbren 1y agoI’m biased [0], but I think we should be scripting around LLM-agnostic open source agents. This technology is changing software development at its foundations—-we need to ensure we continue to control how we work. [0] https://github.com/all-hands-ai/openhands https://github.com/all-hands-ai/openhands
- ProofHouse 1y agoThis 10000%
- handfuloflight 1y agoBut what do we do if the closed models are just better?
- davidmurdoch 1y agoWait?
- handfuloflight 1y agoAnd get superseded by competitors willing to spend on those models?
- bluefirebrand 1y agoSteal from them shamelessly, the same way they stole from everyone else?
- mjrbrennan 1y agoNot trying to be rude here, but that `last_week.md` is horrible to me. I can't imagine having to read that let alone listen to the computer say it to me. It's so much blah blah and fluff that reads like a bad PR piece. I'd much rather scan through commits of the last week. I've found this generally with AI summaries...usually their writing style is terrible, and I feel like I cannot really trust them to get the facts right, and reading the original text is often faster and better.
- block_dagger 1y agoYou can specify desired style in the prompt. The author seems to like PR sounding fluff while making morning coffee.
- fullstackchris 1y agoYeah I was done at "What happened here was more than just code..." -_-
- WD-42 1y agoI felt the same thing about the onboarding. Like what future are we trying to build for ourselves here, exactly? The kind where instead of sitting down with a coworker to learn about a codebase, instead we get an ai generated PowerPoint to read alone???? Im so over this timeline.
- JohnMakin 1y ago
- aussieguy1234 1y agoI played around with agents yesterday, now I'm hooked. I got Claude Code (With CLine and VSCode) to do a task for a personal project. It did it about 5x faster than i'd have been able to do manually including running bash commands e.g. to install dependencies for new npm packages. These things can do real work. If you have things in plain text format like markdown, csv spreadsheets etc, alot of what normal human employees do today could be somewhat automated. You currently still need a human to supervise the agent and what its doing, but that won't be needed anymore in the not so distant future.
- johnwheeler 1y agoI've actually stumbled upon a novel new way of using Claude code that I don't think anybody else is doing that's insanely better. I'll release it soon.
- throwawayoldie 1y ago"...but the proof is too large to fit in this margin."
- tom_m 1y agoWell, there will always be a job for programmers folks.
- fullstackchris 1y agoGonna be a bit blunt here and ask why hooking up an agentic CLI tool to one or more other software tool(s) is the top post on HN right now... sure, some of these ideas are interesting but at the end of the day literally all of them have been explored / revisited by various MCP tools (or can be done more or less in scripted / hacked ways as the author shows here) I don't know, just feels like a weird community response to something that is the equivalent to me of bash piping...
- AstroBen 1y agoSide note but the contrast between background and text here makes this really hard to read
- jvanderbot 1y agoI can't wait until Section 174 changes are repealed and nobody is financially invested in software from AI anymore.
- tra3 1y agoThank you, finally a realistic take.
- jvanderbot 1y agoIt seems I'm in the vast minority. Post hoc ergo proper hoc indeed.
- eru 1y agoAmerica = world?
- jvanderbot 1y agoIs this meant to say "I don't care because I'm not in USA"? Or "it's not a problem because it's only USA?" Or "don't speak of US-specific situations on this forum because it contains people of many nationalities?" It's entirely possible for a world changing tech to be created and steered to match a unique problem inside one country, and for that to change job markets everywhere.
- eru 1y agoSpeaking of US-specific situations is fine. Or in general, speaking of any specific institutions. I objected to the '[...] and nobody is financially invested in software from AI anymore.' That's a rather dubious claim of a universal consequence for a change that only affects the US. If the comment was 'I can't wait until Section 174 changes are repealed and nobody in the US is financially invested in software from AI anymore.' I would have nothing to complain about. To critique the content more specifically and explicitly, and not just the form: people and companies all around the world have plenty of incentives to invest in AI. A tax change in the US might change the incentives in the US slightly. But it won't have much of an impact on the incentives in Europe, China, etc. And even in the US, even with that suggested tax change, I doubt it'll lead to 'nobody [in the US being] financially invested in software from AI anymore.' Basically, the original comment was hyperbole at best and BS at worst.
- dweinus 1y ago> Is it Shakespeare? No. It's at least decent though, right? > "What emerged over these seven days was more than just code..." Yeesh, ok, but is it accurate? > Over time this will likely degrade the performance and truthfulness Sure, but it's cheap right? > $250 a month. Well at least it's not horrible for the environment and built on top of massive copyright violations, right? Right?
- b0a04gl 1y ago[dead]
- distortionfield 1y agoUnrelated; but I am absolutely in love with this blog theme and color scheme.
- citizenpaul 1y ago>openai codex (soon to be rewritten in rust) Lol, I guess their AI is too good for a redactor. Better have humans do it.
- Syzygies 1y agoNo mention of Opus there or here (so far). Having tried everything I settled on a $100/month Anthropic "Max" plan to use Claude Code. Then I learned how Claude Opus 4 is currently their best but most expensive model for my situation (math code and research). I limited out of a five hour session, switched to their API, and burned $20 in an hour. So I upgraded to $200/month "Max" and haven't hit limits yet. Models matter. All these stories are like "I met a person who wasn't that smart." Duh!
- beigebrucewayne 1y agoAll of this was with Opus.
- luckystarr 1y agoI recently investigated some problematic behaviour of both Opus 4 and Sonnet 4. When tasked to develop something more complicated (broker fed task management system, staggered execution scheduler) they would inevitably produce thousands of lines of over engineered, unmaintainable garbage. When Opus was then tasked to simplify it it boiled it down to 300 lines in one shot. The result was brilliant. This happened twice. Moral of the story: I found out that I didn't constrain them enough. I now insist that they keep the core logic to a certain size (e.g. 300 lines) and not produce code objects for each concept but rather "fold them into the code". This improved the output tremendously.
- jumski 1y agoGreat article! I have similar observations and techniques and Claude Code is exceptionally good - most of the days I'm working on multiple things at once (thanks to git worktrees) and each going faster than ever - that's really crazy. For the "sub agents"thing, I must admit, that Claude Code calling o3 via sigoden/aichat saved me countless of times! There are just issues that o3 excells at (race conditions, bug hunting - anything that requires lot of context and really high reasoning abilities). But I'm using it less since Opus 4 came out. And of course its none of the sub-agent thing at all. I use this prompt @included in the main CLAUDE.md: https://github.com/pgflow-dev/pgflow/blob/main/.claude/advanced_ai.md https://github.com/pgflow-dev/pgflow/blob/main/.claude/advan... sigoden/aichat: https://github.com/sigoden/aichat https://github.com/sigoden/aichat
- _1tem 1y agowait what? how do you work on multiple things at once with git worktrees?
- pjm331 1y agoI never had any reason to use it before claude code et al. So I also wasn’t aware Commands for working with copies of your entire repo in a new folder on a new branch https://git-scm.com/docs/git-worktree https://git-scm.com/docs/git-worktree
- noiwillnot 1y agoThis is amazing, I had no idea about this, I have been cloning my repo locally for years.
- jumski 1y agogit worktree uses one repo to laid out multiple branches in separate directories. git worktree add new/path/for/worktree branchname I now refuse to use git checkout to switch branches, always keep my main branch checked out and updated and always use worktrees to work on features. Love this workflow!
- rikschennink 1y agoI tried to read this on mobile but the blinking cursor makes it impossible.
- beigebrucewayne 1y agoRemoved it! I agree it was distracting.
- hoppp 1y agoFirst time I heard about marp, very handy tool