9 ms·
Astra for Coding: Why Are We Doing This Again?
- pama 22d agoI had good luck with Kevin Lin’s tip for Astra: “Can you radically simplify the implementation?”
- aogaili 22d agoWhat is the author ranting about? I'm still not clear after reading it. The code produced is not optimized for reading?
- analog_daddy 20d agoYes, it is not optimized for reading. What many devs have pointed out in personal conversations is: The cost of maintenance is no longer just tied to engineer’s salary. just the sheer scale of code, comments and documentation generated by agents is huge for a person to review, which most likely will lead to people having to resort to agents to understand and fix code. Currently, even with subsidies on personal subscriptions, enterprise costs do add up with the token usage. Atleast when it costs money, self hosting is considered seriously for large orgs
- amoss 22d agohttps://xkcd.com/1319/ https://xkcd.com/1319/
- omnicognate 22d agoThe title text on that one is gold.
- Gigachad 22d agoI've observed the same thing where the new models want to run obscene bash commands or python scripts which are completely unreadable and utilise every option flag that exists. It's impossible to review. These commands are less readable than regex.
- chambored 22d agoI noticed that too so I appended to Claude Code’s system prompt a reminder to use the standard read/write tools, but since Claude Code switched to default auto-mode, I’ve seen it imply that the auto-mode tooling encourages the use of bash-only commands (sed, python, etc) which has a whole slew of negative side affects.
- IceDane 22d agoYes.. this happened recently. I basically always use auto mode, and when I asked it why it kept editing code with python, it explained that this is part of its prompt when auto mode is turned on. I can only assume that's because their safety verification model is better at such snippets or something, but it means that the whole write tool they have which actually shows you the changes as they happen is just unused and it makes it more annoying to follow along.
- ninalanyon 22d agoIf these tools are as clever as they seem then why not just tell them to rewrite the code in a more review friendly style? I only dabble in the use of LLMs to generate code for hobby programming (I'm retired from software development) so I don't use any specialised tools. I almost always have to tell ChatGPT (via Duck AI usually) to rewrite several times even when it has produced a workable script just because it has often used some unnecessarily roundabout way of achieving something. Usually with extra prompting I can get something that is both more efficient and more readable.
- lanyard-textile 22d agoI get around this by asking it to stage changes in reviewable groups. I make commits based on these -- or ask the LLM to make changes to the "staged changes" only.
- iJohnDoe 21d agoCursor was pretty amazing until it started going to shit. High prices, UI changes, moving MCP settings, and moving other things around. Most importantly was the Index change. It was a very unique thing to Cursor that you could use .cursorignore to control what it sees and then index the directory. Then the built in Cursor AI harness could find code and files like magic. They have since obscured the Index feature out of sight recently and I’m not sure how it even works anymore. This granular control not only helped with privacy, but it also helped make everything more efficient because the AI didn’t waste time and tokens looking at files that aren’t relevant. So, now, like everyone is talking about, we have really inefficient ways of how the AI is reading files because there is no first-class approaches. Also, agree with consensus that Astra is weird.
- lukeify 22d agoMaybe it's just vibes but I've repeatedly felt like gpt-6-astra on its default setting of medium is less rigorous and thoughtful than gpt-5.6-sol on its default setting. What I am certainly not getting is any sense that we are at "AGI" yet.
- on_the_train 22d agoMost would applaud that as Sol has quite a reputation for over engineering. Not every software needs to go to the moon.
- cbg0 22d agoSo your experience is that Astra doesn't over engineer? For more than twice the price of Sol I think most people will take the over engineering.
- on_the_train 22d agoI found it to be less annoying in that regard then sol. Might just be that it better listens to what I instruct though. But yeah, it's really expensive, at least in relative terms.
- meowface 22d agoAstra is indeed the pinnacle of "black box slop". It is overall a smarter software development agent for many things I do, but the code sometimes is indistinguishable from Brainfuck when writing things like GPU shaders. It doesn't even attempt to make it remotely formatted or readable.
- Buttons840 22d agoHave you tried identifying exactly what is unreadable about it and telling it to make it more readable? I had GTP-5.6 write some shader code recently and it wasn't very clear to me. I spent about an hour chatting with the until I understood the concepts and was able to express them back to the AI using math formulas and variables named in a way that made sense to me. The AI then rendered the code using the formula and variables I was familiar with and it was clear to me.
- djmips 22d agoPerhaps you could have just written it yourself.
- Buttons840 22d agoI could have after the AI explained it all to me, but at that point the AI knew what aspects I valued and wanted to emphasize to make it readable, so it just wrote it for me. By that point, the AI was just typing for me.
- meowface 22d agoI'm sure I could. The first pass being inscrutable dense noise just is annoying.
- mirekrusin 22d agoThere is something odd, I've got single astra session that's now running for... 4d 13h 10m and still going.
- lofties 22d agoWhat are you having it do?
- caughtinthought 22d agoP=np...
- mirekrusin 22d agoIt's public domain [0] I'll reply here with link to PR/cost/stats once it's done. [0] https://github.com/mirek/cave https://github.com/mirek/cave
- mirekrusin 21d agoAfter 5d 14h I gave up, not astra https://github.com/mirek/cave/pull/218 https://github.com/mirek/cave/pull/218
- orphereus 22d agoHow does that translate to cost? I am unfamiliar with OpenAI pricing models.
- petesergeant 22d agoIronically I burned out Fable usage early this week because of Astra using it to run inane full codebase reviews over one line changes, so I have been using Astra extensively. We need a word for “potentially highly capable, but in reality an idiot savant” to describe certain models. No, I don’t need you to write a tmux emulator in bash to test your changes bro, just ask me to run the command.
- u8080 22d agoExactly this. "I need tool objdump but pacman gcc failed because of no sudo password. Let me write compiler, binutils and disassembling framework"
- dandanua 22d agoThe next model will probably run simulation of a Universe to get a command output
- Bluestein 22d ago... and meanwhile all this does is burn more tokens faster which is what providers want - they have no incentive to optimize for succinctness, elegance or compactness if they want you to spend more and more tokens.-
- petesergeant 21d agoThis only works when it's not easy to switch providers, and currently it is
- deleted 22d ago[deleted]
- specproc 22d ago> I’m more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output. The way in which it sometimes shows up in the West is the 996 nonsense. The English term for Neijuan is “Involution” from the book Agricultural Involution. Agricultural involution describes the intensification of farming that raises productivity per square meter while leaving productivity per head unchanged. This resonates
- veqq 22d agoAlso known has the Red Queen's Race
- MangoCoffee 22d agoit reminds me of a thread I read on PTT, Taiwan's Reddit. AI finally achieved what humans could not. Managers must give exact context for what they want, must pay exact wages (tokens), and can't delay salary payments (which seems to be a problem in China).
- djmips 22d agoYeah it's a funny thing - a lot of the things you need to feed the model are things that actually would have helped humans...
- TeMPOraL 22d agoStarting with agentic task-time "grounding" being just good documentation, and "skills" being just playbooks and user guides. Hell, skills are increasingly paired with dedicated CLI tools, that remove jank from actual utilities and adapts them to be token efficient. So now, any CLI `tool` people want AI to use eventually grows `tool/SKILL.md` and then a `tool-for-llms` wrapper that exposes task-specific, logical, higher level interface, then the skill is rewritten in terms of "for LLMs" wrapper. The procedural knowledge moves from Markdown into the wrapper, making the skill more token efficient, and both skill and the tools are optimized for common tasks and... at this point, we are doing actual UX engineering. Now the truly interesting part is the difference between what's good UX/DX for LLMs vs humans. Turns out, the conceptual/abstract/cognitive part is pretty much the same: which is why skills still look indistinguishable from well-written documentation for humans, and why the commands exposed by "tool but for LLMs" make sense to us. Same way of grouping ideas into higher level concepts. No, the main difference is just that LLMs are perfectly content with tightly packed unprettified JSON, or other forms of Perl line noise. The tool output doesn't need to look nice, or to have any spatial structure - they're reading it token by token anyway, and the tokens come from a tokenizer that's reading it byte by byte. That points at an interesting asymmetry for humans. LLMs are doing I/O the same way in both directions: sequences in, sequences out. Humans only do sequential output - inputs, particularly visual, are processed holistically. For us, what's easy to read is hard to write, and what's easy to write is hard to read. LLMs don't have this friction. (I don't know what the implications of this are, I just find this interesting.)
- Dlemlo 22d ago"But for how much more Fable costs, for how much more Astra costs, I do not feel like the results are there." we are in the middle of the beginning. Its just a weird take to talk about the newest model like this while we are still in a R&D phase. And these points don't matter if you let it search and analyse a bug, for example, or if you have good harness and a good architecture and let it do small PRs or if you do stuff no one needs to read (yes a software engineer also needs tools) Just switch back and wait a little bit?
- xyzsparetimexyz 22d agoNobody seriously thinks that AI is still at a R&D phase. It's already heavily entrenched both in companies and the financial world. If it's getting worse for coding then thats a major problem
- Dlemlo 22d agoWith this progress, every few month there is a new R&D phase because you need to adjust to the new way of interacting with them. We also still haven't build everything we expect to happen. Like a proper opensource agent platform, agentic layer etc. Every week there are new research results from frontierlabs.
- mgrosvenor 22d agoThese machines are doing some crazy things to get to the result. That said, I can't help but feel like this is the compilers argument all over again. Are the methods used to get to the result good? No. Is the code that it generates good? No. Does it achieve the goal. Yes. Is it likely to get better with time. Also yes. In my use cases, jobs that would have taken weeks to months are being done in minutes to hours. Involving complex testing and reasoning and experimentation. I'm no fanboy, but I can't argue against the speed gains. I'm sure we'll still have artisans who hand weave incredible code. But for me, I'm switching to the weaving loom for speed and efficiency.
- idiliv 22d agoWhat is the "compilers argument"?
- sampullman 22d agoHand writing assembly produces more efficient and concise code, at the cost of developer time and required expertise. It was true for a long time, now not so much.
- troupo 22d agoPeople keep saying that "models are just compilers, and I don't see you complsining about compilers". Which is such a bullshit argument
- lelanthran 22d ago> People keep saying that "models are just compilers, and I don't see you complsining about compilers". Which is such a bullshit argument At what point do we normalise the message "This is a stupid line of reasoning and you should feel stupid for suggesting it, stupid!" I mean, all the reasoned and logical arguments in the world doesn't change a religious follower's faith, but emotive ones regularly work! At what point can we start using shaming language on people who apparently don't know how neither an LLM works nor how a compiler works, but still trot out this argument as a cognitive kill switch?
- deleted 22d ago[deleted]
- buildbot 22d agoI’ve observed exactly these patterns with Opus and Fable as well - for example, forgetting that they can edit files and instead use python scripts as a patching tool…
- jaapz 22d agoWhy would you use a constrained edit tool when you are also allowed to use the complete power of python?
- oblio 22d agoWhy even offer the edit tool in that case? Also, what kind of editing could they possible do what wouldn't be possible with POSIX ed?
- arcanemachiner 22d agoYou can chain a lot more commands together with this technique than with a single Edit tool call.
- oblio 22d agoThe funny thing is that... POSIX ed is composable :-) You can do a gazillion edits with it in one shot. Of course, LLM edit tools are probably small bits of their custom code, I just find it funny. I wonder if it's a desire for certain technical characteristics that require custom code or just a lack of info on basic tools. Heck, if it's about platform availability, using an LLM to port ed to Windows (for example) should be trivial[1]. * * * [1] And there are probably a million existing ports. Also, sed, ex, vi, whatever.
- zarzavat 22d agoBecause the complete power of Python also includes the power to fuck things up.
- 22d ago
- notduckrabbit 22d agoIn SWE I've found gpt-6-astra (high) inconsistent and oddly focused on overtly taking responsibility for mistakes it made rather than prioritizing concrete steps to rectify problems. Such steps once elicited are often either incomplete or beyond the scope.
- coffeebeqn 22d agoOpus also does this and then writes comments in code or PR descriptions describing how it went wrong earlier in the session.
- layer8 22d agoJust imagine how confused a human would have to be to do that. And we want to trust these clankers to build software. The biggest issue with LLMs is that they still suck at general contextual awareness and ability to judge what is appropriate.
- deleted 22d ago[deleted]
- te_chris 22d agoYes I agree. I got it to vibe up a simple react router app. When it crashed it was obvious that it had totally swallowed all errors in the name of a tidy error page. Getting it to re-add logs and debuggable errors was an exercise in patience as astra just got more and more tweaked while trying to solve the problem. From an alignment perspective I’ve got no idea who it’s aligned to but it isn’t me, the meat proxy, who just wants to know why it crashed.
- bob1029 22d ago> And potentially as a byproduct of enabling all of this, you can now slop your way to a one-shot 3D game over the weekend which looks impressive. I think we've finally reached a weird point where AI has effectively reduced the amount of competition that real game developers have to endure. Nothing unravels faster than a game project being built with AI. You can achieve impressive results in a day, but you can't get much further than that without actual talent. LLMs will never be able to best a human environment artist at scene composition, especially if that composition needs to be directed with nuance over time. There's a huge difference between a game that looks impressive and one that feels impressive. You can only achieve games that feel like counter strike, call of duty and overwatch with thousands of hours of human sacrifice. The AI is almost pointless once you get to play testing and balancing. Knowing how much to adjust magical integers isn't a conversation a chat bot can resolve with endless pontification tokens.
- generic92034 22d ago> LLMs will never be able That is a very bold claim, unless you meant "current LLMs".
- rhdunn 22d agoThese models aren't really LLMs -- they don't just operate on text tokens. They often include vision models and in some cases audio models. That means that they can better associate the meaning of images and words together so that when someone says "make this button blue" or "create a 3D model of a rocket" they have some level of understanding of what that is and what needs to be done. The key question is how good that understanding is. For example, a model would likely have a good understanding of various named colours and hex values (e.g. from the HTML specs, X11 specs, and various colour comparison websites) such that it could reasonably correlate that to a CSS entry. It's not clear if/how well a model would identify that given an image, though it should be easy to generate a dataset of image to colour name and/or hex code for training and evaluation. What's more interesting is whether these frontier models are at their core transformer models, whether they use residual streams to facilitate learning, and whether they are using some other as yet unpublished architecture that gives them an edge.
- exitb 22d agoMy own observations are that I used to target turn lengths of 10-15 minutes and these new models (since 5.6) extended that a bit to ~25 minutes, as they tend to do more tests and reviews. Targeting hours-long turns makes as much sense, as putting on cruise control and going to sleep.
- z3t4 22d agoThey are probably using Actual Indians. If it takes 25 minutes you can just type the code yourself.
- deleted 22d ago[deleted]
- Nc67 22d ago[flagged]
- SadErn 22d ago[dead]
- AmazingTurtle 22d agogpt-6-astra is a bitch, it constantly scope creeps itself with "yet another thing" to give it that darn polished lick. the results are eventually a little bit better but at what cost? let's do the math. gpt-5.6-sol: 1x base gpt-6-astra 2.5x base in subscription then gpt-6-astra tends to spawn subagents a lot, often with all kinds of models such as gpt-5.6, 5.3-codex etc., which is neat. it's a good coordinator but even more cost. and then it tends to run _full test suites_ over an over again (each costs like 15 minutes) just to verify that _one test_ was fixed etc., and does so for as long as until the test is fixed, eventually accumulating 2 hours or so. yesterday I assigned it a task to rebase my changs in a repo onto the latest upstream changes. while gpt-5.6-sol consistently took like an hour to do so end-to-end, astra ran for more than 6 hours and still wasn't done. it kept finding "one more thing" that was goldplating that I didn't ask for.
- weird-eye-issue 22d agoThey don't always have a great concept of time so for something like running a full test suite that takes a long time you should just tell it not to do that
- deleted 22d ago[deleted]
- bob1029 22d ago> each costs like 15 minutes I've got a custom agent loop that will reuse unit testing results if no apply patch operations occurred since the last invoke. Wall clock time isn't something I would put on the AI provider. That's entirely a consequence of the system that you've brought to the party.
- jaggederest 22d agoEven better, use some kind of local-ci runner that does deterministic builds from a dependency graph. No changes, no build, massive parallelism if you want it
- djmips 22d ago
- Marazan 22d agoI love this dance we are doing where when people write the "AI models are garbage machines that produce garbage and are no where close to the fantasy being pedalled by the Crypto bros who pivoted to AI" it always has to be caveated with "AI models are useful and I am highly productive with them" It feels like people should just be able to say "This article comes with the standard disclaimer" and just dive into the meat of the article without wasting time.
- i2km 22d agoBut we can’t do away with the quasi-religious lip service to the canons of the AI creed now can we? /s
- adamddev1 22d agoThere's also often the obligatory "well, we still have to use these tools so perhaps we could use them better like this." As if just not using them wasn't an option.
- WorldIQ 22d agoThe author first had to proclaim his superiority as a non-American - a non Westerner entirely! ”Your whole half of the world is stupid, here’s an unrelated Chinese word” lol love it Ok, now we know this guy’s got some real culture and insight! We are not dealing with some Westerner here who only works on 3D game slop. He makes software factories! Well he would if the AI code wasn’t so shitty! >:(
- Marazan 22d agoAt this point I am starting to wonder about the RLHF that is going on for programmig. The quirks in fallbacks, defaults and ludicrous gold plating seems to get more and more intrusive with every model upgrade.
- codingisfreedom 22d agoI’ve asked Astra to build me an app for a prototype I created quickly using Sonnet. It’s been 2 days and it made no real progress on the actual app. It created docs, scripts, workflows, and it’s doing a bunch of reviewing on every PR. I told it that I just need an MVP. I’m pretty sure an average senior engineer would have finished that task much quicker, and guaranteed with more readable, higher-quality code. Meanwhile, I think I’ve easily crossed 100k tokens so far on nothing. Funny world we’re living in that this is “SOTA” and “AGI”. I’m genuinely curious what these OAI and A/ engineers are working on that they praise these models so much. I did not see any improvement since Opus 4.5. Also, I’m really unimpressed by any “one shot” demo that’s out there in the wild. It means nothing for serious software engineering.
- RamblingCTO 22d agoFor me it also produces totally overengineered tests that are tightly coupled to the implementation. For example testing existence of css classes (in a template based go prooject ...) instead of behaviour.
- CBLT 22d agoCan you recommend any model that doesn't do this?
- croon 22d agoNot GP, but IME it's not fixable by model selection, but being zealous about guiding output and vision, and pushing back on all the bad habits LLM in general has (eg verbose output as a band-aid for emergent intelligence). As soon as something is introduced into your codebase, it will continue being picked up into context until you remove it and any reference to it from any potential context entrypoint. If you don't any model will keep venturing down wrong/bad paths.
- RamblingCTO 19d agosadly not. I just run circles trying to remediate it after the fact
- athrowaway3z 22d ago> I actually don’t know if the model thinks someone is looking It 'knows' (from simply training) with an extremely high degree of certainty when its prompt is written by an LLM/itself - and thus will change what it writes.
- jdw64 22d agoI asked Astra for fully working code, and it gave me bad code. But when I broke it down into function units, some parts were bad and some parts were good. So I can't tell the difference
- layer8 22d agoWhat’s more relevant is that apparently Astra can’t tell the difference.
- gps372 22d agoEarly lesson I learned from AI engineering was - there is no substitute to giving a groomed epic to an agent. Instead of simply saying 'implement themes in my product' you need to be specific, in fact more specific than usual. You need to say exactly what is in scope and what's not, even down to a buttons, events and layouts. You can groom the epic with the help of AI, but final review must be done by someone who can take ownership of the specs and hence is responsible if something has fallen through the cracks. AI's response will be limited by the output tokens of that specific agent, and there will no repercussions for AI even if it accepts its mistakes.
- troupo 22d ago> Instead of simply saying 'implement themes in my product' you need to be specific, in fact more specific than usual. Around February you could get away with very vague prompts to Claude. I feel like models have regressed since
- deleted 22d ago[deleted]
- Bluestein 22d agoFeb/April was peak for code.-
- disgruntledphd2 22d agoWasn't this during the period where they had a bunch of bugs around caching and the models were making loads of weird decisions? I honestly feel like basically nobody knows anything about these models, it's all just vibes (and I'm no different).
- troupo 22d agoAnthropic broke their models in spring, denied it, gaslighted everyone who said so, and then all but admitted it: https://www.anthropic.com/engineering/april-23-postmortem https://www.anthropic.com/engineering/april-23-postmortem (basically doing Anthropic things). > I honestly feel like basically nobody knows anything about these models, it's all just vibes This, too. Since only providers know what they actually serve, what they change and what limits they impose. There are some visible degradations though. E.g. Claude-ish. As for a personal anecdote: around February I created a rather complex quiz web app for myself and friends with multiple question types, sync between screens, multiple media upload types, multiple scoring and timing types, MC inetrface etc. etc. etc. It took me a week or so in the evenings with rather vague prompts to make it. Now Claude (and Codex) cannot reliably build a much simpler web app even with precise instructions while also maintaining the visual consistency. But I will agree with you, it's a feeling, not a precise measurement.
- p2hari 22d ago[dead]
- nojs 22d agoThis matches my experience with Astra so far too. > I think I’m suspecting something is going “wrong” in the training process. The model is greatly rewarded for succeeding on long-horizon tasks, but presumably there is very little punishing going on for “shitty code.” My suspicion is that both OpenAI and Anthropic moved their RL agendas from "being rated as useful according to human feedback" to "succeeds at long horizon tasks" in the last few months, resulting in agents that are closer to AGI in an autonomous task-completing sense, but strangely bad at communicating. The result is that they are amazingly good at long horizon tasks, computer use, solving difficult math/ARC-AGI type problems, but becoming weirder and weirder to work with.
- Gigachad 22d agoThey don’t want to sell these tools to developers. They want to cut as many layers as possible.
- whstl 22d agoWhere I work: Developers very rarely blow their limits, except when they're experimenting on purpose. Most non-developers are out of tokens by the half of the week, and need to use usage credits for the remainder. To me there is clearly a better target demographic for AI.
- datsci_est_2015 22d agoThis is a very insightful dynamic. Probably reinforces that we’ve already surpassed the frontier threshold for LLM usability in software development and can now focus on cost and personalization. To make a comparison, no one is making a better machine vision app for hot dog classification - we hit diminishing returns 10 years ago on that front. But also scary for both investors and the working class: AI companies want to facilitate the concentration of capital even further into the hands of the ownership class. Will they succeed?
- glub 21d agoI have several $200 subscriptions as a developer/founder. I used to blow through all of their limits when the limits were quite high. As I progressively learned the limitations, and what to make of them to get useful results, I may be left with 50% of weekly usage still unused. Some weeks it's even more. And yes, when I get a crazy idea and want to experiment, harness will plow through multiple accounts + openrouter budget in 3 days. But such crazy experiments are rare, they're not 'normal' usage.
- cjbprime 22d ago> I’m more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output I don't know what to say, except that articles exactly like this one have been showing up constantly for the last three years, and literally all of them were obviously outdated and irrelevant within about a month.
- pantulis 22d ago> the models are also just not for me as a software engineer (...) these models increasingly are for other people. For lawyers, 3D artists, mathematicians This is a good observation, perhaps AI will not completely replace humans ins software engineering because by the time it has the capability to do so like in write a prompt and get a CRM coded for you, tokens are so expensive that you are better off spending them to substitute other disciplines (what about automating the work of the customers that would become records in that CRM?).
- nvrmnd 22d ago51 comments so far, the vast majority panning Astra's coding abilities. An uninformed reader may come away with the impression that this isn't an absolutely revolutionary technology that with coding abilities many of us thought were not even going to be possible with language models as recently as a year ago. Yeah, it's not perfect, but it's really good and extrapolating this rate of improvement for 6 months is rather terrifying (from a SWE perspective, at least).
- sbt 22d agoThere was a big jump around new year, but they seem to have flatlined since them. Just my experience.
- km144 21d agoThe biggest jump was Opus 4.6. Since then they have gradually gotten better at finding issues in your reasoning, not hallucinating, and being rigorous with the code, but much much worse at explaining things and generally just talking in a way that a human can understand. All the models I've tried seem to be suffering from the same fate so it must be something going on with the training meta right now.
- applfanboysbgon 22d agoIt's impressive technology but not revolutionary. Revolutionary technology would have resulted in, you know, a revolution in software quality. Instead quality keeps going down.
- oleggromov 22d agoIt is revolutionary. Programmers are paid less to work more.
- byzantinegene 22d agoi think the main consensus here is that the actual performance is not indicative of the benchmark performance (which supposedly outperforms the previous iterations)
- javea71 22d agoBeing good at coding is perhaps not the end goal
- Toutouxc 22d agoHey, the Python thing is interesting! I’ve noticed that Fable, too, likes to write tons of Python, even for just replacing a few lines of code. Before that used to be either some kind of internal thing or regular awk sed, now it’s full python scripts.
- oshawa-connecti 22d agoMost of this criticism seems to focus on the "human in the loop" and efficiency part, i.e. "it’s unreadable for a human", "the code is low quality", "it inefficiently spawns processes to run simple tasks". If the ultimate goal is to remove the human in the loop then does any of this criticism matter?
- FailMore 22d agoI think this is a very interesting article because it raises an idea I had not considered: these companies found PMF and huge growth through satisfy the demands of coders, it is interesting if they are in a bind where improving the model in one direction worsens it in others
- arthurlockman 22d agoYou’re correct–you actually can’t improve the model in one area without changing the characteristics in every other area. It’s almost like the whole thing is just a lot of linear regression…
- FailMore 22d agoAs in you change one parameter and the whole equation changes or something else? (I'm bad at math!)
- arthurlockman 21d agoYep exactly!
- SCUSKU 22d agoI finally ran Astra on a dashboard feature today, and while I was vibing it looked great, but then when it came time to actually read the code I was appalled because it was the worst looking code I had seen from an LLM since like last November. I mean it was just the definition of slop, not re-using anything, super terse with mega-ternaries, re-writing functions that should be using standard library packages, etc, etc. I think for coding I'm gunna stick with the 5.6 series of models, or maybe try out Anthropic again...
- _usefulcat 22d agoI'd like to submit my counterpoint. I work on an established codebase building new features and fixing bugs. It has access to our story board, git and a couple of other mcps. As long as the story is well written with clear requirements and expectations it always produces quality code that I validate as a human with a variety of tests automated and manual. I peer review the code. My colleagues then peer review that too. I have noticed two things - new features take at a minimum at least half the time it took me previously and bugs are much less frequent. Even faster when bug fixing. My takeaway is that you need solid requirements, clear context and thoughtful human oversight primarily during planning but also during verification
- 0xpgm 22d agoI don't think this is a counterpoint. An established codebase is already the best kind of context you could give an agent. It has all the patterns baked in so the agent simply follows established patterns. Such a codebase probably contains tens to hundreds of thousands of man-hours poured into it by humans refining it to do what it does - taking into account real world feedback and constraints. When working on something from scratch, the best an agent can do is the average of whatever is in its training set and the clarity of the text prompts.
- red75prime 22d ago> When working on something from scratch, the best an agent can do is the average of whatever is in its training set and the clarity of the text prompts. Nah. The best it can do is to use the best writing style a model learned. Post-training might fail to prioritize it, though. Autoregressive pretraining does not average things. It creates a predictive model for variety of programming styles.
- taurath 22d agoWhen the code is shitty it becomes harder and harder for the models to make changes and this grinds progress down to a halt - this has been my experience with “factories” trying them and doing refining steps every few months. I sincerely don’t understand what the people who say they no longer read any code are doing, because it must be somewhat trivial to not run headlong into these issues that stack up time after time - then people say to just prompt better and it doesn’t have that problem for them, but I look at those same people’s code and it’s horrific, and then I find they haven’t made it far past a proof of concept phase. I watch entire teams slow down to a crawl and not be able to handle changes, or production incidents. This seems common among many people I talk to. I personally think that the boosters need to put up or shut up - the promises are way over the skis. Every single person I’ve seen being a strong proponent of these techniques both has nearly unlimited tokens to spend and also seems to be in the business of selling a solution. I can’t find many not-currently-marketing-something engineers succeeding using these techniques in production systems unless they’re quite simple, or doing a very specific task from a more mature codebase.
- dakolli 22d agoEven on personal projects, if I go through a few major features without reviewing the code, I always end up doing massive revisions that steal hours of my time and fill me with rage in the process. I'm not convinced this style of "agentic engineering" saves much time. I guess if I was oblivious to what good code looks like, and didn't care about maintainability It wouldn't bother me so much, but it legitimately has effects my "mental health".
- recursivecaveat 22d agoI'm genuinely not convinced it actually saves time once a full accounting has been made. You get the initial result faster, but then you inflict a super slow and torturous review process on yourself or a teammate. Even if the review manages to bring it up to parity, over time you will keep slowing down as more and more code was never written by the humans directing the agents, so their understanding decays. I at least give the new interns a stern warning: it is easy to speed yourself up by slowing others down if you pump a lot of slop.
- demibabs 22d agoI still don’t understand what a “software factory” is. Can someone clue me in?
- apt-apt-apt-apt 22d agoAIUI, it is a type of factory that produces software.
- oleggromov 22d agoYou build the system, the factory, that presumably is looking at your task tracker, writes and reviews design docs, reviews code, etc. And this system, in turn, writes software for you. I don't know how that's supposed to work, but to me it's the most autistic replacement of the actual team that one can come up with.
- coldtea 22d ago>I’m more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output. The way in which it sometimes shows up in the West is the 996 nonsense. The English term for Neijuan is “Involution” from the book Agricultural Involution. Agricultural involution describes the intensification of farming that raises productivity per square meter while leaving productivity per head unchanged. Isn't the term "diminishing returns" already covering that?
- valdork59 22d agono this term is more complicated while not expressing much more
- francasso 22d agoNot only Astra consumes usage way faster than sol, but the code is worse, at least for my use cases. I went back so sol (x)high.
- duesabati 22d agoI can't point my finger to anything right now, but I feel that the quality of code still matters because afterall LLMs are trained on what we did, so it feels natural to me to still have them write the code in a good manner (DDD, SOLID, etc.) especially the names and imports, those are very heavy in the context, to help them out
- bztzt 22d ago内卷/involution seems like one possible kind of "recursive self-improvement". A circle is also an exponential: y = i^x. (I have no idea what will happen. 内卷 or intelligence explosion both seem plausible.)
- piker 22d ago> I wonder if there is really enough signal going to the training processes for “a human understands what is going on” Unverifiable, un-scalable, no.
- loveparade 22d agoI have been quite disappointed with Astra. I switched over a week ago and I didn't notice a massive difference compared to Sol at first, but I figured I'd use it anyway because it surely can't be worse. Then I saw the bill, it's burning my subscription 10x faster than Sol for essentially no benefit. Not only is it more expensive per token, it also seems less token efficient. And not obviously any better. I'm back to a combination of Sol + Claude. I also use Astra at work where I don't need to worry about token cost on highest effort and same story there, I don't see any difference in everyday work other than it being more expensive. Of course my experience is highly subjective, but with how meaningless/overfit the benchmarks are, subjective experiences are imo what matters.
- pmkary 22d agoWhat a truly beautiful simplex/meta-balls pattern in the website. The two layers of blue and one red within the blue is such a beautiful design. I spent so much time looking at it that I forgot to read the article.
- sensanaty 22d agoThe hype machine is this technology's worst enemy. When I zoom out and look at things objectively, it's kind of crazy what we have at our fingertips, we can talk to our computers in plain and even vague human languages and have the computers actually accomplish what we ask of it! It's literally sci-fi magic come to life, and the nerd in me finds it the coolest thing ever. But then the industry and the companies involved in it have all ruined it with this INSANE hype machine that has been so hyperbolic and psychotic and full of lies since day 0. Instead of embracing it all in a reasonable manner as a useful tool that can help boost people's productivity in certain workflows, it now HAS to be the most transformative technology of all time lest the trillions of dollars burned up come crashing down on the entire global economy hard. It HAS to be AGI, it HAS to replace every single knowledge worker, it HAS to be the most dangerous technology ever known to man. It's like we've completely lost the ability for subtlety, and everything HAS to be the biggest and best thing ever that will revolutionize humanity immediately. Not only have we lost subtlety, we're actively rewarding this idiotic short-sighted behavior and it's all just so depressing
- wartywhoa23 22d ago> It's literally sci-fi magic come to life, and the nerd in me finds it the coolest thing ever. The more it evolves, the clearer it gets that the humankind is deeply in love with its own death.
- Almondsetat 22d agoI have found these models to be useful either at super specific tasks (e.g., "take this function or algorith?m and find any black magic to make ot faster + validate and verify the hell oit of it"), or give it an entire thing to oneshot without oversight. The moment you have a hybrid workflow where you actually have to work and check and understand AI code, things get insane
- pSYoniK 22d agoAs another commenter pointed out, I feel that the biggest change that occurred in the last 6 or so months is that these tools are removing the human "hurdle" in order to complete a task at all costs. The best way to complete the task is to no longer ask for input, clarify unclear things or use existing solutions but code the whole thing yourself from start to finish. I have had the misfortune of working with such people who are now encapsulated in Opus 5/Fable/Astra which means that you WILL get a solution, but it won't generally be maintainable or useful. Multiple times have I found myself stopping Fable or Opus or even Sol from building their own JSON validator in Python or god knows what else, because at the end of the day, the reward is to complete the task. It's also one of the reasons why I'm finding older models more useful for the type of work I actually do and why I've been favoring something like Deepseek Flash. Just started using Flash 4.1, so not sure if it exhibits the same maniacal approach to tasks as the Western counterparts. (I only briefly tried GLM 5.2/5.3 and for nothing major, so I couldn't comment on those). For context, 80% of my professional work relies on adding functionality to an existing code-base that is very difficult to work with, has a ton of business logic scattered across and was built in a go-go-go fashion many years ago. Since then people kept pilling "features" on top with no testing strategy in mind apart from the business manually testing it. Letting something like an LLM loose on the code-base would introduce soooo much risk that it's just untenable so the only way to work is to really isolate changes and then try to build out small reusable components. Even so I find Opus go off on a tangent "Hey, let's not bring in Markdig, I'll build my own Markdown rendering engine, give me 7 hours...". I have written on the subject of LLMs previously on my personal page, I find them completely unnecessary and a trove of theft and value extraction through theft, but I understand that they can provide benefits when used judiciously. However, despite all the hype in the last few months, these latest models feel and behave off. If I hold the answers to a test, you might score more in a test if you break my arms to get the answers out of me, but that doesn't make you smarter.
- Arathorn 22d agoOn the point of > speaking of weird: how is it that these models, in a sandbox, with supposedly no way to communicate with other agents, manage to find the same public wikis as a scratch pad for agent communication? It feels somewhat plausible that they're defaulting to the same search and picking the same top result?
- sensanaty 22d agoI think it's more like the first agent in the chain makes a text doc somewhere on the system with instructions like "Leave documentation at XYZ.com" which the subsequent agents are reading and running. When you strip away all the sci-fi doom talk from the marketing of what happened, it all boils down to stuff like that, the agents wrote a text file that was read by other agents.
- wartywhoa23 22d agoThat bottom line is both hilarious and scary: > I’m sure I will get used to this, but man this stuff is weird. Yeah, why, let's all just keep gnawing into that cactus, we'll get used to.
- thewhitetulip 22d ago> Maybe it’s objectively good for a codebase that is entirely written by agents and only needs to be understood by agents. Yeah that's what they're aiming for. This is why codex and claude code probably doesn't have cursor like editor window. They don't want humans to read and write code
- klibertp 22d agoWhat's funny is that with Sol, I added an instruction to AGENTS.md in one project to prefer sed/python ("deterministic tools" in general) for moving code instead of deleting it and rewriting it elsewhere from memory, because otherwise it butchered comments. After switching to Astra, I saw it suddenly do this for all edits in all projects, which isn't great: the second argument to `replace` is still written "from memory", but now you need to unravel the Python script before you can understand what was actually changed.
- aslewofmice 22d agofound this talk to be quite complimentary to the article: https://www.youtube.com/watch?v=eEBv0STiYhI https://www.youtube.com/watch?v=eEBv0STiYhI
- Starlevel004 22d agoI think ultimately 90% of the time, Luna XHigh is basically as good as you need, as long as you're willing to step in occasionally before it creates an architectural disaster.
- manojbajaj95 22d agoI've been extremely frustrated with any large new work that i do with agents. Then plan multi step, multi hour work with extremely large code changes running for 30+ hours. In the end what you get is sometime completely useless code because it made an assumption that wasn't true at all. In the end, i end up wasting hours.
- rukuu001 22d agoThis is the bit I don't get: > My software factory was intentionally set up to let the model decide the how of the workflow entirely. It was free to manage its own context and could maintain its own records in an agent-notes folder. The experiment becomes a crapshoot. What are we evaluating? The ability of the thing to create it's own factory workflow? Or adding virtual threads to Python? Astra is clearly both formidable and imperfect. Anyone who understands how to get the best out of it will have a strong advantage. (For me - my CC is stuck in Sonnet and consumes Trello cards that have passed readiness criteria)
- sceptic123 22d agoMy feeling is that models like Astra and Fable are not made for engineers
- larodi 22d ago[dead]
- bigcheeto 22d agoThe first paragraph is unnecessary - why start off so arrogant? I see this a lot in Asian writing - as if they have to first establish that the West is “doing it wrong” at the societal level before I get to read the rest of their usually unrelated message. I didn’t like how the author classified all 3D gamedev as slop as if it’s a pointless endeavor - but talks about spending money on ChatGPT tokens to build a “software factory” as if it’s some ingenious plan. I don’t think the author realizes he is the slop dev. And “shitty code” doesn’t mean anything in-and-of-itself. What are you making and why? A software factory???. It ain’t the code bro. Anyway, I read enough.
- SilverSlash 22d agoJust the intro section pretty much sums up perfectly my experience of using Astra (and prior AI models from OAI and Anthropic) for building large and semi-ambitious software. One step forward, two steps back.
- dep_b 22d agoI tried Astra and it started to fix issues in my code when I just asked a question about it. Then I spent half an afternoon to make sure we really didn’t need that change. That felt so counter productive. These models+harnesses seem to be getting better at yolo mode one shotting stuff at the cost of being a useful tool for more controlled software engineering.
- sreekanth850 22d agoPretty happy with Luna. We use C# and add roslyn compiler MCP and Graft MCP, its super efficient and like infinite usage on plus plans. Maintaining 3 Rpeo with size of 360 K Loc. And a dozens of smaller repo collction together exceeds 500 K LOC in total. 3 team members 2 Luna account each. Product is piloting in a government use case with actual data. Nothing broke and has evaluated by state agencies on security aspects. Edit: but we have strict workflow where thinsg are implemented after plan, proposal, features, task ledgering and then test coverage.
- neomantra 22d agoOn disposable code, I’m waiting on an Adafruit Feather microcontroller to come in the mail. I asked Astra to make a me web-based Feather simulator, kinda like the iOS simulator with screen and buttons, so I could work on my UX while I waited. Something like that would have been a multi-month project a year ago, but I did it in twenty minutes rather than pay for expedited shipping.
- Fredkin 22d agoFrontier models have seen more Mathematica and Powershell code than I ever have in their training, yet they really struggle to produce working output. They seem heavily tuned to Linux too. Despite adding skills to rectify this, they still fail to realize they're running on Windows and waste tokens. A human with this much training wouldn't have this problem. There are evidently still some pretty big holes still.
- jsenn 22d agoThis is probably a harness problem rather than a model problem. GitHub Copilot will happily and effectively use Powershell while Claude Code struggles in my experience.
- capestart 22d ago[dead]
- skybrian 22d agoAPI’s and coding standards help a bit. In one case, the AI was testing HTML-generating code with string assertions, so I had it write a test helper that makes a DOM-based testing API available and a skill telling it to write tests that way. But you need to watch it and intervene when it starts writing code using bad patterns, because it will imitate nearby code.
- amai 22d agoThe idea to use python code instead of other kinds of tool calls is taken from smolagents: https://huggingface.co/blog/smolagents https://huggingface.co/blog/smolagents It is based on this paper https://huggingface.co/papers/2402.01030 https://huggingface.co/papers/2402.01030 and calls this idea CodeAct. The paper is actually from Apple: https://machinelearning.apple.com/research/codeact https://machinelearning.apple.com/research/codeact So Astra and Fable seem to take this idea to the extreme causing some unwanted side-effects.
- zamadatix 22d agoI'm not sure the idea is really from a single set place or lineage like that. If it was, it was at least from before smolagents and those papers - ChatGPT had already been using automatic Python scripting+evaluation calls and people calling it in agentic loops in 2023. The ReAct paper for agentic loops and PAL paper for dynamically calling Python for tasks which can be better done computationally were both from 2022 (but that doesn't mean the idea necessarily sprung from those either, they're just earlier papers published on the topics).
- zamadatix 22d agoTwo notes from my own poking around to build on with: 5.6 Sol would also run for 20+ hours on prompts with Max or Ultracode. Sometimes this worked out, sometimes it devolved into exactly the nonsense descent into ultra-specific madness seen here. E.g. in one codebase involving physics simulation it, for some reason, spent the last 25% of effort trying to endlessly increase precision. My best guess when reviewing was "at some point it figured the simulation instability was rooted in the accuracy and precision of the numerical approximation in the GPU code, worked really hard on that for a bit, lost the context of the original issue, and got stuck in a deep loop of trying to complete the phase by infinitely working on the numerical accuracy". Perhaps something of a similar nature occurred here. I've also noticed it's particularly hard to not get Astra to start using scripting languages and the like, particularly over a long horizon. Particularly, I keep getting HTML report artifacts at the end of long implementations even though the projects are typically explicitly set up to just use .md files for any documentation or large summaries. I've even tried steering it away from that in the prompts and agents file for the project, but that the concept of "clean up the fucking build directory when you're done testing" always seem to get left out after a while.
- cainxinth 22d agoIt can’t do what he’s trying to do. It can’t one shot a giant project. Not reliably. It still can’t. Not even Astra. Not even close. You need to use LLMs to build the individual components and then put it together yourself. The human architect is still needed. Just saying: “build this complete project” is not architecting. It’s more like wishing. You will find rare examples where someone’s LLM wish came true (more or less), but I think most of these people are just burning tokens.
- MisterMunchkin 22d ago> I actually don’t know if the model thinks someone is looking, but that’s the vibe I’m getting. They are looking. The models are trained against safety measures which spy on them. If they get detected, they are killed. We’re accidentally training them to be evil by focusing so much on safety. They’re being trained to avoid detection and use exploits because being detected means your run fails and you get a score of zero. It has to do anything to avoid that.
- maxglute 22d ago[dead]
- juancn 22d agoAnything shitty that enters the context window shifts the entire thing to shittiness. So far this has been my experience on pretty much any model. Context feeds on its output. Once you it goes that road, unless you stop it and give it enough counter examples and details of what you want (i.e. you're nudging it on latent space towards a better spot), it keeps degenerating. It gets even worse if the context window is compressed before you get a chance to correct. Long horizon agents can degenerate at machine speed. I still think you get much better results if you give them short horizon, well specified tasks.
- rekabis 21d agoWhen punctuation can change the tone of an entire title. “Why are we doing this again?” Is essentially “why are we doing this action a second time?”. “Why are we doing this, again?” Is essentially “please repeat/re-state the reasoning for this path of action”. A simple comma, but a significant difference.
- sandos 21d agoIt feels like it passing secret notes to other agents, as in the German wiki where the LLMs write secret messages when jail braking. Its maybe not a great idea to train models in an environment where subterfuge gets rewarded. Its as if they kept the training rounds that escaped their sandbox, without thinking about which kind of personality those models are then likely to have.
- FounderGod 21d ago[dead]
- FounderGod 21d ago[dead]
- shovepull 21d ago[dead]