5 ms·
IMHO, Codex with Astra/Sol and Claude with Fable/Opus are all any professional programmer should be using in Sep 2026, if they can afford it. These models are
by jacobgold 17d ago
IMHO, Codex with Astra/Sol and Claude with Fable/Opus are all any professional programmer should be using in Sep 2026, if they can afford it.
These models are still terrible compared to what we'd actually wish for, but they're the best available.
If you can get away with using the $200/mo subscriptions, it's really not even a money thing for most professionals.
Almost all of my work is now plan, generate, review, plan, generate, review, commit, push.
I'm using Claude or Codex (or both), and they're doing all of the testing "inline" rather than through a CI action, etc.
- cyanydeez 17d agoQwen3.8 is all you need.
- varispeed 17d agoAstra is quite crap (enters reasoning loops like Gemini used to and fails to actually work on a task - would say yes this needs fixing, so I say go ahead and then it will spend half an hour coming back with yes this needs fixing and not doing any fix) and Fable/Opus unusable in many instances (they struggle to generate coherent English let alone code). Out of these only Sol is quite useful - actually finishes a task, though you need to interrupt often as it likes to wander into its comfort zone.
- nullc 17d agoUse of closed models is unprofessional, and depending on your field negligent. The fact that it has been widely normalized does not make it less so. You're handing over your (presumably your customer/employers) data to an unaccountable third party which has demonstrated itself willing to commit criminal acts, and to take other people's data without permission. Your ability to continue to perform this work can be withdrawn at any time for any (or no) reason. You have little ability to validate that the work is being performed as expected and isn't being silently nerfed or outright subverted based on competitive considerations, bribes, overactive 'safety', or cost management. Outsourcing to a black box would be a reasonable expectation if you asked a non-professional to perform the work. A professional should be able to account for the tools they use.
- kestrel-robotic 17d ago[flagged]
- tomhow 16d agoPlease omit internet tropes on HN. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- flatline 17d agoI think that's a bit strong of an assertion. How many people use copilot daily under an enterprise agreement? I don't necessarily disagree in spirit, especially given the questionable data sanitization around the recent Navier-Stokes announcement, but most companies disclose huge amounts of data regularly to hopefully-trusted third parties. I think the internet -- and a good share of the world's commerce -- would grind to a halt if we suddenly stopped. Setting up, securing, and maintaining local models for even a small user base is non-trivial and there is way more demand than supply for that skillset right now.
- pbasista 17d ago> Your ability to continue to perform this work can be withdrawn at any time for any (or no) reason. Thanks to the fact that there are no widespread stories about this actually occurring in practice, at least not yet, people do not take it as a relevant risk at the moment. > You have little ability to validate that the work is being performed as expected and isn't being silently nerfed or outright subverted based on competitive considerations, bribes, overactive 'safety', or cost management. Yes, I agree that this is a real concern that many people might rightfully have. And I am unaware of any way to mitigate this concern while using black box AI models. Because the only thing that their creators can do is to tell their customers: "trust us". But there is no way to objectively verify whether they serve tainted AI model responses or not.
- rozenmd 17d agoAnyway here's how Cloudflare orchestrates AI reviews at scale: https://blog.cloudflare.com/ai-code-review/ https://blog.cloudflare.com/ai-code-review/
- deleted 17d ago[deleted]
- chrisguilbeau 17d agoWith everything changing it's great to hear others have the same workflow. I added a snapshot step so I'm doing plan, generate step 1, review, snapshot, generate step 2, review, snapshot... That way I have a chance to diff with the previous iteration and clean up comments, modify skills, etc. also if it bonks on a step I'm one snapshot away from trying again... Is there a place people share their workflows other than HN comments?
- jacobgold 17d agoI wish I knew, I haven't had any place to point people to. I'm going to start sharing this on YouTube since I already spend a few hours each week talking some friend or user through the latest best practices.
- extr 17d agoI just tell everyone to use Fable 5.1 for everything at this point. Astra is unfortunately a dud, I'm sure they will try to fix a bunch of it with GPT-6.1 but OAI has had this issue for awhile now where every other generation has some sort of strange tic, or reward hacking issue, or something. It's almost like they are balancing the RL on the tip of a needle. Opus 5 has issues too, comment-slop, claude-ish, etc. 5.1 on the other hand can seemingly do no wrong. Easy to work with, writes human-level code. Expensive, yes, but even at Low effort it's well worth it.
- triyambakam 17d agoWhat have you noticed about Astra? I haven't used Claude models lately so I can't compare but it seems fine compared to 5.6 Sol
- extr 17d ago- It doesn't write great code. - Occasionally has strange tics around asking for permission for obvious next-steps, implied actions, etc. - It's very expensive, both in terms of tokens and % usage on subscription plans. - Relatedly, effort level is unintuitive. Sometimes it seems like higher effort levels are actually cheaper due to not under-thinking and needing to correct work. But other times they are overkill and send the model into rabbitholes. That said, it's fantastic as a code-reviewer or "hunter seeker". It's better at finding bugs than Fable and "Get this well articulated task done single-mindedly" is an Astra-shaped task.
- chis 17d agoI gave up on fable 5 after I asked it to critique my PR and it spit out a page of complete nonsense technical jargon. Like, to the point that I had to review the feedback with other models and try to parse what it was saying and ultimately it wasn't even right. Compare to Astra and Sol where I can almost forget there's a model and just speak/read naturally. I think I should give 5.1 another chance but I am just so triggered by the way it talks after spending so long battling fable 5. Also I'm starting to wonder if the latest round of models have finally saturated for my personal coding needs. I mean obviously not for taste and judgement, but those barely seem to improve with model generations. For just spitting out a 1000-line feature I've vaguely scoped out, Astra feels basically as good as I need.
- baxtr 17d agoCould you elaborate on your exact setup? Where do you run these models?
- jacobgold 17d agoSince you asked, the answer is that I built and use an agent multiplexer called Clor https://clor.com https://clor.com I have a $200/mo Claude subscription and a $200/mo Codex subscription, and I'm signed in to both. The Docker containers keep each session isolated, so dev servers, browser testing, etc. can work without conflicts. It includes `/ask-claude` and `/ask-codex` skills that I use very frequently to have the Claude or Codex harness call out to the other one for advice on plans, bug repro, code review, etc. The agents run in total "yolo" mode, so there are no permission prompts to approve. The risk is mitigated by the Docker containers (which don't necessarily provide a security barrier but do limit accidents). I was doing this manually in Ghostty tabs for a long time, and it got painful, so I built a much more sophisticated version that I (and my friends/colleagues) could use.
- aschobel 17d agoI have a slightly jankier setup. Generally using Claude Code with Fable 5.1 (high) to plan and implement (Opus 5 (medium) as the implementer subagents), and using Codex with Astra high to review the plan and review the implementers' output. Using the OpenAI codoex plugin thingy: https://github.com/openai/codex-plugin-cc https://github.com/openai/codex-plugin-cc
- jacobgold 17d agoI originally had multiple skills for Claude and Codex but found that "ask" is a great single mechanism. "Ask claude about this" "Ask codex to implement this" "Ask claude to review this plan" etc
- aschobel 17d ago"Ask" is a super clever mechanism. I'll give that a shot. It's also more polite than "tell" or "yell at"!
- ghthor 17d agoI disagree; I just spent 15x dogfooding some Claude setup I rolled out to the org making changes that would have cost me less then a dollar had I used Luna and I would have got the same, if not better results; better because it would have been faster so I could have iterated more.
- epolanski 17d ago[flagged]
- jacobgold 17d agoLow IQ: "Just use Claude and Codex" Midwit: "No, you see, you need a deterministic 12-stage multi-agent orchestration framework with vector embedding semantic routing, and five open weight models with custom harnesses!" Genius: "Just use Claude and Codex"
- epolanski 16d agoNow you only need to figure which of the 3 you really are.
- rationalist 17d agoUnfortunately ChatGPT just stopped allowing people to upgrade to the $200/mo subscription. I started my first paid subscription ($100/mo) last week, and now I want to upgrade and I can't :-(
- oblio 17d agoHardware capacity limitations?
- rationalist 16d agoThey've been losing money on the $200/mo plan, so that's my guess.
- micw 17d agoAlso not available in their business offer. So I have now a $20 business plus a $200 "personal" plan in my account.
- fy20 17d agoFor side projects I pretty much exclusively use Luna xhigh. The $20/mo plan with the recent generous resets is more than enough for me. Sometimes I reach the 5hr limit, but haven't reached the weekly limit yet. The most recent project it finished was a SIP client for an ESP32 in-wall touch panel that I got from AliExpress for $50. It rings when someone is at my doorbell and let's me answer calls and see video. Yes an ESP32 can stream H.264 video :D My only complaint with Luna is it seems to give up when the work is half finished, and I often need to tell it to continue. But I feel this is mainly a harness problem. I just use it in ChatGPT/Codex as it gives me easy remote access to check in on what it's doing. (At my dayjob I usually spend $200+/day with Opus/Fable)
- nickthegreek 16d agowhy not Luna at max?