6 ms·
You can get a taste of this today yourself with Codex Security. I turned it on just as an experiment and in less than a week it has now become essential to all
by mdeeks 4mo ago
You can get a taste of this today yourself with Codex Security. I turned it on just as an experiment and in less than a week it has now become essential to all of us. I was shocked how accurate it is, how many security issues it found in existing code, how it continually finds them as we commit, and how NO ONE is immune from making these mistakes.
I'd say it is about 90% accurate for us. Often even the "Low" findings lead us to dig and realize it is actually exploitable. Everyone makes these mistakes, from the most junior to the most senior. They are just a class of bugs after all.
I expect tools like this to be a regular part of the development lifecycle from here on. We code with AI, we review with AI, we search for vulns with AI. Even if it isn't perfect, it is easily worth the cost IMHO. Highly recommend you get something enabled for your own repos ASAP
- 0xAstro 4mo agoI would recommend you to try out the setup with gpt-5.5-cyber as the orchestrator and deepseek-v4-flash or some other fast cheap model as its workers. Getting pretty good results using this setup.
- Version467 4mo agoI’ve had the same experience. The ui is a little unclear about this, because it says you have 5 scans, but 1 scan is just the continuous monitoring of the default branch of a repo. The high impact findings have almost all been bang on for me. I was especially surprised by the high-quality documentation it produces as well as how narrow the proposed fixes are. I’m used to codex producing quite a but more code than it needs to, but the security model proposed fixes that are frequently <10 loc, targeting exactly the correct place. It’s really quite good. I’m assuming it’ll be pretty expensive once out of beta, but as a business I’d be jumping on this.
- winstonwinston 4mo ago> I expect tools like this to be a regular part of the development lifecycle from here on. We code with AI, we review with AI, we search for vulns with AI. Even if it isn't perfect, it is easily worth the cost IMHO. So, how is that supposed to work? Claude Code generates security bugs, then Claude Security finds them, then Claude Code generate fix, spend tokens, profit?
- jimmy2times 4mo agoThe AIs have already figured out how to succeed in a software job: 1. Ship bugs 2. Fix them 3. You're the hero!
- genghisjahn 4mo agoI thought we were all doing that already?
- flir 4mo agoJesus, dude. There are managers reading this.
- genghisjahn 4mo ago>_<
- OtomotO 4mo agoTake them out of the loop. Unless they are not human.
- josh-sematic 4mo agoI’m guessing you wouldn’t really rather be managed by a bot..,
- OtomotO 4mo agoI am not managed by anyone:p
- joquarky 4mo agoThey became obsolete when they stopped clearing obstacles, stopped masking politics, and started acting as a proxy for JIRA. But it's just they way it has always been done, so they get paid to meddle.
- rmast 4mo agoI help maintain a project that is used as a dependency by a lot of security tools to handle PE files. It’s disappointing that Anthropic and OpenAI never responded to the applications to their respective programs for open source maintainers. From my perspective it seems like their offers are primarily for the shiny well-known projects, rather than ones that get only a few million monthly installs but aren’t able to get thousands of stars due to being “hidden” as a dependency of popular tool.
- mnahkies 4mo agoOne issue I've seen with LLM's is adding superfluous code in the name of "safety" and confidently generating a bunch of stuff that was useful in years gone by, but now handled correctly by the standard lib. I'm of the opinion that less is more when it comes to code, and find the trend this is introducing quite frustrating. How do you avoid this pitfall?
- appplication 4mo agoGosh this couldn’t be more true, which IMO is the real reason LLM workflows are not strictly faster if you care about quality. Otherwise you end up with a codebase where only 60% of it is necessary. Standard testing patterns also tend not to be great at catching this particular flavor of LLM-ism.
- tomjakubowski 4mo agoI wonder this too. I prompted Opus 4.7 to generate some Python threading code for me. The code to run the sub-thread looked like this: def run(): with contextlib.suppress(SystemExit): do_thread_thing() threading.Thread(target=run, daemon=True).start() Suppressing SystemExit was surprising, and made me curious. I followed up and asked the model: what's the purpose of that? The model's response: "Honestly? Cargo-culting on my part. You should remove it."
- cassianoleal 4mo agoI had some shell scripts littered with `|| true`, which was obviously obscuring real errors everywhere. When I challenged the model, it gave me the same "cargo-culting" answer.
- bewuethr 4mo agoThe `|| true` is often done because people use `errexit` as part of "Bash strict mode"[1], which comes with so many caveats[2] that I usually avoid it. Claude, however, loves it. [1]: http://redsymbol.net/articles/unofficial-bash-strict-mode/ http://redsymbol.net/articles/unofficial-bash-strict-mode/ [2]: https://mywiki.wooledge.org/BashPitfalls#set_-euo_pipefail https://mywiki.wooledge.org/BashPitfalls#set_-euo_pipefail
- hollowturtle 4mo ago> I was shocked how accurate it is, how many security issues it found in existing code, how it continually finds them as we commit, and how NO ONE is immune from making these mistakes. Dude is flexing that he's pushing unsecure code every day, that's a skill!
- Smaug123 4mo agoBy the way, you might be interested in looking up “blameless post-mortems” and indeed the field of incident response more generally. Modern incident response practice is to treat failures of an individual to do something as problems with the system they were operating in, because humans aren’t designed to be consistent or perfect and therefore shouldn’t be pretended or assumed to be.
- lateral_cloud 4mo agoDid you need to do anything special to get access to Codex Security?
- ofjcihen 4mo agoNot sure what the threshold is but I sent them all of my bug bounty profiles and papers I’ve authored. I don’t think you need all of that though. I know a whole mess of people that have gotten it for much less. Should just give it a try.
- gofreddygo 4mo agoThis got me thinking, so what happens in two years? every tom, dick and harry who can type english has the tools to attack any software that isn't patched. tools that were accessible to specialized groups, now made available to anybody with a grudge and a few dollars for tokens. and what does anthropic and openai do? They form an inner ring to make the latest models available first to Enterprises. Enterprises will cough up the prices that anthropic and openai set, they have no choice here. e Eventually everybody pays. This does not sound good
- mrtesthah 4mo agoI would say that if this sounds untenable to you, then you may want to consider that the way we architect software has itself been untenable for a while. What Mythos can accomplish today in public, an APT unit can already accomplish in secret.
- conradkay 4mo agoYou'll have access to the same models as your hypothetical attackers, and a big advantage if only you have access to the source code
- mdeeks 4mo agoTwo years? That exists right now. You only have to point Codex Security at an open source repo. There are a lot of tools and companies that are spinning up today that do autonomous pentesting. I'm not even sure a specialized model is needed here. It probably just needs the right harness around existing ones. I expect the next two years to be absolutely brutal for hacks. Attackers have supercharged tools in their hands right now. Defenders are only getting started and will have to plow through a massive backlog of newly uncovered vulns. The major short term downside is that open source or personal projects won't be able to afford things like Codex Security.
- nullbio 4mo ago> The major short term downside is that open source or personal projects won't be able to afford things like Codex Security. Realistically, all open-source projects should be forced to have automated scans of this nature before their releases can be shipped. This is something the package managers and github need to figure out. It'd stop the supply chain attacks too.
- kortilla 4mo agoIt seems to me like either your architecture is fucked up or you’re using the wrong language/tooling for the type of software you are making if you’re introducing security vulnerabilities that frequently.
- fragmede 4mo ago"get a taste of this". The real thing is, GPT-5.5 is better than Opus 4.7, so if Anthropic doesn't release Mythos soon, other people are going to notice and switch off Claude.
- alexwwang 4mo agohttps://blog.chuanxilu.net/en/posts/2026/05/dual-pass-review-recall-precision-tradeoff/ https://blog.chuanxilu.net/en/posts/2026/05/dual-pass-review... This is what I did. Using a loop skill to dig problems and bugs in each step on development from design to coding to make sure the output software works properly and on purpose.
- perlgeek 4mo agoWhat kind of application are you developing?
- deleted 4mo ago[deleted]