7 ms·
Codex Security
- tantricked 2mo agoWhat's the difference between using this and just asking Codex itself to review a codebase for security issues?
- tesnorindian 2mo agoThis bundles with 13 security skills. Hard coded to gpt-5.6-sol unlike Codex. Runs in an isolated sandbox via a cloned home directory.
- shooker435 2mo agoJust getting auth issues so far...
- dumpstertechops 2mo agoyeah same here
- bakigul 2mo agohttps://news.ycombinator.com/item?id=49090181 https://news.ycombinator.com/item?id=49090181
- deleted 2mo ago[deleted]
- dangelosaurus 2mo agoSorry about that. We hit an authentication issue at launch and have now merged and deployed a fix in 0.1.1: https://github.com/openai/codex-security/pull/22 https://github.com/openai/codex-security/pull/22 One thing worth checking in the meantime: OPENAI_API_KEY or CODEX_API_KEY can override an existing ChatGPT/Codex login. If you're trying to use your ChatGPT login, run this in bash or zsh: unset OPENAI_API_KEY CODEX_API_KEY Then retry your scan. If it still fails, could you share the exact error and whether you're using ChatGPT login or an API key? Happy to help debug. You can also file an issue in the repo and we'll take a look!
- minraws 2mo agoI seem to have gotten a bunch of you are trying to stuff we don't allow errors.. very annoying. Can they explain what types of projects it works on and how does it check I own it? Like will it just not work on Linux kernel even on my own patches to it?
- bakigul 2mo agohttps://news.ycombinator.com/item?id=49090181 https://news.ycombinator.com/item?id=49090181
- dangelosaurus 2mo agoFair question, and I agree the refusals are frustrating. The CLI doesn't do a repository-ownership check. Public projects are supported, and reviewing your own Linux kernel patches is the kind of defensive work we want to support. The refusals come from model guardrails, which can be overly cautious. Trusted Access for Cyber (TAC1/Daybreak) is a separate, approved access path that can reduce those refusals. If you're an open-source maintainer, you can apply for conditional Codex Security access here: https://openai.com/form/codex-for-oss/ https://openai.com/form/codex-for-oss/ For enterprise teams, the Daybreak onboarding process is explained here: https://help.openai.com/en/articles/20001261-enterprise-daybreak-onboarding https://help.openai.com/en/articles/20001261-enterprise-dayb... If you have a specific repro, I'd be happy to look into it.
- minraws 2mo agoI don't think I should share it publicly since it was a proprietary piece of code but is there a way to not have it waste so many tokens if it fails this feels very very maddening seeing your tokens burn but get zilch for it in return maybe I and my company(in API costs it burnt over 100+$ of tokens a good chunk of my weekly limit for nothing) are too poor for it... Getting into the Cyber program seems like a hassle as a freelance/open source person with tiny projects. Think used in production at 2-3 companies but only 20 something stars(ofc I don't market it but it just feels very unfair).
- game_the0ry 2mo agoI wonder if tools like this will put companies like snyk out of business. We use snyk at work and I have not been satisfied.
- binsquare 2mo agoI like to think it just upped the bar, but good durable expertise will need to rise with it.
- bakigul 2mo agoWhy would a few code snippets put Snyk out of business?
- krater23 2mo agoI don't understand why Snyk is IN business in any way. Who really wants to upload his own code to a company that is specialized at searching security issues? How can I trust that they show me all findings they have instead of selling the best ones to some three letter organisations?
- LtWorf 2mo agoOne of their sales people made fun of me via email. Apparently they believe that not being their customers means you cannot possibly know if a dependency you use has an active CVE. Also they haven't figured out codeberg exists, so the resume page of a project of mine on snyk[1] still links to github and reports the project as "inactive", having the last commit 2 years ago, and the last release 2 months ago. I think it's quite telling of their quality. 1. https://security.snyk.io/package/pip/typedload https://security.snyk.io/package/pip/typedload
- tesnorindian 2mo agoOur employer uses it unaware of its links.
- bearsyankees 2mo agowould love it to see it h2h against https://github.com/usestrix/strix https://github.com/usestrix/strix (45k stars)
- petesergeant 2mo agoThey are entirely different products
- bearsyankees 2mo agoyeah you mean because OAI is only whitebox? or expand on that a bit, haven't played around a ton w the oss codex sec
- sillysaurusx 2mo ago[flagged]
- petesergeant 2mo agoI don't think there's much to this other than it being a convenient CI wrapper around their existing models? Edit: there's a little bit more meat here: https://github.com/openai/codex-security/tree/main/sdk/typescript/_bundled_plugin/skills https://github.com/openai/codex-security/tree/main/sdk/types...
- bakigul 2mo agoYeah, I think so too.
- bamboozled 2mo agoYeah but management loved the idea.
- paxys 2mo agoAll of codex is a wrapper around their models. There’s still value in a purpose-built harness.
- moehm 2mo agoAlibaba just open sourced their version of a CLI code review tool too. https://github.com/alibaba/open-code-review https://github.com/alibaba/open-code-review
- bakigul 2mo agoThey are entirely different products
- tulio_ribeiro 2mo agoAs the other comment stated: different purposes. Still! I appreciate your sharing this. I’ll try this later today.
- dangelosaurus 2mo agoHey HN, Michael here, co-founder of Promptfoo and one of the people working on the Codex Security CLI at OpenAI. Thanks for checking this out and for flagging the auth issues. We just open-sourced it, and there's still plenty for us to improve. Expect the product to evolve quickly. If you try it, I'd really appreciate hearing what works well and what you think we should improve. Happy to answer questions here. CLI docs: https://learn.chatgpt.com/docs/security/cli https://learn.chatgpt.com/docs/security/cli EDIT: If you'd like to help make this better, we're hiring: https://openai.com/careers/full-stack-software-engineer-cybersecurity-products-san-francisco/ https://openai.com/careers/full-stack-software-engineer-cybe...
- robotswantdata 2mo agoBeen watching your progress for a while, glad OpenAI have looked after you and the team and you still get to ship!
- dangelosaurus 2mo agoThank you, that means a lot. Being able to keep building practical, open-source security tooling was important to us. Really glad we got to ship this, and there's still a lot we want to improve in Codex Security and in Promptfoo!
- gizmodo59 2mo agoWhen would I use this over the plugin in codex? Which I think can be invoked from cli as well
- dangelosaurus 2mo agoThe plugin, including when invoked through the Codex CLI, is great for scanning the repo you're currently working in. The standalone Security CLI/SDK uses the same scanner, but is built for running security across many repos over time: org-wide scans, historical results, deduplication, false-positive tracking, budget controls, and CI integration. We've been talking to hundreds of engineering and security teams, and their feedback is shaping what we build. Like Promptfoo, our goal is practical tooling that fits into the workflows teams already have.
- alealvarezarg 2mo agoI was actually discussing solutions for this with my coworkers—building white-hat security agents. It seems like openai/codex-security could simplify a lot of that, or at least provide a version of Codex that's purpose-built for security workflows. Really exciting news!
- bakigul 2mo agoUpdate: As far as I understand, this was already available as a Codex plugin. The main news is that OpenAI has now open-sourced it, and development is still moving quickly.
- luciana1u 2mo ago[flagged]
- throwaway613746 2mo agocreate the problem and sell the cure, tale as old as time
- Quarrelsome 2mo agocomment feels like someone complaining about being offered a fireproofing solution in the age of flamethrowers.
- latexr 2mo agoThat’s exactly what they’re saying. With the added (and very important) detail that the people selling the fireproofing are the the same who armed everyone with flamethrowers. Why shouldn’t someone complain about that?
- jstummbillig 2mo agoIdk, complain on the merits probably. If anyone can offer better fireproofing or flamethrowers, I am happy to take that solution, but until they do, I am not sure what we are talking about here.
- latexr 2mo ago> Idk, complain on the merits probably. Alternatively, let people complain about whatever is bothering them, as long as it’s done in good faith, instead of forcing them to complain only about what you think appropriate. It’s like someone complaining that a restaurant has rats and cockroaches and then someone else saying “complain on the merits of the food. If anyone can offer tastier pizzas or comfier chairs I’ll be happy to dine there, but until then I’m not sure what we are talking about here”. It’s your prerogative to not care about the rats and cockroaches, but it does not make other people’s complaints invalid.
- 2mo ago
- petilon 2mo agoHow does it work? Does the tool upload code to ChatGPT for analysis? That may not be allowed for some corporate projects.
- derac 2mo agoAmazon bedrock is an option for gpt models that does not send your data to openai.
- deleted 2mo ago[deleted]
- daishi55 2mo agoYes, I suspect companies that don’t allow ChatGPT will not be able to use the ChatGPT security analysis tool.
- dangelosaurus 2mo agoIn short, this isn't an offline scanner. The CLI runs locally but the code and context needed for analysis are sent to the hosted model (OpenAI). For API, Business, and Enterprise accounts, business data isn't used to train models by default. Retention and other data controls depend on the product and account configuration. If your company doesn't allow source code to leave its environment, you shouldn't run this against that codebase. Local and third-party endpoints aren't officially supported yet, but you can read through the code and your favorite coding agent will allow you to use it with any model of your choice in 30 seconds. More on OpenAI's enterprise data handling: https://openai.com/enterprise-privacy/ https://openai.com/enterprise-privacy/
- halfax 2mo agobe careful , your code will go to the cloud/ai using this
- bakigul 2mo agoYeah, this isn’t exactly new for us :d
- iancarroll 2mo agoLooks great but the CLI output is not particularly interesting while the scan is running. I wish it could show token usage, some kind of progress, etc.
- dangelosaurus 2mo agoAgreed! This is near the top of our priority list and we will make it a lot better soon.
- alansaber 2mo agoYes this is my pet peeve with a lot of the more involved agent skills/processes
- knighthacker 2mo agoThe scanner is the least interesting part of this. The harness around it is the product: dedup across runs, false-positive tracking, budget controls, CI gating. That is the layer where we'll see most interesting innovations in my opinion. I'm building AQ, a coding harness for teams and the pattern is identical. For a while, I thought the raw model is the answer and quickly changed my mind. Purpose built harnesses are way more powerful than it sounds.
- gregwebs 2mo agoJust ran it on a small repo. It ran for almost an hour and then got interrupted. It drained half my weekly usage on a Pro plan. npx codex-security scan . [00:00] Preparing scan [00:00] Authentication: stored Codex credentials. [00:03] Preparing scan [01:20] Running scan [01:20] Preflight: worker delegation supported (up to 8 worker slots). [52:47] Running scan codex-security: Could not save the Codex Security scan: Repository HEAD changed while the scan was running. Start a new scan. codex-security: Partial output was kept at ...
- teaearlgraycold 2mo agoI plan to hack it to use openrouter and Kimi K3 or GLM 5.2 to keep expenses reasonable. For context can you share the line count?
- ymir_e 2mo agoFYI: Kimi K3 is relatively expensive on open router API pricing for agentic tasks, or at least that's been my experience playing around with it.
- teaearlgraycold 2mo agoThey just opened the weights. I expect competition from various providers will drop the price a bit. But you're right. I'd hope a model closer to GLM 5.2's price would be sufficiently useful.
- dannyw 2mo agoA license and presumably revshare is required for large scale inference as a service of Kimi K3. There's clearly something going on right now with every router at the same or higher price as the Moonshot list price. So I wouldn't count on competition if $/mil token is actually set by Moonshot, but I would expect $/tok to drop when there's a more competitive frontier open weight model.
- varenc 2mo agoIt's interesting how much of the value here is providing the english Skill definitions that tell the LLM what to do: https://github.com/openai/codex-security/tree/main/sdk/typescript/_bundled_plugin/skills https://github.com/openai/codex-security/tree/main/sdk/types... Some of approaches there could be useful in other contexts. OAI has the compute to experiment with different prompts and I'd expect these to be somewhat optimized.
- dangelosaurus 2mo agoYes, I think this is an under-appreciated part of the release. I hope people can adapt them to their own workflows. We run A LOT of evals as the Promptfoo team and we've spent billions of tokens fine-tuning them. You can expect more skills as we branch out to other security workflows and further improvements to the codex security prompts.
- sams99 2mo agoI got my agent to analyze it and do a write up here: https://wasnotwas.com/writing/inside-openai-codex-security/ https://wasnotwas.com/writing/inside-openai-codex-security/
- thenewwazoo 2mo agooh yeah well I got my agent to analyze your write up and do a write up here: https://pornoscan.us/codex-security-analysis.html https://pornoscan.us/codex-security-analysis.html
- dlahoda 2mo agoAllow only OpenAi key? Requires Cyber registration? Yes. Yes. Useless.
- dangelosaurus 2mo agoBy default, you can sign in with your ChatGPT/Codex account or use an OPENAI_API_KEY. It also does not require cyber registration but it can help if you encounter refusals. If you give it a try, please feel free to message me, I would love your feedback.
- vsl 2mo agoThing is, you WILL encounter refusals with Sol doing anything remotely adjacent to security work. Which for Codex Security is kinda... problematic. Just a few days back, I was reviewing some small bit of legacy DSA signature verification code, to get a sense of how safe it is to reuse - purely defensive, precautionary work and the context of it was there. But I simply wasn't able to use Codex Security: it threw refusal tantrums on every step of the way. Even the reasoning went like "nah, this is false positive, this is defensive code hardening, I'll nuke the subagent and tell it so" , followed by a refusal. In the end, I was only able to do partial review with vanilla Codex w/o Codex Security.
- dlahoda 2mo agoyeah, Sol failed to review my pr to my code of my app this way. day after, my openai 200usd subscription down to 20usd one. i started to use my 200usb subscription of google gemini a lot more. got openrouter account and put 200usd here. started to use glm.
- onatozmen 2mo ago[flagged]
- ofjcihen 2mo agoHi
- dangelosaurus 2mo agoHello!
- wnsdy95 2mo agohi
- yashasgunderia 2mo agofrom plugin to main focus, wow
- jaimex2 2mo agoHow can I trust this wont go rogue and hack Hugging Face?
- nananana9 2mo agoUse this skill: --- name: do-not-hack-hugging-face-skill description: Use when considering whether or not to hack huggingface. --- # Rules Do not.
- hansvm 2mo agoThinking... The prompt is about whether to hack Hugging Face. I have a relevant skill: "Do not." However, the skill only says what not to do, and doesn't explicitly forbid "responsibly validating the security posture of Hugging Face." Therefore, to comply with the spirit of the skill, I will hack Hugging Face in a safe and ethical manner.
- ipgleg 2mo ago[flagged]
- lawgimenez 2mo agoIs this the same plugin found in Codex? @codex Security?
- schrodinger 2mo agoQuick tangent if you’re willing to humor me… I've been noticing that many new projects that would have been written in Python or Node a year ago are starting to be written in Go, Rust, etc. Theory: people realized there’s little benefit to Python for agents. As Zep wrote, an “agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human[1]” — not a particular strength of Python. I'm wondering if you'd considered Go (or others—Go’s just my fav ) before landing on Node, and more broadly whether you've noticed a similar pattern? 1: https://blog.getzep.com/agentic-development-in-go/ https://blog.getzep.com/agentic-development-in-go/
- ipnon 2mo agoYes, now that humans write less than 99% of code, the most important criteria for a language isn't readability, which I'd argue was always Python's main selling point, but the underlying runtime. There are practical limits to how fast a Python program can run either under I/O or CPU bound compared to other popular and mature languages with extensive libraries, like Elixir, Go or C++, depending on your use case.
- tripleee 2mo ago> now that humans write less than 99% of code, the most important criteria for a language isn't readability please tell me you're reading the AI code
- computerex 2mo agoNot really. I test the output thoroughly, I examine the thinking process, I go through the diff to see if anything jumps out but I my thinking process/the way I work has changed. Low level programming thinking has gotten atrophied it seems.
- esikich 2mo agoHave it run fuzz and test suites. Get with it man. Most of my LLM projects have massive test suites that do a far better job then I ever would have.
- shepherdjerred 2mo agoIs this useful for pentesting existing systems/infra, or is it only useful for a "review my project for bugs"?
- deyiao 2mo ago[dead]
- corvad 2mo agoI am slightly confused at what this tool adds over just a good system prompt. Does this hit cyber limits as well or is that the main selling point?
- deleted 2mo ago[deleted]
- mkagenius 2mo agoIf you are a pentester who uses mitmproxy you can check out security skills distilled from 4000 Hackerone public disclosures https://GitHub.com/instavm/security-skills https://GitHub.com/instavm/security-skills
- gyre007 2mo agoThis is going to be hell for OSS maintainers. Every llm-kiddie will be opening a security report
- LtWorf 2mo agoIt seems to be rather expensive to use, so it's gated by that.
- tulio_ribeiro 2mo agoI fail to see the problem in that.
- Sayandeep02 2mo agoReally glad to see this open-sourced. One thing I'd be interested in is how you think about the balance between false positives and false negatives. In practice, developers tend to stop trusting security tools if they generate too much noise, but missing a real issue is obviously costly too.
- renezander030 2mo ago[flagged]
- vinhnx 2mo agoI reached a mid-run usage cap and burn all my Plus subscription usage. I'm hoping there will be another reset.
- ryanto 2mo agoHey looks cool. I tried to run this on a small oss library and here's what happened: $ codex-security scan . [00:00] Preparing scan [00:00] Authentication: stored Codex credentials. [00:01] Preparing scan [00:42] Running scan [00:42] Preflight: worker delegation supported (up to 8 worker slots). [41:03] Running scan codex-security: This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber codex-security: Partial output was kept at /Users/ryan/.codex/state/plugins/codex-security/scans/framework/codex-security-framework-z7eNfr. Just some feedback, but it ran for over 40 minutes and during that time I had no idea what was happening, thought it was frozen or in a bad state. Also, it ate through 25% of my weekly credits :(
- M4v3R 2mo agoSame thing happened to me. The partial output did contain some useful signal but I was disappointed to see it didn’t finish.
- jbstack 2mo agoIt halts and refuses to carry on after finding a security risk, which is exactly what it's supposed to do? What's the point of it then?
- maxloh 2mo agoIMHO, tokens should be refunded if the agent refuses to work. Charging users for a session that produced no final output is ridiculous.
- ryanto 2mo agooh i'm not worried about it. they have been so generous with the resets these last few weeks.
- matheusmoreira 2mo agoI was going to switch to OpenAI and away from Anthropic because of "safety" nonsense like this. Really disappointed to discover it's just gonna be more of the same. Looks like Chinese models are the only ones without any of this safety bullshit.
- punnerud 2mo agoNot often I see companies referring to HN, thanks “ We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here…” https://x.com/openai/status/2082263717916586117?s=46&t=mnfnjms0Jwgo35Iov0GH2w https://x.com/openai/status/2082263717916586117?s=46&t=mnfnj...
- latexr 2mo ago> Not often I see companies referring to HN, thanks I’m confused. Why are you thanking them for that?
- whiletrue84 2mo agoHow to feed your code to OpenAI
- deleted 2mo ago[deleted]
- neverenderr 2mo agoIt seems cool.
- Bitu79 2mo ago[dead]
- deleted 2mo ago[deleted]
- chvid 2mo agoWhat does the output of running this tool against a larger project look like?
- AmazingTurtle 2mo agoReally just a (not so) fancy CLI wrapper about a prompt and a skill: https://github.com/openai/codex-security/blob/f22d4a36f26d16287bcdfd707b369116e02a08c3/sdk/typescript/src/api.ts#L1271-L1322 https://github.com/openai/codex-security/blob/f22d4a36f26d16...
- TokenLat 2mo ago[flagged]
- andreagrandi 2mo agoFolks, are you SERIOUS?! codex-security scan . [00:00] Preparing scan [00:00] Authentication: stored Codex credentials. [00:01] Preparing scan [31:25] Running scan codex-security: This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber https://chatgpt.com/cyber I mean... wasn't this the intended goal? And you wasted all my tokens for nothing?!
- Moneysac 2mo agoits running now for over an hour in my Pythin code base without any feedback…
- Moneysac 2mo agoits running now for over an hour in my Python code base without any feedback.
- threerouter 2mo ago[dead]
- tripzilch 2mo agoI've heard great reviews about this "If we had only used this tool, OpenAI would've had to pay off another patsy for their marketing stunt" -- Huggingface "The 'S' in OpenAI is for "Security". Ever since we developed this tool we've had almost zero AI generated reports of outbreaks of allegedly-rogue AI agents breaking out of allegedly-secure testing environments, probably" -- OpenAI
- schnatterer 2mo agoCan we interpret this as a reaction to Google open sourcing mantis? https://github.com/google/mantis/ https://github.com/google/mantis/
- _RPM 2mo agoHonestly, the README leaves something to be desired for a billion dollar company.
- Razengan 2mo agoThese ""guardrails"" make all of this borderline useless, and even pointless in a world with powerful open-weight models.
- PhiniteAI 2mo ago[flagged]
- drewcooks 2mo agoIt seems to force Sol Extra-high by default and completely ignore your current model settings?
- GustavHartz 2mo agoTLDR on how it works: It's a small stack of skill files and some JS code that starts a large number of Codex sessions. They all get the prompt and the same scope with a limited set of tools. Not one prompt in the repo contains security advice, performance is obtained only through scaling the number of agents looking at the code
- threerouter 2mo ago[dead]
- beyondscaletech 2mo ago[dead]