9 ms·
XBOW, an autonomous penetration tester, has reached the top spot on HackerOne
- deleted 1y ago[deleted]
- ikmckenz 1y agoRelated: https://arstechnica.com/gadgets/2025/05/open-source-project-curl-is-sick-of-users-submitting-ai-slop-vulnerabilities/ https://arstechnica.com/gadgets/2025/05/open-source-project-...
- moyix 1y agoThe main difference is that all of the vulnerabilities reported here are real, many quite critical (XXE, RCE, SQLi, etc.). To be fair there were definitely a lot of XSS, but the main reason for that is that it's a really common vulnerability.
- ikmckenz 1y agoAll of them are real? You have a 100% rate of reports closed as valid?
- andrewstuart 1y agoAll the fun vanishes.
- tptacek 1y agoGood. It was in the way.
- kiitos 1y agoIn the way of what?
- tptacek 1y agoGetting more bugs fixed.
- kiitos 1y ago> Getting more bugs fixed. OK.. but "getting more bugs fixed" isn't any kind of objective success metric for, well, anything, right? It's fine if you want to use it as a KPI for your specific thing! But it's not like it's some global KPI for everyone?
- tptacek 1y agoIt's very specifically the objective of a security bug bounty.
- nottorp 1y ago[flagged]
- ryandrake 1y agoReceiving hundreds of AI generated bug reports would be so demoralizing and probably turn me off from maintaining an open source project forever. I think developers are going to eventually need tools to filter out slop. If you didn’t take the time to write it, why should I take the time to read it?
- triknomeister 1y agoEventually projects who can afford the smugness are going to charge people to be able to talk to open source developers.
- tough 1y agoisnt that called enterprise support / consulting
- deleted 1y ago[deleted]
- triknomeister 1y agoThis is without the enterprise.
- tough 1y agogotchu, maybe i could see github donations enabling issue creation or wahtever in the future idk but foss is foss, i guess source available doesnt mean we have to read your messages see sqlite (wont even take PR's lol)
- jgalt212 1y agoOne would think if AI can generate the slop it could also triage the slop.
- err4nt 1y agoHow does it know the difference?
- jekwoooooe 1y agoThey should ban this or else they will get swallowed up and companies will stop working with them. The last thing I want is a bunch of llm slop sent to me faster than a human would
- fredfish 1y agoAs long as they maintain a history per account and discourage gaming with new accounts, I don't see why anyone would want slop that performed lower just because the slop was manual. (I just had someone tell me that they wished the nonsensical bounty submissions they triaged were at least being fixed up with gpt3.)
- danmcs 1y agoHackerOne was already useless years before LLMs. Vulnerability scanning was already automated. When we put our product on there, roughly 2019, the enterprising hackers ran their scanners, submitted everything they found as the highest possible severity to attempt to maximize their payout, and moved on. We wasted time triaging all the stuff they submitted that was nonsense, got nothing valuable out of the engagement, and dropped HackerOne at the end of the contract. You'd be much better off contracting a competent engineering security firm to inspect your codebase and infrastructure.
- tptacek 1y agoMoreover, I don't think XBOW is likely generating the kind of slop beg bounty people generate. There's some serious work behind this.
- tecleandor 1y agoStill they're sending hundreds of reports that are being refused because they are not following the rules of the bounties. So they better work on that.
- tptacek 1y agoIf you thought human bounty program participants were generally following the rules, or that programs weren't swamped with slop already... at least these are actually pre-triaged vetted findings.
- mkagenius 1y ago> XBOW submitted nearly 1,060 vulnerabilities. Yikes, explains why my manually submitted single vulnerability is taking weeks to triage.
- tptacek 1y agoThe XBOW people are not randos.
- lcnPylGDnU4H9OF 1y agoThat's not their point, I think. They're just saying that those nearly 1060 vulnerabilities are being processed so theirs is being ignored (hence "triage").
- tptacek 1y agoIf that's all they're saying then there isn't much to do with the sentiment; if you're legit-finding #1061 after legit-findings #1-#1060, that's just life in the NFL. I took instead the meaning that the findings ahead of them were less than legit.
- croes 1y agoWhether it is legit-finding is precisely what needs to be checked, but you’re at spot 1061. >130 resolved >303 were classified as Triaged >33 reports marked as new >125 remain pending >208 were marked as duplicates >209 as informative >36 not applicable 20% bind a lot of resources if you have a high input on submissions and the numbers will rise
- tptacek 1y agoI think some context I probably don't share with the rest of this thread is that the average quality of a Hacker One submission is incredibly low. Like however bad you think the median bounty submission is, it's worse; think "people threatening to take you to court for not paying them for their report that they can 'XSS' you with the Chrome developer console".
- tecleandor 1y agoFirst: > To bridge that gap, we started dogfooding XBOW in public and private bug bounty programs hosted on HackerOne. We treated it like any external researcher would: no shortcuts, no internal knowledge—just XBOW, running on its own. Is it dogfooding if you're not doing it to yourself? I'd considerit dogfooding only if they were flooding themselves in AI generated bug reports, not to other people. They're not the ones reviewing them. Also, honest question: what does "best" means here? The one that has sent the most reports?
- jamessinghal 1y agoTheir success rates on HackerOne seem widely varying. 22/24 (Valid / Closed) for Walt Disney 3/43 (Valid / Closed) for AT&T
- thaumasiotes 1y ago> Their success rate on HackerOne seems widely varying. Some of that is likely down to company policies; Snapchat's policy, for example, is that nothing is ever marked invalid.
- jamessinghal 1y agoYes, I'm sure anyone with more HackerOne experience can give specifics on the companies' policies. For now, those are the most objective measures of quality we have on the reports.
- moyix 1y agoThis is discussed in the post – many came down to individual programs' policies e.g. not accepting the vulnerability if it was in a 3rd party product they used (but still hosted by them), duplicates (another researcher reported the same vuln at the same time; not really any way to avoid this), or not accepting some classes of vuln like cache poisoning.
- pclmulqdq 1y ago
- bgwalter 1y ago"XBOW is an enterprise solution. If your company would like a demo, email us at info@xbow.com." Like any "AI" article, this is an ad. If you are willing to tolerate a high false positive rate, you can as well use Rational Purify or various analyzers.
- moyix 1y agoYou should come to my upcoming BlackHat talk on how we did this while avoiding false positives :D https://www.blackhat.com/us-25/briefings/schedule/#ai-agents-for-offsec-with-zero-false-positives-46559 https://www.blackhat.com/us-25/briefings/schedule/#ai-agents...
- tptacek 1y agoYou should publish the paper quietly here (I'm a Black Hat reviewer, FWIW) so people can see where you're coming from. I know you've been on HN for awhile, and that you're doing interesting stuff; HN just has a really intense immune system against vendor-y stuff.
- moyix 1y agoYeah, it's been very strange being on the other side of that after 10 years in academia! But it's totally reasonable for people to be skeptical when there's a bunch of money sloshing around. I'll see if I can get time to do a paper to accompany the BH talk. And hopefully the agent traces of individual vulns will also help.
- tptacek 1y agoJ'accuse! You were required to do a paper for BH anyways! :)
- moyix 1y agoWait a sec, I thought they were optional? > White Paper/Slide Deck/Supporting Materials (optional) > • If you have a completed white paper or draft, slide deck, or other supporting materials, you can optionally provide a link for review by the board. > • Please note: Submission must be self-contained for evaluation, supporting materials are optional. > • PDF or online viewable links are preferred, where no authentication/log-in is required. (From the link on the BHUSA CFP page, which confusingly goes to the BH Asia doc: https://i.blackhat.com/Asia-25/BlackHat-Asia-2025-CFP-Preparation.pdf https://i.blackhat.com/Asia-25/BlackHat-Asia-2025-CFP-Prepar... )
- mellosouls 1y agoHave XBow provided a link to this claim, I could only find: https://hackerone.com/xbow?type=user https://hackerone.com/xbow?type=user Which shows a different picture. This may not invalidate their claim (best US), but a screenshot can be a bit cherry-picked.
- zndr 1y agoIf you scroll down on [the leaderboard](https://hackerone.com/leaderboard?year=2025&quarter=2&owasp=a1&country=US https://hackerone.com/leaderboard?year=2025&quarter=2&owasp=...) page to Country and select United States, xbow is currently on top
- mellosouls 1y agoAh thanks, I think it would be useful for them to perhaps add it as a footnote or something.
- chc4 1y agoI'm generally pretty bearish on AI security research, and think most people don't know anything about what they're talking about, but XBOW is frankly one of the few legitimately interesting and competent companies in the space, and their writeups and reports have good and well thought out results. Congrats!
- wslh 1y agoI'm looking forward to the LLM's ELI5 explanation. If I understand correctly, XBOW is genuinely moving the needle and pushing the state of the art. Another great reading is [1](2024). [1] "LLM and Bug Finding: Insights from a $2M Winning Team in the White House's AIxCC": https://news.ycombinator.com/item?id=41269791 https://news.ycombinator.com/item?id=41269791
- hinterlands 1y agoXbow has really smart people working on it, so they're well-aware of the usual 30-second critiques that come up in this thread. For example, they take specific steps to eliminate false positives. The #1 spot in the ranking is both more of a deal and less of a deal than it might appear. It's less of a deal in that HackerOne is an economic numbers game. There are countless programs you can sign up for, with varied difficulty levels and payouts. Most of them pay not a whole lot and don't attract top talent in the industry. Instead, they offer supplemental income to infosec-minded school-age kids in the developing world. So I wouldn't read this as "Xbow is the best bug hunter in the US". That's a bit of a marketing gimmick. But this is also not a particularly meaningful objective. The problem is that there's a lot of low-hanging bugs that need squashing and it's hard to allocate sufficient resources to that. Top infosec talent doesn't want to do it (and there's not enough of it). Consulting companies can do it, but they inevitably end up stretching themselves too thin, so the coverage ends up being hit-and-miss. There's a huge market for tools that can find easy bugs cheaply and without too many false positives. I personally don't doubt that LLMs and related techniques are well-tailored for this task, completely independent of whether they can outperform leading experts. But there are skeptics, so I think this is an important real-world result.
- absurdo 1y ago> so they're well-aware of the usual 30-second critiques that come up in this thread. Succinct description of HN. It’s a damn shame.
- normie3000 1y ago> Top infosec talent doesn't want to do it (and there's not enough of it). What is the top talent spending its time on?
- hinterlands 1y agoVulnerability researchers? For public projects, there's a strong preference for prestige stuff: ecosystem-wide vulnerabilities, new attack techniques, attacking cool new tech (e.g., self-driving cars). To pay bills: often working for tier A tech companies on intellectually-stimulating projects, such as novel mitigations, proprietary automation, etc. Or doing lucrative consulting / freelance work. Generally not triaging Nessus results 9-to-5.
- martinald 1y agoThis does not surprise me. In a couple of 'legacy' open source projects I found DoS attacks within 10 minutes, with a working PoC. It crashed the server entirely. I suspect with more prompting it could have found RCE but it was an idle shower thought to try. While niche and not widely used; there are at least thousands of publicly available servers for each of these projects. I genuinely think this is one of the biggest near term issues with AI. Even if we get great AI "defence" tooling, there are just so many servers and (IoT or otherwise) devices out there, most of which is not trivial to patch. While a few niche services getting pwned isn't probably a big deal, a million niche services all getting pwned in quick succession is likely to cause huge disruption. There is so much code out there that hasn't been remotely security checked. Maybe the end solution is some sort of LLM based "WAF" that inspects all traffic that ISPs deploy.
- ActorNightly 1y agoLegacy code (especially C++) does suffer from a lot of bugs, however modern code is generally much better. There is also a BIG hurdle between crashing something (which generally will be detected), versus RCE which requires a lot more work.
- Sytten 1y agoSince I am the cofounder of a mostly manual based testing in that space we do follow the new AI hackbots closely. There is a lot of money being raised (Horizon3 at 100M, Xbow at 87M, Mindfort will probably soon raise). The future is definitely a combination of human and bots like anything else, it won't replace the humans just like coding bots won't replace devs. In fact this will allow humans to focus ob the fun/creative hacking instead of the basic/boring tests. What I am worried about is on the triage/reproduction side, right now it is still mostly manual and it is a hard problem to automate.
- jp0001 1y agoI want to know how much they made in bounties versus how much they spent on compute. The thing about bug bounties, the only way to win is to not play the game.
- eddd-ddde 1y agoI think you have to factor in the training value. Each submission is a new data point that will only improve the overall performance.
- skeptrune 1y agoI'm confused on whether or not this actually outperformed humans. The more interesting statistic would be how much money it made versus the average hacker one top ranked contributor.
- lallysingh 1y agoI think a better one would be the average compute cost for xbow per exploit found, if you're interested in the shift in security economics this represents.
- moktonar 1y agoWhile impressive, a lot of manual human work was involved both to filter the input and the output, this is not a “fully” automated workflow, sorry. But, yeah, kudos to them.
- vmayoral 1y agoIt’s humans who: - Design the system and prompts - Build and integrate the attack tools - Guide the decision logic and analysis This isn’t just semantics — overstating AI capabilities can confuse the public and mislead buyers, especially in high-stakes security contexts. I say this as someone actively working in this space. I participated in the development of PentestGPT, which helped kickstart this wave of research and investment, and more recently, I’ve been working on Cybersecurity AI (CAI) — the leading open-source project for building autonomous agents for security: - CAI GitHub: https://github.com/aliasrobotics/cai https://github.com/aliasrobotics/cai - Tech report: https://arxiv.org/pdf/2504.06017 https://arxiv.org/pdf/2504.06017 I’m all for pushing boundaries, but let’s keep the messaging grounded in reality. The future of AI in security is exciting — and we’re just getting started.
- vasco 1y ago> It's humans Who would it be, gremlins? Those humans weren't at the top of the leaderboard before they had the AI, so clearly it helps.
- vmayoral 1y agoActually, those humans (XBOW's) were already top rankers. Just look it up. What's being critized here is the hype, which can be misleading and confusing. On this topic, wrote a small essay: “Cybersecurity AI: The Dangerous Gap Between Automation and Autonomy,” to sort fact from fiction -> https://shorturl.at/1ytz7 https://shorturl.at/1ytz7
- TZubiri 1y agoThis would be impressive even as a human assisted project. But there's a claim that it is unsupervised, which I doubt. See how these two claims contradict each other. >"XBOW is a fully autonomous AI-driven penetration tester. It requires no human input, " >"To ensure accuracy, we developed the concept of validators, automated peer reviewers that confirm each vulnerability XBOW uncovers. Sometimes this process leverages a large language model; in other cases, we build custom programmatic checks." I mean, I doubt you deploy this thing collecting thousands of dollars in bounties and you sit there twiddling your thumbs. Whatever work you put into the AI, whether fine tuned or generic and reusable, counts as supervised, and that's ok. Take the win, don't try to sell the automated dream to get investors or whatever, don't get caught up in fraud. As I understand it, when you discover a type of vulnerabilities, it's very common to automate the detection and find other clients with such vulnerability, these are usually short lived and the well dries up fast, you need to constantly stay on top of the latest trends. I just don't buy that if you leave this thing unattended for even 3 months it would keep finding gold, that's a property of the engineers that is not scaleable (and that's ok).
- deleted 1y ago[deleted]
- spacecadet 1y agoIve also invested some time in this space over the last several years. A group Im in took the approach of custom building agents for each CTF we approached. Our best so far was an agent participating in an AI CTF against current top injection/jailbreak and leakage defense techniques, the agent autonomously completed 22 of the 40 challenges and at one point held 8th place out of 380 teams. It eventually plateaued and slipped to 12th by the end. The tooling and models are maturing quickly and there is definitely some value in autonomous security agents, both offensive and defensive- but also still requires alot of work, knowledge(my group is all ML people), skill, planning- if you want to approach anything more than bug bashing. This recent paper from Dreadnode discusses a benchmark for this sort of challenge: https://arxiv.org/abs/2506.14682 https://arxiv.org/abs/2506.14682
- imglorp 1y agoAnd so we've arrived at William Gibson's black ice and ice-breaker (Russian military) systems. https://en.wikipedia.org/wiki/Burning_Chrome https://en.wikipedia.org/wiki/Burning_Chrome
- keisborg 1y ago«XBOW submitted nearly 1,060 vulnerabilities. All findings were fully automated, though our security team reviewed them pre-submission to comply with HackerOne’s policy on automated tools» That seems a bit unethical. I’ve thought companies specifically deny usage of automated tools. A bit too late ey…?
- 8200_unit 1y agoThey acknowledge that in the article and all submissions are human reviewed before they are submitted.
- keisborg 1y agoThe policies states it’s not allowed to use automated tools, not to submit report using automated tools alone. Human review does not really change that.
- slt2021 1y agoif a human reviewer can repro the bug, there is no difference between automated or human found bug. bug works and is repro - as a software owner, do you care if human or ai found it?
- keisborg 1y agoI cannot answer for all the program owners, but I imagine that there are other concerns than reproducibility
- billy99k 1y agoWhile I think this is great progress, It will be a hard sell for many business owners. I've been in this space (infosec) for awhile now and customers can be very apprehensive about non-AI/humans looking for vulnerabilities on their networks/systems.