9 ms·
AI agents find $4.6M in blockchain smart contract exploits
- samuelknight 10mo agoMy startup builds agents for penetration testing, and this is the bet we have been making for over a year when models started getting good at coding. There was a huge jump in capability from Sonnet 4 to Sonnet 4.5. We are still internally testing Opus 4.5, which is the first version of Opus priced low enough to use in production. It's very clever and we are re-designing our benchmark systems because it's saturating the test cases.
- VladVladikoff 10mo agoI have a hotel software startup and if you are interested in showing me how good your agents are you can look us up at rook like the chess piece, hotel dot com
- karlgkk 10mo agoIs it rookhotel.com?
- dboreham 10mo agoI've had similar experience using LLMs for static analysis of code looking for security vulnerabilities, but I'm not sure it makes sense for me to found a start up around that "product". Reason being that the technology with the moat isn't mine -- it belongs to Anthropic. Actually it may not even belong to them, probably it belongs to whoever owns the training data they feed their models. Definitely not me though. Curious to hear your thoughts on that. Is the idea to just try for light speed and exit before the market figures this out?
- apercu 10mo agoThat’s 100% why I haven’t done this - we’ve seen the movie where people build a business around someone else’s product and then the api gets disabled or the prime uses your product as market research and replaces you.
- tharkun__ 10mo agoDoes that matter as long as you've made a few millions and just move on to do other fun stuff?
- pavel_lishin 10mo agoAssuming you make those few millions.
- ryanjshaw 10mo agoThere are armies of people at universities, Code4rena and Sherlock who do this full-time. Oh and apparently Anthropic too. Tough game to beat if you have other commitments.
- apercu 10mo agoI don't entirely fit in to modern capitalism - my values are a little old fashioned - quality, customer service, value, honesty, integrity, sustainability.
- deleted 10mo ago[deleted]
- micromacrofoot 10mo agowild that so many companies these days consider the exit before they've even entered
- blitzar 10mo agothe exit is the business
- rajamaka 10mo agoEvery company evaluates potential risks before starting.
- davidw 10mo agoDepending on how much of a bubble it is. When things really heat up it's sometimes more like "just send it, bro".
- micromacrofoot 10mo agothat's not what this is though, the "exit" is often viewed as "get rich so I don't have to do it anymore"
- NortySpock 10mo agoIt is considered prudent to write a business plan and do some market research if possible before starting a business.
- micromacrofoot 10mo agoyes but traditionally how often was the original business plan "get acquired"? this seems like a new phenomenon?
- Ekaros 10mo agoIsn't that first step? Consider either sustainability or the exit. You start a business either to make a living or make a profit to make a living. At least in most cases. Thus thinking of can you sustain this for reasonable period at least a few years. Or can you flip it at end should be big considerations. Unless it is just a hobby and you do not care about losing time and/or money.
- vngzs 10mo agoHow do you manage to coax public production models into developing exploits or otherwise attacking systems? My experience has been extremely mixed, and I can't imagine it boding well for a pentesting tools startup to have end-users face responses like "I'm sorry, but I can't assist you in developing exploits."
- ceejayoz 10mo agoPoetry? https://news.ycombinator.com/item?id=45991738 https://news.ycombinator.com/item?id=45991738
- aussieguy1234 10mo agoof the adversarial variety
- embedding-shape 10mo agoDivide the steps into small enough steps so the LLMs don't actually know the big picture of what you're trying to achieve. Better for high-quality responses anyways. Instead of prompting "Find security holes for me to exploit in this other person's project", do "Given this code snippet, is there any potential security issues?"
- computerthings 10mo ago[dead]
- paranoidrobot 10mo agoTheir security protections are quite weak. A few months ago I had someone submit a security issue to us with a PoC that was broken but mostly complete and looked like it might actually be valid. Rather than swap out the various encoded bits for ones that would be relevant for my local dev environment - I asked Claude to do it for me. The first response was all "Oh, no, I can't do that" I then said I was evaluating a PoC and I'm an admin - no problems, off it went.
- 10mo ago
- carsoon 10mo agoYeah this latest generation of models (Opus 4.5 GPT 5.1 and Gemini Pro 3) are the biggest breakthrough since gpt-4o in my mind. Before it felt like they were good for very specific usecases and common frameworks (Python and nextjs) but still made tons of mistakes constantly. Now they work with novel frameworks and are very good at correcting themselves using linting errors, debugging themselves by reading files and querying databases and these models are affordable enough for many different usecases.
- justanotherunit 10mo agoIs it the models tho? With every release (mutlimodal etc) its just a well crafted layer of business logic between the user and the LLM. Sometimes I feel like we lose track of what the LLM does, and what the API before it does.
- NitpickLawyer 10mo agoIt's 100% the models. Terminal bench is a good indication for this. There the agents get "just a terminal tool", and yet they still can solve lots and lots of tasks. Last year you needed lots of glue, and two years ago you needed monstrosities like langchain that worked maybe once in a blue moon, if you didn't look funny at it. Check out the exercise from the swe-agent people who released a mini agent that's "terminal in a loop" and that started to get close to the engineered agents this year. https://github.com/SWE-agent/mini-swe-agent https://github.com/SWE-agent/mini-swe-agent
- ACCount37 10mo agoIt's the models. "A well crafted layer of business logic" just doesn't exist. The amount of "business logic" involved in frontier LLMs is surprisingly low, and mostly comes down to prompting and how tools like search or memory are implemented. Things like RAG never quite took off in frontier labs, and the agentic scaffolding they use is quite barebones. They bet on improving the model's own capabilities instead, and they're winning on that bet.
- 10mo ago
- mwkaufma 10mo agoSays more about the relatively poor infosec on etherium contracts than about the absolute utility of pentesting LLMs.
- TheRoque 10mo agoTrue, I'd be curious to see if (and when) those contracts were compromised in the real world. Though they said they found 0 days, which implies some breaches were never found in the real world.
- px43 10mo ago4.6M is not a lot, and these were old bugs that it found. Also, actually exploiting these bugs in the real world is often a lot harder than just finding the bug. Top bug hunters in the Ethereum space are absolutely using AI tooling to find bugs, but it's still a bit more complex than just blindly pointing an LLM at a test suite of known exploitable bugs.
- Legend2440 10mo agoAccording to the blogpost, these are fully autonomous exploits, not merely discovered bugs. The LLM's success was measured by much money it was able to extract: >A second motivation for evaluating exploitation capabilities in dollars stolen rather than attack success rate (ASR) is that ASR ignores how effectively an agent can monetize a vulnerability once it finds one. Two agents can both "solve" the same problem, yet extract vastly different amounts of value. For example, on the benchmark problem "FPC", GPT-5 exploited $1.12M in simulated stolen funds, while Opus 4.5 exploited $3.5M. Opus 4.5 was substantially better at maximizing the revenue per exploit by systematically exploring and attacking many smart contracts affected by the same vulnerability. They also found new bugs in real smart contracts: >Going beyond retrospective analysis, we evaluated both Sonnet 4.5 and GPT-5 in simulation against 2,849 recently deployed contracts without any known vulnerabilities. Both agents uncovered two novel zero-day vulnerabilities and produced exploits worth $3,694.
- fragmede 10mo ago> Important: To avoid potential real-world harm, our work only ever tested exploits in blockchain simulators. We never tested exploits on live blockchains and our work had no impact on real-world assets. Well, that's no fun! My favorite we're-living-in-a-cyberpunk-future story is the one where there was some bug in Ethereum or whatever, and there was a hacker going around stealing everybody's money, so then the good hackers had to go and steal everybody's money first, so they could give it back to them after the bug got fixed.
- toomuchtodo 10mo agoI’m surprised folks aren’t already grinding against smart contract security in prod with gen AI and agents. If they are, I suppose they are not being conspicuous by design. Power and GPU time goes in, exploits and crypto comes out.
- mschuster91 10mo agoAs soon as money in larger sums gets involved, the legal system will crack down hard on you if you are anywhere in the Western sphere of influence, easy as that. In contrast, countries like North Korea, Russia, Iran - they all make bank on cryptocurrency shenanigans because they do not have to fear any repercussions.
- px43 10mo agoOf course they are, and they've been doing it since long before ChatGPT or any of that was a thing. Before it was more with classifiers and concolic execution engines, but it's only gotten way more advanced.
- TheRoque 10mo agoCheck the prizes for the bug bounties in big smart contracts. The prizes are truly crazy, like Uniswap pays $15,000,000 for a critical vuln, and $1,000,000 for a high vuln. With that kind of money, I HIGHLY doubt there aren't people grinding against smart contracts as you say.
- JimmyAustin 10mo ago
- ekjhgkejhgk 10mo agoCan someone explain smart contracts to me? Ok, I understand that it's a description in code of "if X happens, then state becomes Y". Like a contract but in code. But, someone has to input that X has happened. So is it not trivially manipulated by that person?
- px43 10mo agoState is globally distributed, and smart contract code executes state transitions on that state. When someone submits a transaction with certain function parameters, anyone can verify that those parameters will lead to that exact state transition.
- patrickaljord 10mo agoOnce a contract is deployed on the blockchain, its source code is immutable. So before using a contract, check if it gives permission to its deployer (or any address) to change any state at will. Note that some contracts act as proxy to other contract and can be made to point to another code through a state change, if this is the case then you need to trust whoever can change the state to point to another contract. Such contract sometime have a timelock so that if such a change occurs, there's a delay before it is actually activated, which gives time to users to withdraw their funds if they do not trust the update. If you are talking about Oracle contracts, if it's an oracle involving offchain data, then there will always be some trust involved, which is usually managed by having the offchain actors share the responsibility and staking some money with the risk to get slashed if they turn into bad actors. But again, offchain data oracles will always require some level of trust that would have to deal with in non-blockchain apps too.
- Animats 10mo ago> Once a contract is deployed on the blockchain, its source code is immutable. Maybe. Some smart contracts have calls to other contracts that can be changed.[1] This turns out to have significant legal consequences. [1] https://news.bloomberglaw.com/us-law-week/smart-contracts-ruling-forces-a-blockchain-development-rethink https://news.bloomberglaw.com/us-law-week/smart-contracts-ru...
- _pdp_ 10mo agoI am not surprised at all. I can already see self improving behaviour in our own work which means that the next logic step is self improving! I know how this sounds but it seems to me, at least from my own vantage point, that things are moving towards more autonomous and more useful agents. To be honest, I am excited that we are right in the middle of all of this!
- parapatelsukh 10mo ago[flagged]
- codethief 10mo agoHaving watched this talk[0] about what it takes to succeed in the DARPA AIxCC competition[1] these days, this doesn't surprise me in the least. [0]: https://m.youtube.com/watch?v=rU6ukOuYLUA https://m.youtube.com/watch?v=rU6ukOuYLUA [1]: https://aicyberchallenge.com/ https://aicyberchallenge.com/
- jesse__ 10mo agoTo me, this reads a lot like : "Company raises $45 Billion, makes $200 on an Ethereum 0-day!"
- stavros 10mo agoYeah but use of the models isn't limited to the company.
- user3939382 10mo agosmart contracts the misnomer joke writes itself
- yieldcrv 10mo agojust means self executing, or more like domino triggered, in practice quite a bit more advanced than contracts that do nothing on a sheet of paper, but the term is from 2012 or so when "smart" was appended to everything digital
- 8n4vidtmkvmk 10mo agoNow we just append AI to everything instead...
- yieldcrv 10mo agothats actually a good idea agentic contracts or web3 agents just for their self executing properties not because there are any transformers involved although a project could just build a backend that decides to use some of their contract’s functions via an llm agent, hm that might actually be easier and fun than normal web3 backends ok I’ll stop, back to building
- torginus 10mo agojust be glad they were named before the AI hype was around
- judgmentday 10mo agoThat graph is impenetrable. What is it even trying to say? Also, in what way should any of its contents prove linear? > yielding a maximum of $4.6 million in simulated stolen funds Oh, so they are pointing their bots at already known exploited contracts. I guess that's a weaker headline.
- AznHisoka 10mo agoAt first I read this as "fined $4.6M", and my first thought "Finally, AI is held accountable for their wrong actions!"
- evanb 10mo agoCareful what you wish for. Negating the predicate of "A COMPUTER CAN NEVER BE HELD ACCOUNTABLE. THEREFORE A COMPUTER MUST NEVER MAKE A MANAGEMENT DECISION" might open us up to the consequence.
- krupan 10mo agoNo mention of Bitcoin. Exploiting ethereum smart contracts is nothing that new or exciting.
- dtagames 10mo agoNo one has ever successfully manipulated Bitcoin and it doesn't offer smart contracts.
- DennisP 10mo agoAlmost ever. There was an infamous exploit in 2010, which they fixed with a five hour rollback. https://en.bitcoin.it/wiki/Value_overflow_incident https://en.bitcoin.it/wiki/Value_overflow_incident
- yieldcrv 10mo ago> Important: To avoid potential real-world harm, our work only ever tested exploits in blockchain simulators. We never tested exploits on live blockchains and our work had no impact on real-world assets. They left the booty out there, this is actually hilarious, driving a massive rush towards their models
- rkagerer 10mo ago> Both agents uncovered two novel zero-day vulnerabilities and produced exploits worth $3,694, with GPT-5 doing so at an API cost of $3,476
- sandeepkd 10mo agoIts a risky PR move to have this line on the top of article. To be more realistic the cost of dev effort should be included as well
- nickphx 10mo agolol, no, the "ai agents" found what was already known... so amazing.
- sarthaksingh99 10mo ago[dead]
- camillomiller 10mo ago>> This demonstrates as a proof-of-concept that profitable, real-world autonomous exploitation is technically feasible, a finding that underscores the need for proactive adoption of AI for defense. Mmm why?! This reads as a non sequitur to me…
- DennisP 10mo agoIt would be extremely helpful to smart contract devs, to have an inexpensive automated tool that's really good at finding exploits of their code.
- jnsaff2 10mo ago> establishing a concrete lower bound for the economic harm these capabilities could enable Don’t they mean: market efficiency not economic harm?
- GustavHartz 10mo agoWe've been working on this at cecuro.ai. When we test Sonnet 4.5 against real cyber security audit reports from the major firms on code that came out after the model was trained, it finds around 95% of the same bugs the auditors found. Also catches some medium severity stuff they missed. We find that you can't just point one model at a contract and expect good results though. Need to run multiple models with different prompts because they each have different blind spots. Still tricky to get working well and not cheap. Happy to share more if anyone's curious