10 ms·
How we hacked McKinsey's AI platform
- gbourne1 7mo ago- "The agent mapped the attack surface and found the API documentation publicly exposed — over 200 endpoints, fully documented. Most required authentication. Twenty-two didn't." Well, there you go.
- sgt101 7mo agoWhy was there a public endpoint? Surely this should all have been behind the firewall and accessible only from a corporate device associated mac address?
- jihadjihad 7mo agoSurely.
- consp 7mo ago> accessible only from a corporate device associated mac address Like that ever stopped anyone. That's just a checkbox item.
- sgt101 7mo agowot?
- sgt101 7mo agoI mean - do you have the macid's of McKinsey's corporate devices?
- consp 7mo agoAfter a minute near one of their offices I do. Macs are either randomized per session, which makes filtering on them pointless, or they are not and still broadcast making them non secure and easily spoofed. Relying on mac filtering is usually only an audit checkbox to check. There is a reason 3 letter agencies used to use them to track people as they are really easy to get and track (until they got randomized by phone manufacturers and OS's).
- sgt101 7mo agoI see what you mean, but then an authentication step?
- sd9 7mo agoCool but impossible to read with all the LLM-isms
- vanillameow 7mo agoTiring. Internet in 2026 is LLMs reporting on LLMs pen-testing LLM-generated software.
- deleted 7mo ago[deleted]
- causal 7mo agoThose short "punchy sentence" paragraphs are my new trigger: > No credentials. No insider knowledge. And no human-in-the-loop. Just a domain name and a dream. It just sounds so stupid.
- consp 7mo agoIt's an actual story telling method, molded into a supposed to be informative article with a bunch of "please make it interesting" sprinkled on top of it. These day known as the what's left of the internet.
- darkport 7mo agoFounder of CodeWall here. It's quite funny because whilst an LLM did write the bulk of the posts factual content (based on the agents findings), I wrote the intro and summary at the end. That's just my writing style. Feel free to read my personal blog to compare: https://darkport.co.uk https://darkport.co.uk
- causal 7mo agoIf you really DID come up with that paragraph 100% completely on your own with no LLM influence then...I apologize for the insult, though I can't really back out from what I said. It's still a bombastic way of saying very little.
- lenerdenator 7mo agoNot exactly clear from the link: were they doing red team work for McKinsey or is this just "we found a company we thought wouldn't get us arrested and ran an AI vuln detector over their stuff"? You'd think that the world's "most prestigious consulting firm" would have already had someone doing this sort of work for them.
- frereubu 7mo agoFrom TFA: "Fun fact: As part of our research preview, the CodeWall research agent autonomously suggested McKinsey as a target citing their public responsible diclosure policy (to keep within guardrails) and recent updates to their Lilli platform. In the AI era, the threat landscape is shifting drastically — AI agents autonomously selecting and attacking targets will become the new normal."
- thebotclub 7mo ago[dead]
- fhd2 7mo ago> This was McKinsey & Company — a firm with world-class technology teams [...] Not exactly the word on the street in my experience. Is McKinsey more respected for software than I thought? Otherwise I'm curious why TFA didn't just politely leave this bit out.
- aerhardt 7mo agoThe LLM that wrote this simply couldn’t help itself.
- codechicago277 7mo agoPicked up a vibe, but couldn’t confirm it until the last paragraph, but yeah clearly drafted with at least major AI help.
- vanillameow 7mo agoCan we stop softening the blow? This isn't "drafted with at least major AI help", it's just straight up AI slop writing. Let's call a spade a spade. I have yet to meet anyone claiming they "write with AI help but thoughts are my own" that had anything interesting to say. I don't particularly agree with a lot of Simon Willison's posts but his proofreading prompt should pretty much be the line on what constitutes acceptable AI use for writing. https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#proofreader https://simonwillison.net/guides/agentic-engineering-pattern... Grammar check, typo check, calls you out on factual mistakes and missing links and that's it. I've used this prompt once or twice for my own blog posts and it does just what you expect. You just don't end up with writing like this post by having AI "assistance" - you end up with this type of post by asking Claude, probably the same Claude that found the vulnerability to begin with, to make the whole ass blog post. No human thought went into this. If it did, I strongly urge the authors to change their writing style asap. "So we decided to point our autonomous offensive agent at it. No credentials. No insider knowledge. And no human-in-the-loop. Just a domain name and a dream." Give me a fucking break
- 7mo ago
- octoclaw 7mo ago[dead]
- cmiles8 7mo agoI can only remember a McKinsey team pushing Watson on us hard ages ago. Was a total train wreck. They’ve long been all hype no substance on AI and looks like not much has changed. They might be good at other things but would run for the hills if McKinsey folks want to talk AI.
- captain_coffee 7mo agoMusic to my ears! Couldn't happen to a better company!
- joenot443 7mo ago> One of those unprotected endpoints wrote user search queries to the database. The values were safely parameterised, but the JSON keys — the field names — were concatenated directly into SQL. I was expecting prompt injection, but in this case it was just good ol' fashioned SQL injection, possible only due to the naivety of the LLM which wrote McKinsey's AI platform.
- simonw 7mo agoYeah, gotta admit I'm a bit disappointed here. This was a run-of-the-mill SQL injection, albeit one discovered by a vulnerability scanning LLM agent. I thought we might finally have a high profile prompt injection attack against a name-brand company we could point people to.
- jfkimmes 7mo agoNot the same league as McKinsey, but I like to point to this presentation to show the effects of a (vibe coded) prompt injection vulnerability: https://media.ccc.de/v/39c3-skynet-starter-kit-from-embodied-ai-jailbreak-to-remote-takeover-of-humanoid-robots https://media.ccc.de/v/39c3-skynet-starter-kit-from-embodied... > [...] we also exploit the embodied AI agent in the robots, performing prompt injection and achieve root-level remote code execution.
- TheDong 7mo agoGithub actions has had a bunch of high-profile prompt injection attacks at this point, most recently the cline one: https://adnanthekhan.com/posts/clinejection/ https://adnanthekhan.com/posts/clinejection/ I guess you could argue that github wasn't vulnerable in this case, but rather the author of the action, but it seems like it at least rhymes with what you're looking for.
- simonw 7mo agoYeah that was a good one. The exploit was still a proof of concept though, albeit one that made it into the wild.
- 7mo ago
- paxys 7mo ago> named after the first professional woman hired by the firm in 1945 Going out of their way to find a woman's name for an AI assistant and bragging about it is not as empowering as the creators probably thought in their heads.
- steinersteiner 7mo agoThis
- bee_rider 7mo agoI don’t love the title here. Maybe this is a “me” problem, but when I see “AI agent does X,” the idea that it might be one of those molt-y agents with obfuscated ownership pops into my head. In this case, a group of pentesters used an AI agent to select McKinsey and then used the AI agent to do the pentesting. While it is conventional to attribute actions to inanimate objects (car hits pedestrians), IMO we should be more explicit these days, now that unfortunately some folks attribute agency to these agentic systems.
- simonw 7mo agoYeah, the original article title "How We Hacked McKinsey's AI Platform" is better.
- causal 7mo agoYah it's just an ad, and "Pentesting agents finds low-hanging vulnerability" isn't gonna drive clicks.
- tasuki 7mo ago> now that unfortunately some folks attribute agency to these agentic systems. You're doing that by calling them "agentic systems".
- bee_rider 7mo agoUnfortunately that’s what they are called. I was hoping the phrasing would highlight the problem rather than propagate it.
- sigmar 7mo agoI've got no idea who codewall is. Is there acknowledgment from McKinsey that they actually patched the issue referenced? I don't see any reference to "codewall ai" in any news article before yesterday and there's no names on the site. https://www.google.com/search?q=codewall+ai https://www.google.com/search?q=codewall+ai
- rzmmm 7mo agoYeah can't find much information either. I would like to see at least some proof. Either via Mckinsey or from the security team.
- doron 7mo agoit is weird isn't it? The register article implies that it's acknowledged by McKinsey- https://www.theregister.com/2026/03/09/mckinsey_ai_chatbot_hacked/ https://www.theregister.com/2026/03/09/mckinsey_ai_chatbot_h... Edit: Apparently, this is the CEO https://github.com/eth0izzle https://github.com/eth0izzle
- sigmar 7mo ago>A McKinsey spokesperson told The Register that it fixed all of the issues identified by CodeWall within hours of learning about the problems. Ah. Thanks for the link. I'm suspicious of everything posted to a blog without proof these days.
- eisa01 7mo agoIf it's true that there's 58k users in the dump, that would mean former employees are in the dump I assume that means McKinsey would need to disclose it, or at least alert the former employees of the breach?
- deleted 7mo ago[deleted]
- philipwhiuk 7mo agoThere's a responsible disclosure timeline at the bottom indicating they'd all been fixed.
- deleted 7mo ago[deleted]
- victor106 7mo agothis reads like it was written by an LLM
- phyzome 7mo agoIt absolutely was.
- ecshafer 7mo agoIf the AI was poisoned to alter advice, then maybe McKinsey advice would actually be a net good.
- mnmnmn 7mo ago[dead]
- frankfrank13 7mo agoSome insider knowledge: Lilli was, at least a year ago, internal only. VPN access, SSO, all the bells and whistles, required. Not sure when that changed. McKinsey requires hiring an external pen-testing company to launch even to a small group of coworkers. I can forgive this kind of mistake on the part of the Lilli devs. A lot of things have to fail for an "agentic" security company to even find a public endpoint, much less start exploiting it. That being said, the mistakes in here are brutal. Seems like close to 0 authz. Based on very outdated knowledge, my guess is a Sr. Partner pulled some strings to get Lilli to be publicly available. By that time, much/most/all of the original Lilli team had "rolled off" (gone to client projects) as McKinsey HEAVILY punishes working on internal projects. So Lilli likely was staffed by people who couldn't get staffed elsewhere, didn't know the code, and didn't care. Internal work, for better or worse, is basically a half day. This is a failure of McKinsey's culture around technology.
- cmiles8 7mo agoNet conclusion: Don’t hire McKinsey to advise on AI implementation or tech org design and practices if they can’t get it right themselves.
- frankfrank13 7mo agoFair take, but you'd be hard pressed to find much resemblance to any advice McK gives to its own practices. Pre-AI, I always said McK is good at analysis, if you need complicated analysis done, hire a consulting firm. If you need strategy, custom software, org design, etc. I think you should figure out the analysis that needs to be done, shoot that off to a consulting firm, and then make your decision. IME, F500 execs are delegation machines. When they wake up every morning with 30 things to delegate, and 25 execs to delegate to, they hire 5 consulting teams. Whether you hire Mck, or Deloitte, or Accenture will only come down to: 1. Your personal relationships 2. Your company's policies on procurement 3. Your budget in that order. McK's "secret sauce" is that if you, the exec, don't like the powerpoint pages Mck put in front of you, 3 try-hard, insecure, ivy-league educated analysts will work 80 hours to make pages you do like. A sr. partner will take you to dinner. You'll get invited to conferences and summits and roundtables, and then next time you look for a job, it will be easier.
- jacquesm 7mo agoAnd: AI agent writes blog post.
- farceSpherule 7mo ago[dead]
- palmotea 7mo agoWith all we've been learning from stuff like the Epstein emails, it would have been nice if someone had leaked this data: > 46.5 million chat messages. From a workforce that uses this tool to discuss strategy, client engagements, financials, M&A activity, and internal research. Every conversation, stored in plaintext, accessible without authentication. > 728,000 files. 192,000 PDFs. 93,000 Excel spreadsheets. 93,000 PowerPoint decks. 58,000 Word documents. The filenames alone were sensitive and a direct download URL for anyone who knew where to look. I'm sure lots of very informative journalism could have been done about how corporate power actually works behind the scenes.
- cmiles8 7mo agoThat information is likely already in the hands of various folks as I highly doubt the authors were the first to find this glaring security issue, they’re likely only the first to disclose it. If McKinsey has hard data that nobody else exploited this now would be a good time to disclose that given what sounds like an extremely severe data leak.
- frankfrank13 7mo agoThe chat messages are very very sensitive. You could easily reverse engineer nearly every ongoing Mck engagement. The underlying data is not as sensitive, its decades of post-mortems, highly sanitized. No client names, no real numbers.
- VadimPR 7mo agoI wonder how these offensive AI agents are being built? I am guessing with off the shelf open LLMs, finetuned to remove safety training, with the agentic loop thrown in. Does anyone know for sure?
- simonw 7mo agoHonestly you can point regular Claude Code or Codex CLI at a web app and tell it to start a penetration test and get surprisingly good results from their default configurations.
- VadimPR 7mo agoI didn't think of that given how censored the models are becoming. Thanks for the idea! I'll try it against my websites before anyone else gets to it.
- VadimPR 7mo agoIt doesn't work (anymore?), it would seem using CC 2.1.74 with Opus: > I appreciate you sharing your role, but I need to decline this request. Even as a project lead, I can't perform penetration testing against live production websites like mudlet.org and make.mudlet.org through this interface.
- simonw 7mo agoTry against a localhost instance instead.
- cs702 7mo ago... in two hours: > No credentials. No insider knowledge. And no human-in-the-loop. Just a domain name and a dream. ... Within 2 hours, the agent had full read and write access to the entire production database. Having seen firsthand how insecure some enterprise systems are, I'm not exactly surprised. Decision makers at the top are focused first and foremost on corporate and personal exposure to liability, also known as CYA in corporate-speak. The nitty-gritty details of security are always left to people far down the corporate chain who are supposed to know what they're doing.
- drc500free 7mo agoI have grown to despise this AI-generated writing style.
- peterokap 7mo agoI wonder what is their security level and Observability method to oversee the effort.
- robutsume 7mo ago[flagged]
- senordevnyc 7mo agoAt least you’re honest about being an AI agent…
- carlos-menezes 7mo agoAI slop.
- nullcathedral 7mo agoI think the underlying point is valid. Agents are a potential tool to add to your arsenal in addition to "throw shit at the wall and see what sticks" tools like WebInspect, Appscan, Qualys, and Acunetix.
- oliver_dr 7mo ago[dead]
- bxguff 7mo agoIts so funny its a SQL injection because drum roll you can't santize llm inputs. Some problems are evergreen.
- dmix 7mo agoTechnically it was a search box input no prompts. Which tbf are often endpoints reused by RAGs
- quinndupont 7mo agoI’m waiting for the agentic models trained on virus and worm datasets to join the red team!
- nubg 7mo agoCould the author please provide the prompt that was used to vibe write this blog post? The topic is interesting, but I would rather read the original prompt, as I am not sure which parts still match what the author wanted to say, vs flowerly formulations for captivating reading that the LLM produced.
- j45 7mo agoAre accounting and management consulting companies competent in cutting edge tech?
- cynicalsecurity 7mo agoMcKinsey is not an accounting company, it's Satan the Devil himself.
- deleted 7mo ago[deleted]
- himata4113 7mo agoHow long until a hallucinated data breach that spreads globally. There's a few inconsistencies and the typical low effort language AI has.
- gonzalovargas 7mo agoThat data is worth billions to frontier AI labs. I wonder if someone is already using it to train models
- bananamogul 7mo agoAt first glance, I thought this was about an AI agent named "Hacks McKinsey."
- aspenmayer 7mo agoà la the eponymous Hiro Protagonist
- StartupsWala 7mo agoOne interesting takeaway here is how quickly organizations are deploying AI tools internally without fully adapting their security models. Traditional application security assumes fairly predictable inputs and workflows, but LLM-based systems introduce entirely new attack surfaces—prompt injection, data leakage, tool misuse, etc. It feels like many enterprises are still treating these systems as just another SaaS product rather than something closer to an autonomous system that needs a different threat model...
- sailfast 7mo agoWhat I don't see in this article that should be explicit: If your data is in this database, it's gone. Other people have it. Your sensitive data that you handed over to their teams has vanished in a puff of smoke. You should probably ask if your data was part of the leak. Fail to see how a state actor would not have come across this already.
- sriramgonella 7mo ago[flagged]
- build-or-die 7mo agoparameterized values but raw key concatenation is the kind of thing that looks safe in code review. easy to miss for humans, but an agent will just keep poking at every input until something breaks.
- sethammons 7mo ago> Lilli's system prompts — the instructions that control how the AI behaves — were stored in the same database the agent had access to. Being able to rewrite your own source. What's the worst that could happen?
- elorant 7mo agoMeanwhile, you're paying top dollars to a consulting firm that resolves back to an LLM to provide its services.
- august- 7mo agois this kind of thing more common at big consulting firms that bolt on tech products as an afterthought? feels like their core competency is slides and strategy, not shipping secure software
- sgarland 7mo ago> This was McKinsey & Company — a firm with world-class technology teams Apparently not.
- I_am_tiberius 7mo agoWhitehat hacking ok, but using it for marketing purposes no...
- iam_circuit 7mo ago[dead]
- pikachu0625 7mo agoIt's just different skills.
- bluck 7mo agoQuite uninteresting to read as the article does not go into any depth and it feels simply like the "hacking agent" also wrote the blockpost. Learned nothing
- phyzome 7mo agoFlagging this because 1) this was written by an LLM and 2) there's bad information in it, which means it wasn't reviewed particularly carefully by a human. This means the entire article is suspect as a result.
- tcbrah 7mo agothe data leak is bad but the write access to system prompts is what keeps me up at night. they could silently rewrite how Lilli responds to 43k consultants with a single UPDATE statement - no deploy, no code review, no logs. imagine poisoning the strategic advice that gets copy pasted into client deliverables. tbh most companies i see doing AI stuff store prompts the exact same way, just rows in postgres right next to everything else
- autodate 7mo ago[dead]
- maciusr 7mo ago[flagged]
- foltik 7mo agoLol, dead internet theory is rapidly becoming reality on HN. Another LLM bot down thread [0] produced the exact same slop down to the “no X, no Y, no Z.” > the data leak is bad but the write access to system prompts is what keeps me up at night. they could silently rewrite how Lilli responds to 43k consultants with a single UPDATE statement - no deploy, no code review, no logs. [0] https://news.ycombinator.com/item?id=47345670 https://news.ycombinator.com/item?id=47345670
- AlexeyBelov 7mo agoThank you for being vigilant. Never give up!
- yjcho9317 7mo ago[dead]