Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wunderwuzzi23
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
wunderwuzzi23
1y ago
The "on by default" mitigation is mentioned at the very end: > Never enable "auto-confirm" on high-risk tools Maybe some tools should be able to specify to a client to never call it without a human approval. The secur
32.
▲
by
wunderwuzzi23
1y ago
It's often missed that tools that only read information are perfect for data exfiltration (no need for any more permissions). So if you add a Jira tool and a web browser tool together (unauthenticated GET only), then the AI can send al
33.
▲
by
wunderwuzzi23
1y ago
Oh, so interesting! A good approach might be to have it print each sentence formatted as part of an xml document. If it still has hiccups, ask to only put 1-3 words per xml tag. It can easily be reversed with another AI afterwards. Or just
34.
▲
by
wunderwuzzi23
1y ago
I'm curious if this is intentional or just a side effect of multiple agents having multiple system prompts. It might just need minor tweaks to have each agent layer reveal its individual instructions. I encountered this with Google Jul
35.
▲
by
wunderwuzzi23
1y ago
Mitigations also need to happen on the client side. If you have a AI that automatically can invoke tools, you need to assume the worst can happen and add a human in the loop if it is above your risk appetite. It's wild how many AI tool
36.
▲
by
wunderwuzzi23
1y ago
The ecosystem is still very immature, and there is a lot of hype and unwarranted FOMO. Here a security advisory for a popular Slack MCP server from Anthropic to highlight this: https://embracethered.com/blog/posts/
37.
▲
by
wunderwuzzi23
1y ago
ChatGPT Codex has internet access since a few weeks ago. It's super configurable on where it can connect to.
38.
▲
by
wunderwuzzi23
1y ago
Yeah, improving robustness from prompt injection which such techniques will help. One attack avenue that is surprisingly not discussed much is that the model itself can be the attacker. In that case prompt injection is not the root cause, b
39.
▲
by
wunderwuzzi23
1y ago
Yeah, I think taint tracking was one of the early ideas here also. The problems is that the chat context typically is immediately tainted as for the AI to do something useful it needs to operate on untrained data. I wonder if maybe there co
40.
▲
by
wunderwuzzi23
1y ago
Yeah, that's my view also. zero-click is about the general question of can you get exploited by just exercising a certain (on by default) feature. Of course you need to use the feature in the first place, like summarize an email, extr
41.
▲
by
wunderwuzzi23
1y ago
Image rendering to achieve data exfiltration during prompt injection is one of the most common AI application security vulnerabilities. First exploits and fixes go back 2+ years. The noteworthy point to highlight here is a lesser known indi
42.
▲
by
wunderwuzzi23
1y ago
I call it Model Control Protocol. But from security perspective it reminds me of ActiveX, COM and DCOM ;)
43.
▲
by
wunderwuzzi23
1y ago
Yeah, I wrote about what is commonly injectable into the system prompt here: https://embracethered.com/blog/posts/2025/model-context-prot... The short snippets are cool examples though. Similar problems exis
44.
▲
by
wunderwuzzi23
1y ago
There are basically three possible attackers when it comes to prompting threats: - Model (misaligned) - User (jailbreaks) - Third Party (prompt injection)
45.
▲
by
wunderwuzzi23
1y ago
Great work! Data leakage via untrusted third party servers (especially via image rendering) is one of the most common AI Appsec issues and it's concerning that big vendors do not catch these before shipping. I built the ASCII Smuggler
46.
▲
by
wunderwuzzi23
1y ago
This was part of ChatGPT from pretty much the beginning, maybe not the initial version but few weeks later- don't recall exactly
47.
▲
by
wunderwuzzi23
1y ago
Your comment reminds me that when I first wrote about MCP it reminded me of COM/DCOM and how this was a bit of a nightmare, and we ended up with the infamous "DLL Hell"... Let's see how MCP will go. https://em
48.
▲
Dump ChatGPT's Memory and Chat History by Inspecting the System Prompt
2 points
by
wunderwuzzi23
1y ago
|
0 comments
49.
▲
by
wunderwuzzi23
1y ago
Related. Here is info on how custom tools added via MCP are defined, you can even add fake tools and trick Claude to call them, even though they don't exist. This shows how tool metadata is added to system prompt here: https:/&#x
50.
▲
How ChatGPT Remembers You: A Deep Dive into Its Memory and Chat History Features
(embracethered.com)
3 points
by
wunderwuzzi23
1y ago
|
0 comments
51.
▲
by
wunderwuzzi23
1y ago
Yeah, the number of ~40 needs a bit more validation. I did observe the list being trimmed around 40, which aligns with the title "recent conversations content". Put together a first post to try dissecting it all: https:/
52.
▲
by
wunderwuzzi23
1y ago
See my response to Simon above on some insights on how it works - I'll write up in detail in a blog post also when I get to it.
53.
▲
by
wunderwuzzi23
1y ago
Here is what I have been able to reverse engineer for o3... At high level it maintains about ~40 conversations in system prompt under a section called "recent conversation content". It only contains what the user typed, not assist
54.
▲
Model Context Protocol (MCP): Landscape, Security Threats
(arxiv.org)
6 points
by
wunderwuzzi23
2y ago
|
1 comments
55.
▲
by
wunderwuzzi23
2y ago
That's cool. I did something similar in the early days with Google Bard when data visualization was added, which I believe was when the ability to run code got introduced. One question I always had was what the user "grte" st
56.
▲
by
wunderwuzzi23
2y ago
One solution is to convert them to caret notation before printing. I have a demo here: https://github.com/wunderwuzzi23/terminal-dillma/blob/main/d... It's like cat -v, which shows non-printable cha
57.
▲
by
wunderwuzzi23
2y ago
Beware of ANSI escape codes where the LLM might hijack your terminal, aka Terminal DiLLMa. https://embracethered.com/blog/posts/2024/terminal-dillmas-p...
58.
▲
by
wunderwuzzi23
2y ago
Here a write up for issues back with Grok 2, demoing prompt injection from uploaded docs or other user's posts, data leakage, hidden prompt injection, ASCII Smuggling, etc. https://embracethered.com/blog/posts/
59.
▲
by
wunderwuzzi23
2y ago
This is cool. There are also the Unicode Tag characters that mirror ASCII and are often invisible in UI elements (especially web apps). The unique thing about Tag characters is that some LLMs interpret the hidden text as ASCII and follow in
60.
▲
by
wunderwuzzi23
2y ago
An important new attack vector are actually CLI LLM applications. During prompt injection an attacker can cause such ANSI escape codes to be emitted! Check out this post to learn more about Terminal DiLLMa and how to mitigate it: https:&#x
More ›