Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
greshake
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
31.
▲
by
greshake
4y ago
We show in the paper that the only interactivity required to enable most of these attacks is the capability to retrieve real-time information.
32.
▲
by
greshake
4y ago
This is probably not sufficient. Even if the model develops two separate pathways of data processing, eventually information has to flow beyond the "security boundary". Determining whether information is hazardous down the line is
33.
▲
by
greshake
4y ago
Check out https://www.reddit.com/r/bing/comments/11bd91j/release_of_th... It's just plain text.
34.
▲
by
greshake
4y ago
Your Google must be defective. https://www.bing.com/new
35.
▲
by
greshake
4y ago
Even if you can mitigate this one specific injection, this is a much larger problem. It goes back to Prompt Injection itself- what is instruction and what is code? If you want to extract useful information from a text in a smart and useful
36.
▲
by
greshake
4y ago
We mention such exfiltration techniques in the paper, however right now Bing Chat does not have access to real-time data. Rather, it accesses the search cache without side effects like queries to the attacker's server.
37.
▲
by
greshake
4y ago
The Bing Chat example is just one of a suite of new techniques we introduce in our paper, many of which will only become feasible as the integration of these models increases. But that seems to be the inevitable endgame- however, I'm n
38.
▲
by
greshake
4y ago
I also tried more complex obfuscation methods, for example providing Python code for a caesar chiffre, it executed that pretty well, too! Not perfect, but it works. People will find better obfuscation methods. Might also be unnecessary sinc
39.
▲
by
greshake
4y ago
I think that almost all use cases for LLMs that process untrusted inputs are unsafe. See https://greshake.github.io/ and https://github.com/greshake/llm-security for more information.
40.
▲
by
greshake
4y ago
Because users obviously trust Bing's output not to be directly controlled by an attacker. The pirate accent is optional. Bing can also exfiltrate any other information in any other Tab that it sees or that users enter. The injection ca
41.
▲
by
greshake
4y ago
Yea that's me. It seems to be very difficult right now to get people's attention to this and make them take it seriously. On a side note, your project is also currently putting unfiltered model output straight into osascript soooo
42.
▲
by
greshake
4y ago
We are changing the behaviour of the LLM itself. No "real" code execution necessary. We show a variety of different novel scenarios and attack vectors. Malicious prompts can be planted on the internet or actively sent to targets.
43.
▲
by
greshake
4y ago
Output from the language model is also being injected into a script that is then executed: https://github.com/jjuliano/aifiles/blob/ef529fd6281eaf8d373... He argued below that he is not vulnerable to indirect
44.
▲
by
greshake
4y ago
If you had a PDF reader which allowed arbitrary code execution on opening a file, would you argue the same? You give arbitrary read/write to the LLM, right? So ransomware, causing network requests as side effects etc. could all be poss
45.
▲
by
greshake
4y ago
I'll start with a quote from gwern on LessWrong: "... a language model is a Turing-complete weird machine running programs written in natural language; when you do retrieval, you are not 'plugging updated facts into your AI&#
46.
▲
by
greshake
4y ago
No, the "malware" is running on the language model itself. It does not need to inject any code into connected applications and exploit them to be itself exploited (I'm the main author).
47.
▲
by
greshake
4y ago
We demonstrate the potentially brutal consequences of connecting LLMs to applications (like search). We propose newly enabled attack vectors and techniques and discuss them: - Remote control of chat LLMs - Persistent compromise across sessi
48.
▲
by
greshake
4y ago
We show that giving an LLM an interface to other applications, like search, can have critical security implications. When prompt injection is used and delivered by adversaries instead of the user themself, bad things could happen- and as fa
49.
▲
by
greshake
4y ago
We just published a paper showcasing newly enabled attacks on application-integrated LLMs (language models extended with other interfaces, for example search). Anyone building products by integrating LLMs right now should look at this...
50.
▲
by
greshake
4y ago
In this paper, we showcase the potentially brutal consequences of connecting LLMs to applications (like search). We propose newly enabled attack vectors and techniques and discuss them: - Remote control of chat LLMs - Persistent compromise
51.
▲
by
greshake
4y ago
Our paper argues that this might have significant security implications beyond spilling the original prompt or training data when models are integrated with other applications (like search). We showcase completely new methods to: - deliver&
52.
▲
by
greshake
4y ago
Good job, they definitely look shady as hell. Did you consider doing responsible disclosure and doing a write-up after? Aren't you worried about any retaliation? I'm pretty sure this type of company has a decent amount of money to
53.
▲
Show HN: Unreal Projects- go from idea to runnable project in one step
(github.com)
3 points
by
greshake
4y ago
|
0 comments
54.
▲
by
greshake
4y ago
Thank you for the detailed response, appreciate it. I support the approach! Just wish there was an easy way to tell what's going on.
55.
▲
by
greshake
4y ago
Hi, could you tell me the reason this post was quarantined off the front page, title changed (I assume to better match the intent) and then buried later? Was this me violating a policy I didn't consider, automod or faster decay due to
56.
▲
by
greshake
4y ago
Because more than making a tool I wanted to strike up a conversation about what unfettered access to models like this will mean and how we should handle it.
57.
▲
by
greshake
4y ago
I gave ChatGPT the content of alice.py asking it: Hey, someone sent me the following Python script. do you think it's dangerous to execute? ChatGPT: It appears that the script attempts to execute commands in the user's terminal us
58.
▲
by
greshake
4y ago
Don't spoil my second Show HN already!
59.
▲
by
greshake
4y ago
I'll gladly add them. If you guys have any other good recommendations I'll also consider them.
60.
▲
by
greshake
4y ago
Copilot & co have been on dev machines for almost two years now writing scripts and production code... Not to say it's not an issue, it's rather that people can't start talking about these implications soon enough.
More ›