4 ms·
hey there Joey here from Archestra. Good question. I recently was evaluating what you mention, against the latest/"smartest" models from the big LLM providers,
by joeyorlando 1y ago
hey there
Joey here from Archestra. Good question. I recently was evaluating what you mention, against the latest/"smartest" models from the big LLM providers, and I was able to trick all of them.
Take a look at https://www.archestra.ai/blog/what-is-a-prompt-injection https://www.archestra.ai/blog/what-is-a-prompt-injection which has all the details on how I did this.
- magicalhippo 1y agoThanks. Interesting and scary such blatant attempts succeed. After all, all external data is evil, we all know that right?
- ildari 1y agoexternal data is unavoidable for the properly functioning agent, so we have to learn to cook it
- magicalhippo 1y agoTrue, however this seems like such basic stuff. Download arbitrary text and inject it into your prompt? Why on earth would you not consider that as a very dangerous operation that needs to be carefully managed? It's like parking your bike downtown hoping it wont be stolen. Like, at least use a zip tie or something. That said, I agree with your post that this won't catch everything. So something else, like a quarantined LLM like you suggest is likely needed. However I just didn't expect such blatant attacks to pass.