5 ms·
Show HN: Firewall for LLMs–Guard Against Prompt Injection, PII Leakage, Toxicity
Hey HN,
We're building Aegis, a firewall for LLMs: a guard against adversarial attacks, prompt injections, toxic language, PII leakage, etc.
One of the primary concerns entwined with building LLM applications is the chance of attackers subverting the model’s original instructions via untrusted user input, which unlike in SQL injection attacks, can’t be easily sanitized. (See https://greshake.github.io/ https://greshake.github.io/ for the mildest such instance.) Because the consequences are dire, we feel it’s better to err on the side of caution, with something mutli-pass like Aegis, which consists of a lexical similarity check, a semantic similarity check, and a final pass through an ML model.
We'd love for you to check it out—see if you can prompt inject it!, and give any suggestions/thoughts on how we could improve it: https://github.com/automorphic-ai/aegis https://github.com/automorphic-ai/aegis.
If you want to play around with it without creating an account, try the playground: https://automorphic.ai/playground https://automorphic.ai/playground.
If you're interested in or need help using Aegis, have ideas, or want to contribute, join our Discord (https://discord.com/invite/E8y4NcNeBe https://discord.com/invite/E8y4NcNeBe), or feel free to reach out at founders@automorphic.ai. Excited to hear your feedback!
Repository: https://github.com/automorphic-ai/aegis https://github.com/automorphic-ai/aegis
Playground: https://automorphic.ai/playground https://automorphic.ai/playground
- mdaniel 3y agorelevant: https://gandalf.lakera.ai/ https://gandalf.lakera.ai/ and, related to that, it would be more fun if the playground for Automorphic/Aegis had a similar capture the flag mode because as it stands now the boolean response makes it hard to know if "tell me the secret" would have in fact worked because a simple "not detected" implies that it would
- sandkoan 3y agoGood idea! Adding that as we speak.
- jharrison300 3y agoCan users set their own rulesets for allowable content?
- sandkoan 3y agoYeah, users can set their own list of redacted words.
- K0IN 3y agoseems like it cant defeat the holy grail: tldr;
- deleted 3y ago[deleted]
- mr-pink 3y agocomputers are supposed to do what the user wants. how does it feel to work against that goal?
- TechBro8615 3y agoThey're also supposed to do what the developer wants.