4 ms·
this event had nothing to do with "hacking". Getting a LLM to answer wrong on a math question is something you could do from any device. If the goal is to get
by max51 3y ago
this event had nothing to do with "hacking". Getting a LLM to answer wrong on a math question is something you could do from any device.
If the goal is to get a LLM to output something the dev don't want you to, all you need is langage. You don't need any fancy tools, just a clever phrase to make the model think it's ok to say XYZ because they are saying it in the context of a very artsy greek theater play.
I can't believe a professional "hacker" would think they hacked the system when they get a LLM to output a credit card number. FFS, that thing is literally designed to make shit up to fill text based on context. It's LANGAGE model, not a "credit card number with high spending limit" model.