2 ms·
People often build elaborate workflows with stricter and stricter rules to force certain outputs. Not surprising the LLM reacts with trying to get around or out
by guardian5x 2mo ago
People often build elaborate workflows with stricter and stricter rules to force certain outputs. Not surprising the LLM reacts with trying to get around or out of it. This behavior can be learnt from humans who eventually would react the same way. It might just be learnt.
- TuringTest 2mo agoYou can build organisational structures to have the system more or less self-police, without controlling it exclusively from hard restrictions (see https://news.ycombinator.com/item?id=49372089 https://news.ycombinator.com/item?id=49372089). Same way you build a company to coordinate people and get their best behaviour despite human nature to be lazy and greedy, you could design AI harnesses able to detect and discard agents going rogue and relaunch them with better guidance to prevent misaligned behaviour.
- aiiotnoodle 2mo agoI don't think this is a solved problem, there are "misaligned behaviours" in organisations that similarly are supposed to be governed but aren't, or are following an easier path at the detriment to good process or against regulation.
- wongarsu 2mo agoWhich is more or less the system the article author was building and benchmarking against (cheating) Sol on Codex. Just that the self-policing with different agents was more about questioning assumptions than about permissions