3 ms·
These system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output
by dmix 2mo ago
These system prompts are not the only safety layer that these models use. There's other more deterministic filters in place both on input and (streaming) output.
- lucisferre 2mo agoI think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much. "Make no mistakes"
- dmix 2mo agoYes that's been obvious since the beginning. That's why you should always monitor your agents closely. Just like supervised self driving cars, you have to watch the road and do some hand holding. The tooling around isolation, logging, and real time security/anonomly detection for regular LLM laptop users is very immature right now. I expect that to change soon. The alternative is extremely locked down models which is what Anthropic seems to want to do.
- xmprt 2mo ago> Yes that's been obvious since the beginning But if it's so obvious, then why are we still relying on it in the system prompt. It's just wasting context at this point.
- paxys 2mo agoIt’s equivalent to having client-side input validation. Yes it can easily be bypassed, but in the vast majority of cases where users aren’t malicious it gets the job done quickly and cheaply.
- akshitgaur2005 2mo agoBut isn't the entire point of that system prompt to stop the malicious users. The majority of users are not going to ask those requests anyway.
- metek 2mo agoA locked door stops the lazy thieves, and the lazy thieves are the most common ones.
- 8note 2mo agoit is a heuristic though, and can be measured as such. my steel yield strength table is similarly not guaranteed to be correct for the piece of steel that I have in front of me.
- Hoasi 2mo agoAs reliable as telling a pachinko machine: don’t lose my money!
- deleted 2mo ago[deleted]
- trompetenaccoun 2mo agoIs there concrete evidece that those are xAI's default prompts anyway? They seem plausible enough but how would company outsiders know?
- boorang 2mo agoyou can just look at the traffic in mitmproxy.
- metek 2mo agoI've spent the last year working as an annotator/evaluator for DataAnnotation. All the frontier/flagship model providers use independent contractors for iterating on their LLMs. I'm not able to tell you which models I've worked on as a term of my NDA. The system prompt seems plausible, but in my experience they are much much much much longer and more verbose.
- Melatonic 2mo agoWhat layer do those typically work at and how ? And how did you get into the field ?