Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
csemple
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
csemple
9mo ago
Yep, I use em dashes all the time—still a human typing this.
2.
▲
by
csemple
9mo ago
If a tool is in the context window, the model assigns a non-zero probability to using it. By filtering it out upstream, you entirely remove that path from the inference tree. Instead of asking the model to ignore an affordance, you remove t
3.
▲
by
csemple
9mo ago
I guess working in government has put me ahead of the curve sounding like a robot.
4.
▲
by
csemple
9mo ago
Yes, thanks. That's the one-sentence summary.
5.
▲
by
csemple
9mo ago
You’re totally right—it's ultimately just probabilistic tokens. I’m thinking that by physically removing the tool definition from the context window, we avoid state desynchronization. If the tool exists in the context, the model plans
6.
▲
by
csemple
9mo ago
Ya, makes sense—if the model is trained just to "be helpful," removing the tool forces it to improvise. I’m thinking this is where the architecture feeds back into the training/RLHF. We train the model to halt reasoning in th
7.
▲
by
csemple
9mo ago
I appreciate the feedback. Let me address the key technical point: On enforcement mechanism: You've misunderstood what the system does. It's not asking the LLM to determine security. The Capacity Gate physically removes tools befo
8.
▲
by
csemple
9mo ago
Yes, on your first point "layer 1" isn't fundamentally new. It's applying standard systems administration principles, because we're currently trusting prompts to do the work of permissions. With the pattern I'm
9.
▲
by
csemple
9mo ago
Yep, you nailed the problem: context drift kills instruction following. That's why I’m thinking authority state should be external to the model. If we rely on the System Prompt to maintain constraints ("Remember you are read-only&
10.
▲
by
csemple
9mo ago
You're exactly right—treating the LLM as an untrusted user is the security baseline. The distinction I'm making is between Execution Control (Firewall) and Cognitive Control (Filter). Standard RBAC catches the error after the mode
11.
▲
by
csemple
9mo ago
OP here. *** I'm seeing comments about AI-generated writing. This is my voice—I've been writing in this style for years in government policy docs. Happy to discuss the technical merits rather than the prose style. *** At Ontario
12.
▲
Why Ontario Digital Service couldn't procure '98% safe' LLMs (15M Canadians)
(rosetta-labs-erb.github.io)
40 points
by
csemple
9mo ago
|
48 comments