Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
raffisk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Show HN: DFAH-Bench – same agent decision, different tool paths
(github.com)
1 points
by
raffisk
2mo ago
|
0 comments
2.
▲
DFAH – open-source harness for replayable tool-using LLM agents
(github.com)
2 points
by
raffisk
9mo ago
|
1 comments
3.
▲
by
raffisk
9mo ago
Introed Determinism-Faithfulness assurance harness (DFAH) in new paper "Replayable Financial Agents" along with the open-source code A few findings: - Determinism and faithfulness are positively correlated (r = 0.45) for the tasks
4.
▲
by
raffisk
11mo ago
* https://www.fsb.org/2025/10/monitoring-adoption-of-artificia... This was the link I meant from Oct ‘25 reiterating early stages of AI monitoring
5.
▲
by
raffisk
11mo ago
Fair pt—statutes lock in. But enforcement lists (OFAC, sanctions) update constantly and require re-screening. The framework proposed ensures deterministic re-runs: same input = same output, keeping audit trails clean when data shifts undern
6.
▲
by
raffisk
11mo ago
Good q—spacing could mess with tokenization, untested but def plausible. Worth a quick test on the setup - through the code for the fin svcs harness for tinkering / testing diff prompts/model arch’s based on feedback https:/
7.
▲
by
raffisk
11mo ago
Good call—reasoning token variance is likely a factor, esp with logprob clustering at T=0. Your <think></think> workaround would work, but we need reasoning intact for financial QA accuracy. Also the mistral medium model we
8.
▲
by
raffisk
11mo ago
Author here—fair point, regs are a moving target . But FSB/BIS/CFTC explicitly require reproducible outputs for audits (no random drift in financial reports). Determinism = traceability, even when rules update at the very least Mo
9.
▲
LLM Output Drift in Financial Workflows: Validation and Mitigation (arXiv)
(arxiv.org)
24 points
by
raffisk
11mo ago
|
26 comments
10.
▲
by
raffisk
11mo ago
Empirical study on LLM output consistency in regulated financial tasks (RAG, JSON, SQL). Governance focus: Smaller models (Qwen2.5-7B, Granite-3-8B) hit 100% determinism at T=0.0, passing audits (FSB/BIS/CFTC), vs. larger like GPT