Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anulum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
anulum
7mo ago
https://anulum.github.io/director-ai/benchmarks/
2.
▲
by
anulum
7mo ago
@soletta — you're right, and I've deferred this three times now, which isn't useful.
3.
▲
by
anulum
7mo ago
Hey HN — huge thanks for the thoughtful comments yesterday! I shipped *v1.2.0* overnight with everything you asked for: • Full end-to-end benchmark notebook (600+ real RAG/agent traces, HaluEval + TruthfulQA, head-to-head vs Claude sel
4.
▲
by
anulum
7mo ago
@soletta Got it — thanks for the extra clarity, that’s an important distinction. You’re absolutely right: modern frontier models (Claude 3.5/Opus-class, GPT-4o, etc.) have become extremely good at maintaining internal consistency dur
5.
▲
by
anulum
7mo ago
@soletta Great question — this is exactly why we built it this way. *Short answer*: frontier LLMs are excellent at static self-critique, but terrible for *real-time token-by-token streaming guardrails* because of latency, cost, and lack o
6.
▲
Show HN: Director-AI – token-level NLI+RAG
(github.com)
2 points
by
anulum
7mo ago
|
7 comments
7.
▲
Show HN: I just shipped the canonical neuro-symbolic control demo
(github.com)
1 points
by
anulum
7mo ago
|
0 comments
8.
▲
Show HN: SC-NeuroCore – Rust neuromorphic compiler, 512× speedup
(github.com)
2 points
by
anulum
8mo ago
|
0 comments
9.
▲
Show HN: SCPN Fusion Core – Tokamak plasma SIM and neuromorphic SNN control
(github.com)
2 points
by
anulum
8mo ago
|
0 comments