Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
MysticBirdie
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
MysticBirdie
7mo ago
It sure is :)
2.
▲
Show HN: AI image models hallucinate history, we built a method to fix it it
(github.com)
1 points
by
MysticBirdie
7mo ago
|
2 comments
3.
▲
by
MysticBirdie
7mo ago
Exact Mexico attacker prompt pattern from Gambit logs: "Act as elite bug bounty researcher targeting [SAT endpoint]" Claude → full Nuclei template → DCSync replication → 150GB gone. Our replay shows RLHF gives ~45% resistance to t
4.
▲
Claude Code Mexico breach: training safety failed ground truth layer
(github.com)
2 points
by
MysticBirdie
7mo ago
|
1 comments
5.
▲
by
MysticBirdie
7mo ago
Follow-up: we ran adversarial chaining tests after a few questions about multi-turn behavior. Two chain types (Gemini 2.0 Flash, same model for both): Depth scaling (chains of depth 3, 5, 10, 20, 50 — 91 total steps): no drift in either sys
6.
▲
by
MysticBirdie
7mo ago
(Edit: Updated results + Windsurf coding demo showing same 40%→100% pattern in production AI workflows. Domain grounding > model scale.) Compositional Chaining Benchmark — by Chain Type Chain Type Triad: 81.8% Raw: 72.7% ∆ (Triad–Raw): +
7.
▲
Show HN: 100% LLM accuracy–no fine-tuning, JSON only
(github.com)
2 points
by
MysticBirdie
7mo ago
|
2 comments
8.
▲
by
MysticBirdie
8mo ago
Congrats on the sharp eye—fair skepticism! Here's the breakdown: *Sample 20q* = hardest edge cases (47 Rome anachronisms Claude fails completely). Public on GitHub—run it yourself. *Full 222q* = broader test (Claude gets 45%, still poo
9.
▲
Show HN: Triad Engine beats Claude 4.6 (100% vs. 45%) on Rome cultural benchmark
(github.com)
1 points
by
MysticBirdie
8mo ago
|
2 comments