Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wizeyone
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
wizeyone
5mo ago
Yes, we do. Dr Aigerim Bissenova, our CMO. Internal medicine residency at Mass General and digital health fellowship. She designed the panel selection and reviewed all outputs before publication
2.
▲
by
wizeyone
5mo ago
2 things. The headline math 0.95^20 = 0.358 assumes independent errors. "The body argues the opposite - every subsequent action operates on flawed foundations." Real long chain failure is worse than the math predicts, not equal to
3.
▲
by
wizeyone
5mo ago
Author here - we're the team behind Wizey, one of the two AIs in the comparison. A few things up front: * Methodology was fixed before the runs. * All outputs are quoted verbatim, including Case 2 (MGUS) where ChatGPT beat us cleanly.
4.
▲
ChatGPT vs. a specialized medical AI on 5 clinical cases (verbatim outputs)
(wizey.one)
3 points
by
wizeyone
5mo ago
|
3 comments
5.
▲
by
wizeyone
5mo ago
Arms race framing misses it. Insurers have used algorithmic denial scoring for years (ProPublica/Cigna-EviCore, StatNews/UnitedHealth-NaviHealth). Denial works because appealing is expensive for patients and near-free for insurers
6.
▲
by
wizeyone
5mo ago
"Spending more on AI than humans" tells you nothing about whether it works. Cost per-output is the metric and by that I've watched startups do worse than last year, just more expensively. Feels like investor signal: "we
7.
▲
by
wizeyone
5mo ago
"AI-generated and approved by engineers" is doing a lot of work there. If accepting a 4-character Gemini autocomplete counts, Copilot users hit >90% last year. The useful metric is % of functions where >50% of the body was A