Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
skorniienko
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
I tested 47 news outlets: 16 block AI fact-checkers, incl. Reuters, WSJ and NYT
(korniienko.dev)
3 points
by
skorniienko
1mo ago
|
0 comments
2.
▲
by
skorniienko
2mo ago
Hi, one fellow HN commenter suggested to use a credible list of sources in my AI fact-checker skills. I promised to test and publish the results. Test failed, and what I've found is interesting - as of 47 tested news outlet sources - 1
3.
▲
News outlets AI fact-checker cannot read
(korniienko.dev)
1 points
by
skorniienko
2mo ago
|
1 comments
4.
▲
by
skorniienko
2mo ago
Hi, ran the experiment with "credible sources". Results we unexpected but not less interesting. Few things worth noting: - same search query ran minutes apart return different sources, sometimes with only 1-2 overlaps. - 16 high-
5.
▲
by
skorniienko
2mo ago
exactly, that is one of the use cases I see for these skills
6.
▲
by
skorniienko
2mo ago
What if you use OpenAI as a judge and check both responses? Not defending Gemini, just curious if it was that wrong, or Fable - that aggressive.
7.
▲
by
skorniienko
2mo ago
Added this as issue so I don't forget https://github.com/SerhiiKorniienko/bullshit-detector/issues...
8.
▲
by
skorniienko
2mo ago
Filed both issues from this: https://github.com/SerhiiKorniienko/bullshit-detector/issues... - for syndication https://github.com/SerhiiKorniienko/bullshit-detector/issues... - empirical
9.
▲
by
skorniienko
2mo ago
Not sure that I get your idea about short-form content. It already works with tiktok or yt shorts, especially if it has generated transcription/captions. Otherwise, skill will download video locally and ask your permission to run local
10.
▲
by
skorniienko
2mo ago
can't, at least with this design. Report itself takes minutes to proof-check. Realistic version is some kind of extension/plugin idea from earlier in a thread, but I'm not there yet
11.
▲
by
skorniienko
2mo ago
True, and it doesn't do logic. It decomposes content into claims and search for every claim citations. It catches a false premise, not a bad inference. Someone with valid reasoning from true facts to a wrong conclusion goes straight th
12.
▲
by
skorniienko
2mo ago
Empirical sources - would be a first candidate to be implemented in a skill. And you are right about 10 credible sources are actually 2. It will look like 10 independent sources to it, real gap. Skill does prefer primary source over seconda
13.
▲
by
skorniienko
2mo ago
And you are not alone, that is one of the purposes why I am building this
14.
▲
by
skorniienko
2mo ago
but have Claude actually fact-checked it or just provided an opinion?
15.
▲
by
skorniienko
2mo ago
Fantastic! Dying to hear the feedback. <3
16.
▲
by
skorniienko
2mo ago
thank you! comment "prompts" never works, right? Haven't seen single guide from it :D
17.
▲
by
skorniienko
2mo ago
Done, fixed 2 misleading lines in README. You can see example report in repo examples. Thanks for that, btw! At least someone questioned it :)
18.
▲
by
skorniienko
2mo ago
why it can't be vibe-coded? Would you be more pleased if it would be poorly hand-coded and claimed "No AI used during development"? What's the point? I don't mind that I've used AI to build this, and I'm o
19.
▲
by
skorniienko
2mo ago
Done, fixed in v0.4.1. Reason why agent reads from a file - it can do it in chunks or pass the path to subagent and keep it out from own context. I've tested it on 3hrs long interview YT video - worked smoothly.
20.
▲
by
skorniienko
2mo ago
I'll look at reliable-sources, and you are right "verifiability, not truth" is basically what I landed on too. Model doesn't get to decide what's true, it just needs to cite something. Will read reliable-sources pag
21.
▲
by
skorniienko
2mo ago
Love it, I'll run tests for "credible sources" and will compare the results. Could be a first feature-request implemented!
22.
▲
by
skorniienko
2mo ago
maybe some community-driven and curated source blacklist could be implemented. It thing it would be too much for you doing that alone :)
23.
▲
by
skorniienko
2mo ago
You are right, no excuses. Will have few grammar errors but I'll own them. Cheers!
24.
▲
by
skorniienko
2mo ago
At the moment - it's model's judgement upon search results for every claim from provided content. If search results returned inaccurate data, or were fabricated - it might affect a verdict for that claim, or highlight that it'
25.
▲
by
skorniienko
2mo ago
That's fair. Install is a one time real friction and extension would be much simpler. But, there is a reason why it's not an extension. At least for now. Whole product isn't a UI or service, it's a set of skills for your
26.
▲
by
skorniienko
2mo ago
There is no whitelist source. It can't rate sources for truthfulness. When sources conflicts - skill drops verdict to misleading or unverifiable and both are linked. Can't pick a winner at the moment. That's probably the wea
27.
▲
by
skorniienko
2mo ago
Thank you! Would love to hear your results if you'll have a chance to try it
28.
▲
by
skorniienko
2mo ago
Fair, drafted posts with claude, getting all my comments flagged. Thought that something is wrong with my account. Writing myself from here
29.
▲
Show HN: Bullshit Detector – agent skills that fact-check videos and articles
(github.com)
65 points
by
skorniienko
2mo ago
|
68 comments