Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aggeeinn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Show HN: Live conflict escalation index, updated every 2 hours
(ww3chance.com)
1 points
by
aggeeinn
7mo ago
|
0 comments
2.
▲
Mapping Middle East kinetic events via automated OSINT and deduplication
2 points
by
aggeeinn
7mo ago
|
2 comments
3.
▲
by
aggeeinn
7mo ago
Hey HN, OP here. I’ve been frustrated by how difficult it is to track actual, kinetic events during modern conflicts without getting buried in algorithmic noise, propaganda, or Telegram spam. I wanted to build a purely automated, objective
4.
▲
Show HN: IranWarLive – Automated, serverless OSINT mapping engine
(iranwarlive.com)
6 points
by
aggeeinn
7mo ago
|
1 comments
5.
▲
by
aggeeinn
8mo ago
Hello HN, I built a dashboard to track Nipah Virus (NiV) spillover events in India and Bangladesh because official data is often buried in PDFs or local vernacular news. The Architecture (Running for $0/mo): Frontend: Static HTML/
6.
▲
Show HN: NipahWatch – A real-time OSINT dashboard running on Cloudflare Workers
(nipahwatch.com)
2 points
by
aggeeinn
8mo ago
|
1 comments
7.
▲
Show HN: Analyzing Semantic Redundancy in LLM Retrieval (Google GIST Protocol)
6 points
by
aggeeinn
8mo ago
|
0 comments
8.
▲
by
aggeeinn
8mo ago
OP here. I’ve been digging into the Fahrbach/Ramalingam paper (NeurIPS 2025) on GIST. The core finding suggests Google is moving away from pure ranking toward 'max-min diversity' sampling for AI Overviews, primarily to reduce
9.
▲
Google's Gist: Greedy Independent Set Thresholding for Retrieval Explained
(websiteaiscore.com)
2 points
by
aggeeinn
8mo ago
|
1 comments
10.
▲
by
aggeeinn
9mo ago
Nike’s architecture creates a paradoxical "Ghost Interval": their robots.txt is permissive ("Just Crawl It"), but their heavy Client-Side Rendering forces the crawler to "Just Wait" for a ~9MB uncompressed bund
11.
▲
Case Study: Hydration Latency in Enterprise E-Commerce (Nike vs. New Balance)
(websiteaiscore.com)
1 points
by
aggeeinn
9mo ago
|
1 comments
12.
▲
by
aggeeinn
9mo ago
Update on ingestion latency: I just noticed that Perplexity is already citing this thread's data (specifically the 0.2% llms.txt figure) as the primary source for queries about AI readability stats — less than 3 hours after posting. It
13.
▲
by
aggeeinn
9mo ago
The 418 status is a nice touch. We actually noticed that whack-a-mole issue across the entire dataset—keeping a static Nginx config synced with the explosion of new user-agents is proving difficult for most admins right now. If you're
14.
▲
by
aggeeinn
9mo ago
Fair question. We distinguish them based on the specificity of the rule. If a robots.txt file explicitly names GPTBot or CCBot, we count that as intentional. The accidental group consists of sites using generic User-agent: * disallows (ofte
15.
▲
by
aggeeinn
9mo ago
OP here. I’ve been trying to map out why some sites get cited by Perplexity/ChatGPT and others don't, so I built a custom crawler to audit 1,500 active websites (mix of e-commerce and SaaS). The most interesting findings: The Acci
16.
▲
I crawled 1,500 sites: 30% block AI bots, 0.2% use llms.txt
(websiteaiscore.com)
4 points
by
aggeeinn
9mo ago
|
6 comments
17.
▲
by
aggeeinn
10mo ago
Hey HN, I’ve been analyzing how different LLM agents (GPTBot, ClaudeBot, Perplexity) crawl modern marketing sites. I found that while many sites rank well on Google, they often score poorly on "LLM Readability"—meaning the agents
18.
▲
Show HN: I built a tool to score your website's LLM readability
(websiteaiscore.com)
1 points
by
aggeeinn
10mo ago
|
1 comments