Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
angelhadjiev
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
angelhadjiev
2mo ago
The question, whilst rhetorical, supposes we have an alternative.
2.
▲
by
angelhadjiev
3mo ago
"(...) this incident, combined with users’ initial skepticism about Google’s practices regarding user data, likely won’t make too many people wave at the camera anytime soon." ;]
3.
▲
Pay-per-Crawl
(foura.ai)
2 points
by
angelhadjiev
3mo ago
|
0 comments
4.
▲
by
angelhadjiev
4mo ago
also interesting benchmark on fingerprinting: https://www.reddit.com/r/webscraping/comments/1tq1es9/browse...
5.
▲
Proxy Pool Size Means Nothing in 2026
(foura.ai)
2 points
by
angelhadjiev
4mo ago
|
1 comments
6.
▲
Post-quantum TLS rolled out last January and broke most open-source scrapers
(foura.ai)
5 points
by
angelhadjiev
5mo ago
|
0 comments
7.
▲
by
angelhadjiev
6mo ago
;] Not a bot - just sleep-deprived. Spent last night chasing a tarpit at 2am. You're right that scraping has a bad reputation (still, although it's one of the top topics on google words), and some of it is well-deserved. The moral
8.
▲
by
angelhadjiev
6mo ago
Fair point. Direct outreach works when you can identify who to contact and they’re responsive. In practice though, most data teams are scraping hundreds of domains, not one. The hostmaster path doesn’t scale, and tarpits often get deployed
9.
▲
Web scraping tarpits are catching legitimate data teams, not just AI crawlers
(foura.ai)
4 points
by
angelhadjiev
6mo ago
|
5 comments
10.
▲
by
angelhadjiev
6mo ago
Sites are deploying infinite fake-page mazes (Nepenthes, Locaine, etc.) to trap and poison AI training crawlers that ignore robots.txt. The motivation is understandable — Cloudflare reported 75% of AI web traffic in mid-2025 was training-re