Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andrew_zhong
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Show HN: Resurf – realistic, reproducible test framework for AI browser agents
(github.com)
5 points
by
andrew_zhong
5mo ago
|
0 comments
2.
▲
Show HN: HTML to Markdown with CSS selector & XPath annotations for LLM Scraper
(github.com)
4 points
by
andrew_zhong
6mo ago
|
0 comments
3.
▲
by
andrew_zhong
6mo ago
You can use a browser automation library for rendering JS + interaction (like click collapse button) and then use this library to extract the HTML after interaction. Here is an example of using a AI browser automation library with prompt to
4.
▲
by
andrew_zhong
6mo ago
[Update]] I will replace the stealth browser with plain playwright and remove anti-bot as a feature.
5.
▲
by
andrew_zhong
6mo ago
I hear you loud and clear - will replace the stealth browser with plain playwright and remove anti-bot as a feature.
6.
▲
by
andrew_zhong
6mo ago
[Update] I will replace the stealth browser with plain playwright and remove anti-bot as a feature.
7.
▲
by
andrew_zhong
6mo ago
I agree that When pages have similar structure, for one time extraction as it is (not reasoning from context), scraping with selectors is the way to go. This library also supports HTML as input so running a browser is not required.
8.
▲
by
andrew_zhong
6mo ago
What kind of LLMs are you using? In structured output mode? In this library we recover nullable and optional fields, invalid elements in nested array, bad urls, repair incomplete JSONs. If these issues are what you see, yes it should work f
9.
▲
by
andrew_zhong
6mo ago
Put things to perspective - Gemini 2.5 flash is 0.3/1M tokens - assuming each page is 700 tokens and output is not much you are looking at $210 for 1M pages
10.
▲
by
andrew_zhong
6mo ago
Agreed - in this project I did a one path sanitation to recover invalid optional / nullable fields or discard invalid objects in nested array. I know multi path LLM approaches exist: e.g. generating JSON patches https://gith
11.
▲
by
andrew_zhong
6mo ago
We do see fewer invalid JSONs on latest bigger LLMs but still can happen on smaller and cheaper models. There is also case when input is truncated or a required field not found, which are inherently difficult. On XML vs JSON, I think the go
12.
▲
by
andrew_zhong
6mo ago
HTML -> markdown -> LLM is standard practice. We strip elements like aside, embed, head , iframe etc. the criteria is conservatively set to avoid removing too many elements (especially in extractMain mode) https://github.co
13.
▲
by
andrew_zhong
6mo ago
In context of e-commerce web extraction, invalid JSON can occur especially in edge cases, for example: price: z.number().optional() -> price: “n/a” url: z.string().url().nullable() -> url: “not found” It can also be one invalid o
14.
▲
by
andrew_zhong
6mo ago
I will add a PR to enforce robots.txt before the actual scraping.
15.
▲
by
andrew_zhong
6mo ago
We do respect robots.txt production - also scraping browser providers like BrightData enforces that. I will add a PR to enforce robots.txt before the actual scraping.
16.
▲
by
andrew_zhong
6mo ago
The anti-bot patches here (via Patchright) are about preventing the browser from being detected as automated — fixing CDP leaks, removing automation flags, etc. For sites behind Cloudflare or Datadome, that alone usually isn't enough —
17.
▲
by
andrew_zhong
6mo ago
Yeah that's a good observation. XML's closing tags give the model structural anchors during generation — it knows where it is in the nesting. JSON doesn't have that, so the deeper the nesting the more likely the model loses t
18.
▲
by
andrew_zhong
6mo ago
Good point. The anti-bot patches here (via Patchright) are about preventing the browser from being detected as automated — things like CDP leak fixes so Cloudflare doesn't block you mid-session. It's not about bypassing access res
19.
▲
by
andrew_zhong
6mo ago
Good point. The anti-bot patches here (via Patchright) are about preventing the browser from being detected as automated — things like CDP leak fixes so Cloudflare doesn't block you mid-session. It's not about bypassing access res
20.
▲
Show HN: Robust LLM extractor for websites in TypeScript
(github.com)
72 points
by
andrew_zhong
6mo ago
|
50 comments
21.
▲
Show HN: Robust LLM Extractor for HTML/Markdown in TypeScript
(github.com)
14 points
by
andrew_zhong
1y ago
|
0 comments
22.
▲
by
andrew_zhong
2y ago
Thank you for the feedback. Our intention is to let users try for 10 days without commitment (no credit card required), so that after 10 days they can choose to subscribe or not. I see your confusion on pricing. It will be a paid service gi
23.
▲
Show HN: LLM-powered News Hub, e.g "Games, science or open source" on HN
(lightfeed.ai)
5 points
by
andrew_zhong
2y ago
|
3 comments
24.
▲
by
andrew_zhong
2y ago
Have you used them? They offers it for free and I don't find their website mentioning API limit or whether they support dynamic JS
25.
▲
by
andrew_zhong
2y ago
Great work! I’ve worked on the same problem and used LLM to extract feed into structured data (in my case have to use a more affordable model like GPT3.5 for a Saas app, looking at llama3 now) Have you thought about automatically extract sc
26.
▲
by
andrew_zhong
2y ago
I’ve worked on this exact problem when extracting feeds from news websites. Yes calling LLM each time is costly so I use LLM for the first time to extract robust css selectors and the following times just relying on those instead of incurri
27.
▲
by
andrew_zhong
2y ago
If you are looking for broader less heuristic based filter, you can try LightFeed (full disclosure, I am the creator of LightFeed) It filters and summarizes stories using LLM and your prompt on any site including HN!
28.
▲
Ask HN: Best practice to use Llama 3 8B on production server
3 points
by
andrew_zhong
2y ago
|
1 comments
29.
▲
by
andrew_zhong
2y ago
I got reply from someone in Reddit who mentioned that feed readers will just ignore the custom namespaces. SO I think the best option (but hacky) will be to put the summary into the item description HTML... Hope there are cleaner ways
30.
▲
Ask HN: Where to put AI summary and filter in custom RSS feed
4 points
by
andrew_zhong
2y ago
|
1 comments
More ›