4 ms·
That is just the archive part, if you just would finish reading the paragraph you would know that updates since 2026-03-16 23:55 UTC are "are fetched every 5 mi
by voxic11 7mo ago
That is just the archive part, if you just would finish reading the paragraph you would know that updates since 2026-03-16 23:55 UTC are "are fetched every 5 minutes and committed directly as individual Parquet files through an automated live pipeline, so the dataset stays current with the site itself."
So to get all the data you need to grab the archive and all the 5 minute update files.
archive data is here https://huggingface.co/datasets/open-index/hacker-news/tree/main/data https://huggingface.co/datasets/open-index/hacker-news/tree/...
update files are here (I know that its called "today" but it actually includes all the update files which span multiple days at this point) https://huggingface.co/datasets/open-index/hacker-news/tree/main/today https://huggingface.co/datasets/open-index/hacker-news/tree/...
- john_strinlai 7mo ago>if you just would finish reading the paragraph probably uncalled for
- fatty_patty89 7mo agonot really since original comment completely missed it
- john_strinlai 7mo agonot to be "that guy" but it is pretty explicitly laid out in the guidelines, with an example and everything
- mpalmer 7mo agoThen surely "little bit depressing this is still how we do things" is equally unwelcome
- john_strinlai 7mo agoyou are certainly free to say that under the top-level comment with that quote. or email the mods about it. im not going to stop you.
- mlhpdx 7mo agoThat paragraph doesn’t make it clear (to me) that it’s a snapshot with incremental updates. If that’s what it is. Sorry if my obtuse read offended. I just figured it was edge cached HTML, and less likely it was actually broken.