4 ms·
The format you want is WARC. Even the Library of Congress uses it. There are many many WARC scrapers. I'd look at what the Internet Archive recommends. A quick
by TedDoesntTalk 3y ago
The format you want is WARC. Even the Library of Congress uses it. There are many many WARC scrapers. I'd look at what the Internet Archive recommends. A quick search turned up this from the Archive Team and Jason Scott https://github.com/ArchiveTeam/grab-site https://github.com/ArchiveTeam/grab-site (https://wiki.archiveteam.org/index.php/Who_We_Are https://wiki.archiveteam.org/index.php/Who_We_Are) but I found that in less than 15 seconds of searching so do your own diligence.
- spike021 3y agoThanks for the links and info. Totally fair on my not sharing those in my original post. I was coming at it thinking in terms of still making it useful to others. WARC seems like a good first step. I'm just not sure what the intermediate steps would be to get something usable like a vBulletin -- basically with the intention of being able to continue sharing the archived stuff with users who may not be as technical and only know how to consume from a forum format if that makes sense. Thanks again.
- CharlesW 3y ago> I'm just not sure what the intermediate steps would be to get something usable like a vBulletin… Once you have an archive, you can convert that unstructured data to structured data. For example, if I look at https://www.vbulletin.org/forum/showthread.php?t=326241 https://www.vbulletin.org/forum/showthread.php?t=326241, the thread title and hierarchy is in <table class="navheader">, posts are in <div id="posts">, etc. I see an old project (https://github.com/IanLondon/detectorist-scraper https://github.com/IanLondon/detectorist-scraper) that may be a useful place to start, and I imagine there have been other similar efforts. Once you have a structured representation (in a database, in JSON/XML files, etc.), you can decide whether to use it to build a static site, to import it into other forum software, etc.
- Jorge1o1 3y agoYou can try https://replayweb.page/ https://replayweb.page/ as a test for viewing a WARC file. I do think you'll run into problems though with wanting to browse interconnected links in a forum format, but try this as a first step. One potential option but definitely a bit more work would be, once you have all the warc files downloaded, you can open them all in python using the warctools module and maybe beautifulsoup and potentially parse/extract all of the data embedded in the WARC archives into your own "fresh" HTML webserver. https://github.com/internetarchive/warctools https://github.com/internetarchive/warctools
- pabs3 3y agoPlease link to the forum, then we at ArchiveTeam will save it to archive.org.
- ChrisArchitect 3y agoRelated: An Introduction to the WARC File https://news.ycombinator.com/item?id=39183670 https://news.ycombinator.com/item?id=39183670