10 ms·
Make Your Own Internet Archive with Archive Box
- evc 6y agoYou will need a lot of disk storage right?
- Ace_Archer 6y agoThat probably depends on the scope of what you're looking to archive. If you're looking to make up local backup of your bookmarks folder (as one of the intentions seems to be), probably not an unreasonable amount of storage. Maybe a few GB at most(if you have a moderate to large bookmarks folder), depending on how many sites/heavy the sites are?
- flas9sd 6y agoit doesn't show in the Screenshot in the article, but ArchiveBox in Aug 2020 implemented the "readability article text extractor", see description in the release notes: https://github.com/pirate/ArchiveBox/releases/tag/v0.4.14 https://github.com/pirate/ArchiveBox/releases/tag/v0.4.14 and the module that does the work https://github.com/pirate/readability-extractor https://github.com/pirate/readability-extractor By only extracting text and article images you could go deep into an archive. If you skip images, much more so
- reefab 6y agoFor reference, archivebox uses 250GB for 5000 links in my setup.
- mosselman 6y agoThat is an insane amount of storage for so few links. Is your setup somehow very greedy? Saving article only view (images + text) should probably do better I suspect your numbers come from JavaScript and css, etc? Is there a way for archivebox to not download react 5000 times but share source files? Most likely custom bundles that sites compile will not make this possible most of the time. Just thinking out loud here.
- nikisweeting 6y agoIt's recommended to run it on a compressed filesystem like ZFS. On mine it's using ~75GB for ~3000 URLs. It varies greatly depending on the content, usually the vast majority of storage is from video/audio ripped with youtube-dl.
- LEARAX 6y agoThere are different extractors/services, and you can toggle them pretty easily. By default it screenshots everything, exports a PDF, saves like 4 different HTML copies and submits the link to the wayback machine. It also tries to extract important text, and stores that separately. You could easily configure it to only extract text, turn off some HTML extractors, or disable the PDF and screenshot captures if you want to prioritize disk space.
- remirk 6y agoThis article is blogspam. The repository has enough information on its own: https://github.com/ArchiveBox/ArchiveBox https://github.com/ArchiveBox/ArchiveBox
- blastro 6y agoi use this every single day and think very highly of it. thanks for reminding me - i'm going to sponsor this developer on github...
- m-s-sripati 6y agoIt is the right thought, aligned to the spirit of open source.
- mikece 6y agoThis would be a nice thing to be able to run on a Synology NAS or other kind of device that typically has terabytes of storage.
- greypowerOz 6y agoso.. you CAN have a box that is "the internet"....
- mosselman 6y agoYes Jen
- jedimastert 6y agoIs there a list of web page archive formats I could look at? There are a few things I'd love to do where it would be very handy to have one file per page
- nikisweeting 6y agoThe main archive formats for web content are WARC, ZIM, Memento, and static HTML (e.g. from a tool like wget or Singlefile). If you want 1 page per URL I recommend Singlefile. Lots more info here if you want to compare different software options: https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community#Web-Archiving-Projects https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-...
- lazyjeff 6y agoI feel like a simple automatic capture of timestamp + url + screenshot would already be very useful. This gives you a visual memory of the things you've seen on the web. I've wanted to develop this for a while, as a browser plugin. Being able to skim the past month or two click around the thumbnails would already be amazing. I've wanted to do that many times before to check if my memory was correct, or if a page changed since I last saw it, or figure out when I last saw something online. You don't need a special viewer for it, as your operating system's file explorer can view the screenshots already, and you don't need to set up a crawl. Screenshots also compress well, as webp or png after crunching it.
- jumploops 6y agoI've dreamed about this as well, basically a personalized FullStory that allows you to search and replay all of your sessions across sites. Easy block list for sensitive things like banking, internal sites, email, etc. I currently use the Session Buddy Chrome Plugin, which helps in some cases (I was able to find a hard to Google repo today, for example), but the historic context is largely missing.
- bravura 6y agoThis doesn't allow full text search easily, though.
- 0x426577617265 6y agoI use this with an automated script that watches my Twitter activity. If I like a tweet it determines if it contains a URL then archives it.
- ketamine__ 6y agoHow does archive.is trick news sites into showing content without the paywall? Is it pure user agent spoofing? I'm wondering if this could be applied here.
- nikisweeting 6y agoYeup, just the reason why we expose the USER_AGENT options in ArchiveBox config ;) https://github.com/ArchiveBox/ArchiveBox/wiki/Security-Overview#private-mode https://github.com/ArchiveBox/ArchiveBox/wiki/Security-Overv... I don't want to officially endorse using the Google bot user agent, but you're welcome to try it on your own and see if it improves the experience.
- mycall 6y agoHow does ArchiveBox function compared to https://archivarix.com https://archivarix.com? I recently used Archivarix to backup a large website (93k pages), but it messed up the js/css.
- mikiem 6y agoHow can I use this to archive sites/pages that require logging in to see?
- CodeWriter23 6y agoFrom the blog comments, I think this is what you’re after https://github.com/c9fe/22120 https://github.com/c9fe/22120
- nikisweeting 6y agohttps://github.com/ArchiveBox/ArchiveBox/wiki/Security-Overview#private-mode https://github.com/ArchiveBox/ArchiveBox/wiki/Security-Overv...
- unnouinceput 6y agoQuote: "..even if you instruct it to begin archiving a site then it can easily fail if that site’s robots.txt prevents crawling" Huh? Does actually the big corporations care anymore about robots.txt? Nowadays is more of a "netiquette" than anything else. Google definitely ignores it. Dunno DuckDucGo what it does
- zeckalpha 6y agoHow long until this is a feature baked into a mainstream web browser? Archive, prefetch, cache, all variants on a theme. History, bookmarks, local search engine, all the same.
- andai 6y agoI often wish that I could do a full text search of every page I've already visited.
- mail2merge 6y agoI'm working in that in my "self host the internet offline from your browsing history" project https://github.com/c9fe/22120 https://github.com/c9fe/22120 It makes a web archive from everything you browse, and lately I've been working on the full text search
- Moru 6y agoSeems it requires chrome to work?
- mail2merge 6y agoWow, your reading comprehension is amazingly good. Yep, that's correct.
- andai 6y agoThere's a link to more info on the Chrome thing but it 404s https://github.com/c9fe/22120/issues/57 https://github.com/c9fe/22120/issues/57
- nikisweeting 6y agohttp://web.archive.org/web/20201206153345/https://github.com/c9fe/22120/issues/57 http://web.archive.org/web/20201206153345/https://github.com...
- matt_f 6y agoInteresting side note: It seems like a lot of people in this thread have an interest in retaining a "replayable timeline" of their own browsing/reading history. There's probably enough support here to gather a few contributors for an open source project.
- hobo_mark 6y agoI seem to remember Google's Larry Page once proposed a similar thing in the early days, a product that would record all you read on your computer (to make it searchable later), but now I can't find it mentioned anywhere, am I imagining things?
- dgeiser13 6y agoIf you use Google Chrome as your primary browser this exists at https://myactivity.google.com/item https://myactivity.google.com/item
- nikisweeting 6y agoA "remember everything for me" tool is often called a "Memex" https://en.wikipedia.org/wiki/Memex https://en.wikipedia.org/wiki/Memex
- nikisweeting 6y agoThere are a bunch of projects trying to do different flavors of this already, check out some of these: https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community#other-archivebox-alternatives https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-...
- williesleg 6y agoJust do the fucking needful
- throwawaysea 6y agoCan you configure this tool to login to websites (for paid news subscriptions) and get past those paywalls?
- frombody 6y agoLikely not without some modification, but you could try this: https://www.jacoduplessis.co.za/bypass-paywall/ https://www.jacoduplessis.co.za/bypass-paywall/
- ernesth 6y agoThat is the default for the screenshot, pdf and one of the html archives: they use your chrome cookies.
- nikisweeting 6y agoYeah, it supports it but there are security considerations if you're doing it for anything more serious than news content. See here: https://github.com/ArchiveBox/ArchiveBox/wiki/Security-Overview#private-mode https://github.com/ArchiveBox/ArchiveBox/wiki/Security-Overv...
- throwawaysea 6y agoThanks, much appreciated. This is a very informative set of things to watch out for that I wouldn't have thought of otherwise.
- nikisweeting 6y agoMake sure you read through this section as well to fully understand the security concerns: https://github.com/ArchiveBox/ArchiveBox#caveats https://github.com/ArchiveBox/ArchiveBox#caveats
- dirtyid 6y agoTried this a while ago, disappointed at HD usage. My solution as heavy TTS user who has balabolka setup to read copied text which naturally leaves a log for future reference. There's extentions to auto copy highlighted text and append urls which makes entire flow straight forward. Log each day is around 1-5mbs of text saved in a big folder. Biggest limitation is trying to advance search unstructured text files by complex keywords within dates. I'm sure I can setup each clip with delimiters so logs can be imported into a searchable DB, just too lazy.
- nikisweeting 6y agoI think you tried a very old version ;) all that has long since changed. As of v0.5 ArchiveBox has everything in a Sqlite3 DB and full-text search is implemented with Sonic.
- nikisweeting 6y agoHey all, @pirate (ArchiveBox maintainer) here, thanks for posting this @adamhearn. If you like ArchiveBox check out our new Twitter account for the project, https://twitter.com/ArchiveBoxApp https://twitter.com/ArchiveBoxApp we just opened it and we'll be posting announcements and prerelease sneak-peeks on there in the future.
- egberts1 6y agoA real OSINT archive box would also capture all non-inline JavaScript, CSS and blob: files.