6 ms·
HTTrack Website Copier
- jregmail 2y agoI recommend to try also https://crawler.siteone.io/ https://crawler.siteone.io/ for web copying/cloning. Real copy of the netlify.com website for demonstration: https://crawler.siteone.io/examples-exports/netlify.com/ https://crawler.siteone.io/examples-exports/netlify.com/ Sample analysis of the netlify.com website, which this tool can also provide: https://crawler.siteone.io/html/2024-08-23/forever/x2-vuvb0oi6qxkr-ku79.html https://crawler.siteone.io/html/2024-08-23/forever/x2-vuvb0o...
- alberth 2y agoI always wonder if this gives false positives for people just using the same WordPress template.
- zazaulola 2y agoThe archive saved in HTTrack Website Copier can be opened in https://replayweb.page https://replayweb.page locally or they have different save formats?
- suriya-ganesh 2y agoThis saved me a ton when back in college in rural India without Internet in 2015. I would download whole websites from a nearby library and read at home. I've read py4e, ostep, Pgs essays using this. I am who I am because of httrack. Thank you
- superjan 2y agoI have tried the windows version 2 years ago. The site I copied was our on-prem issue tracker (fogbugz) that we replaced. HTTrack did not work because of too much javascript rendering, and I could not figure out how to make it login. What I ended up doing was embedding a browser (WebView2) in a C# Desktop app. You can intercept all the images/css, and after the Javascript rendering was complete, write out the DOM content to a html file. Also nice is that you can login by hand if needed, and you can generate all urls from code.
- chirau 2y agoI use it to download sites with layouts that I like and want to use for landing pages and static pages for random projects. I strip all the copy and stuff and leave the skeleton to put my own content. Most recently link.com, column.com and increase.com. I don't have the time nor the youth to start with all the JavaScript & React stuff.
- j0hnyl 2y agoScammers love this tool. I see it used in the wild quite a bit.
- deleted 2y ago[deleted]
- xnx 2y agoGreat tool. Does it still work for the "modern" web (i.e. now that even simple/content websites have become "apps")?
- alganet 2y agoNope. It is for the classic web (the only websites worth saving anyway).
- freedomben 2y agoEven for classic web, if it's behind cloudflare, then HTTrack no longer works. It's a sad point to be at. Fortunately, the single file extension still works really well for single pages, even when they are built dynamically by JavaScript on the client side. There isn't a solution for cloning an entire site though, at least that I know of
- knowaveragejoe 2y agoI'm aware of this tool, but I'm sure there are caveats in terms of "totally" cloning a website: https://github.com/ArchiveTeam/grab-site https://github.com/ArchiveTeam/grab-site
- alganet 2y agoIf it is cloudflare human verification, then httrack will have an issue. But in the end it's just a cookie, you can use a browser with JS to grab the cookie, then feed it to httrack headers. If cloudflare ddos protection is an issue, you can throttle httrack requests.
- _lvbh 2y ago> you can use a browser with JS to grab the cookie, then feed it to httrack headers They also check your user agent, IP and JA3 fingerprint (and ensures it matches with the one that got the cookie) so it's not as simple as copying some cookies. This might just be for paying customers though since it doesn't do such heavy checks for some sites
- dark-star 2y agooh wow that brings back memories. I have used httrack in the late 90s and early 2000's to mirror interesting websites from the early internet, over a modem connection (and early DSL) Good to know they're still around, however, now that the web is much more dynamic I guess it's not as useful anymore as it was back then
- dspillett 2y ago> now that the web is much more dynamic I guess it's not as useful anymore as it was back then Also less useful because the web is so easy to access, I remember using it back then to draw things down over the university link for reference in my room (1st year, no network access at all in rooms) or house (or per-minute costed modem access). Sites can vanish easily of course still these days, so having a local copy could be a bonus, but they just as likely go out of date or get replaced, and if not are usually archived elsewhere already.
- Alifatisk 2y agoGood ol' days
- corinroyal 2y agoOne time I was trying to create an offline backup of a botanical medicine site for my studies. Somehow I turned off depth of link checking and made it follow offsite links. I forgot about it. A few days later the machine crashed due to a full disk from trying to cram as much of the WWW as it could on there.
- rkhassen9 2y agoThat is awesome.
- Felk 2y agoFunny seeing this here now, as I _just_ finished archiving an old MyBB PHP forum. Though I used `wget` and it took 2 weeks and 260GB of uncompressed disk space (12GB compressed with zstd), and the process was not interruptible and I had to start over each time my hard drive got full. Maybe I should have given HTTrack a shot to see how it compares. If anyone wanna know the specifics on how I used wget, I wrote it down here: https://github.com/SpeedcubeDE/speedcube.de-forum-archive https://github.com/SpeedcubeDE/speedcube.de-forum-archive Also, if anyone has experience archiving similar websites with HTTrack and maybe know how it compares to wget for my use case, I'd love to hear about it!
- begrid 2y agowget2 has an option por paralel downloading. https://github.com/rockdaboot/wget2 https://github.com/rockdaboot/wget2
- criddell 2y agoIs there a friendly way to do this? I'd feel bad burning through hundreds of gigabytes of bandwidth for a non-corporate site. Would a database snapshot be as useful?
- dbtablesorrows 2y agoIf you want to customize the scraping, there's scrapy python framework. You would always need to download the html though.
- squigz 2y agoIsn't bandwidth mostly dirt cheap/free these days?
- oriettaxx 2y agoI don't get it: last release 2017 while in github I see more releases... so, did developer of the github repo took over and updating/upgrading? very good!
- subzero06 2y agoi use this to double check which of my web app folder/files are publicly accessible.
- woutervddn 2y agoAlso known as: static site generator for any original website platform...