15 ms·
10% of the top million sites are dead
- yajjackson 4y agoTangential, but I love the format for your site. Any plans to do a "How I built this blog" post?
- kozziollek 4y agoMost of cities in Poland have their own $city.pl domain and allow websites to buy $website.$city.pl. That might not be well known. And cities have theri websites, so I guess it's OK. But info.pl and biz.pl? Did nobody hear about country variants of gTLDs?!
- drdaeman 4y agoThose are called Public Suffixes or effective TLDs (eTLDs): https://en.wikipedia.org/wiki/Public_Suffix_List https://en.wikipedia.org/wiki/Public_Suffix_List And you're entirely correct that author should've referred to such list.
- slyall 4y agoI think the problem is that the original source needs to use that list as well. Just looked though .nz and they list several sites ( govt.nz , school.nz, gen.nz ) that don't exist since all the domains are one level below. They even list.govt.nz as the top site. In fact that doesn't exist (although www.govt.nz does since it is a a kinda government portal ) I see they list an old employer of mine who got bought 15 years ago and whose website has been redirecting for 10 years.
- crikeyjoe 4y ago
- macintux 4y agoTitle is misleading: that’s the outcome, but the bulk of the story is the data processing to reach that conclusion.
- hinkley 4y agoIt happens. Most of the stuff we do these days invokes a number of disciplines. I forget sometimes that maybe ten percent of us just play with random CS domains for “fun” and that most people are coming into big problems blind, even sometimes the explorers (though having comfort with exploring random fields is a skill set unto itself). Before the Cloud, when people would ask for a book on distributed computing, which wasn’t that often, I would tell them seriously “Practical Parallel Rendering”. That book was almost ten years old by then. 20 now. It’s ostensibly a book about CGI, but CGI is about distributed work pools, so half the book is a whirlwind tour of distributed computing and queuing theory. Once they start talking at length about raytracing, you can stop reading if CGI isn’t your thing, but that’s more than halfway through the book. I still have to explain some of that stuff to people, and it catches them off guard because they think surely this little task is not so sophisticated as that… I think this is where the art comes in. You can make something fiddly that takes constant supervision, so much so that you get frustrated trying to explain it to others, or you can make something where you push a button and magic comes out.
- 5350-uiop-1130 4y agowhenever i go through my bookmarks, i tend to find maybe 5-10% are now 404. this is why i like the archive.ph project so much and using it more as a kind of bookmarking service.
- syedkarim 4y agoWhat’s the benefit to using archive.ph instead of archive.org (Internet Archive)? Seems like the latter is much more likely to be around for awhile.
- 5350-uiop-1130 4y agoi find archive.ph does a better job of preserving the page as is (it also takes a screenshot) compared to internet archive which can be flaky at best. i also find archive.ph much faster at searching, and the browser extension is really useful too. the faq does a great job of explaining too https://archive.ph/faq https://archive.ph/faq
- yellow_lead 4y agoIsn't archive.ph/today the one with questionable funding sources and backing? Who is behind it and can it be trusted for longevity?
- 5350-uiop-1130 4y agoyeah funding is a grey area... fwiw the website is only accessible by VPN in a lot of countries, which is say a lot for me..and i don't think they've taken down any content, although i cant say for sure.
- NavinF 4y agoIn this case the less we know, the longer it will last. Notice how this site ignores robots.txt and copyright claims by litigious companies that would like to see their past erased. The data saved on your NAS will outlast this site regardless of who owns/funds it.
- bioemerl 4y agoI'm honestly amazed that out of the top million sites, which probably includes a ton of tiny tiny sites that are idle or abandoned, only ten percent are offline.
- MonkeyMalarky 4y agoHow many are placeholder pages thrown up by registrars like Network Solutions?
- denton-scratch 4y agoIf they're placeholder pages, they're not dead. Those 10% are not responding at all; the requests aren't reaching any HTTP server.
- winddude 4y agoat least from his computer/script. A number could have been blocked simply detecting him as a bot.
- zamadatix 4y agoNot all placeholder pages will forever stay placeholder pages though. Some may get sold, become a site, then stop being a site again. Some may not get sold, come up for renewal and be deemed unlikely to be worth trying to sell anymore (renewal is cheap for a registrar but the registry will still charge a small fee). Of course the vast majority with enough interest to make this list will either be sold and be an active page or still be an active placeholder but I wouldn't rule out there being a good count of pages towards the lower end of the top million being placeholders that were eventually deemed not worth trying for anymore.
- MonkeyMalarky 4y agoExactly. Afaik, there's even a whole auction based side industry trading expired domains. When you see a placeholder page, the original site is for all intents and purposes, dead. The domain just happens to be interesting enough that someone wants to stick an ad on it rather than let it resolve to nothing.
- MonkeyMalarky 4y agoLast time I tried to crawl that many domains, I ran into problems with my ISP's DNS server. I ended up using a pool of public DNS servers to spread out all the requests. I'm surprised that wasn't an issue for the author?
- wumpus 4y agoYou have to run your own resolver. Crawling 101.
- MonkeyMalarky 4y agoThis is of course the correct answer. It just felt like shaving a big yak at the time.
- mh- 4y agoA properly configured unbound running locally can be a decent compromise.
- denton-scratch 4y agoThat is running your own resolver. Unbound is a resolver.
- mh- 4y agowell, yes, but I guess I think of unbound in a different category from setting up (e.g.) bind. but, my experience configuring bind is probably more than 20 years out of date. you're right to make that correction though, so thank you. :)
- fullstop 4y agoBIND is odd in that it combines a recursive resolver with an authoritative name server, and this has actually led to a number of security vulnerabilities over the years. Other alternatives, such as djb's dnscache/tinydns and NLNet Labs' Unbound/nsd separate the two to avoid this entirety.
- superb-owl 4y agoOne of the few things I like about blockchain is the promise of a less ephemeral web.
- bergenty 4y agoIs that actually true? Don’t most nodes hold heavily compressed pointers only while there are only a percentage of nodes that host the entire blockchain. I mean if what you’re saying is true then each node needs to have a copy of the entire internet which isn’t reasonable.
- superb-owl 4y agoI'm thinking about things like Filecoin, a blockchain which is meant to power IPFS. To be fair though, IPFS itself is not a blockchain
- deltree7 4y agospoken like someone who is clueless about Blockchain
- matkoniecz 4y agoOne of many things I dislike about cryptoscams is making promises which are lies.
- zzzeek 4y agoirony that the site is not responding?
- ocdtrekkie 4y agoI've been working on trying to migrate sites I ran in 2008 or so into my new preferred hosting strategy lately: I know zero people look at them, since many were functionally broken at present, but I don't like the idea of actually removing them from the web. So I'm patching them up, migrating them to a more maintainable setting, and keeping them going. Maybe someday some historian will get something out of it.
- tete 4y agoThe biggest problem I find is that it seems to be pretty "outdated" to keep redirects in place, if you move stuff. So many links to news websites, etc. will cause a redirect to either / or a 404 (which is a very odd thing to redirect to in my opinion). If you are unlucky an article you wanted to find also completely disappeared. This is scary, because it's basically history disappearing. I also wonder what will happen to text on websites that are some ajax and javascript breaks because a third party goes down. While the internet archive seems to be building tools for people to use to mitigate this I found that they barely worked on websites that do something like this. Another worry is the ever-increasing size of these scripts making archiving more expensive.
- Kye 4y agoYou can often pop the URL into the Wayback Machine to bring up the last live copy. It's better at handling dynamic stuff the more recent it is. Older stuff, especially early AJAX pages, are just gone because the crawler couldn't handle it at the time. It's far from a perfect solution, especially in light of the big publishers finally getting their excuse to go after the Internet Archive legally. It's a good silo, but just as vulnerable as any other.
- nikisweeting 4y agoArchiveWeb.page + ReplayWeb.page are the best I've found at handling ajax loaded content.
- phkahler 4y agoRead that again folks: "a very reasonable but basic check would be to check each domain and verify that it was online and responsive to http requests. With only a million domains, this could be run from my own computer relatively simply and it would give us a very quick temperature check on whether the list truly was representative of the “top sites on the internet”. " This took him 50 minutes to run. Think about that when you want to host something smaller than a large commercial site. We live in the future now, where bandwidth is relatively high and computers are fast. Point being that you don't need to rent or provision "big infrastructure" unless you're actually quite big.
- deleted 4y ago[deleted]
- jayd16 4y agoThe flip side is anyone can run these kinds of tools against your site easily and cheaply.
- stevemk14ebr 4y agoyour point has a truth behind it for sure, but there's a large difference between serving requests and making requests. Many sites are simple html and css pages, but many others also have complex backends. It's those that often are hard to scale and why the cloud is hugely popular, maintaining and scaling the backend is hard
- phkahler 4y agoOh absolutely, but he also said this: I found that my local system could easily handle 512 parallel processes, with my CPU @ ~35% utilization, 2GB of RAM usage, and a constant 1.5MB down on the network. Another thing that happened in the early web days was Apache. People needed a web server and it did the job correctly. Nobody ever really noticed that it had terrible performance, so early on infrastructure went to multiple servers and load balancers and all that jazz. Now with nginx, fast multi-core, and speedy networks even at home, it's possible to run sites with a hundred thousand users a day at home on a laptop. Not that you'd really want to do exactly that but it could be done. Because of this I think an alternative to github would be open source projects hosted on peoples home machines. CI/CD might require distributing work to those with the right hardware variants though.
- zinekeller 4y agoTLDR: Campbell's methodology is flawed, does not consider edge cases (one of which (equating apex-only and www-prefixed domains) I consider reckless), and didn't understand how Majestic collects and processes its data. Longer version: This isn't comprehensive, but I think of two main reasons why: - The Majestic Million lists only the registrable part (with some exceptions), and this sometimes lead to central CDNs being listed. For example, the Majestic Million lists wixsite.com (for those who are unaware is a CDN domain used by Wix.com with separate subdomains), but if you visit wixsite.com you wouldn't get anything. Same with Azure, subdomains of azureedge.net and azurewebsites.net do exist (for example https://peering.azurewebsites.net/ https://peering.azurewebsites.net/) but azureedge.net and azurewebsites.net themselves don't exist. Without similar filtering, using the Cisco list (https://s3-us-west-1.amazonaws.com/umbrella-static/index.html https://s3-us-west-1.amazonaws.com/umbrella-static/index.htm...) would quickly lead you to this precise problem (mainly because the number one is "com", but phew at least http://ai./ http://ai./ does exist!) - Also, shame on the author considering www-prefixed and apex-only as one and the same. For some websites, it isn't. Take this example: jma.go.jp (Japan Meteorological Agency), which doesn't respond (actually NODATA) on http://jma.go.jp/ http://jma.go.jp/ but is fine on https://www.jma.go.jp/ https://www.jma.go.jp/. Similarly, beian.gov.cn (Chinese ICP Licence Administrator) wouldn't respond at all but www.beian.gov.cn will. And for ncbi.nlm.nih.gov (National Center for Biotechnology Information) ? I can't blame Majestic: https://www.ncbi.nlm.nih.gov/ https://www.ncbi.nlm.nih.gov/ and https://ncbi.nlm.nih.gov/ https://ncbi.nlm.nih.gov/ don't redirect to a canonical domain, and unless you've compared the HTTP pages there's no way you would know that they are the same website! Edit: I've downloaded out the CSV to check my claims, and it shows: wixsite.com 0 beian.gov.cn 0 Please, for the love of sanity, consider what the Majestic Million (and similar lists) criterion on inclusion. I can't believe it to say, but can we crowd-source "Falsehoods programmers believe about domains"? Also addendum to crawling but I consider "probably forgivable": - Some websites are only available in certain countries (internal Russian websites don't respond at all outside Russia for example). This can skew the numbers a little bit.
- deleted 4y ago[deleted]
- zepearl 4y ago
- the_biot 4y agoBy what possible criteria are these the "top" million sites, if 10% are dead? I'd start with questioning that data.
- deltree7 4y agoExactly! Garbage In == Garbage Out
- kjeetgill 4y agoDude, it's the second sentence of the first paragraph: > For my purposes, the Majestic Million dataset felt like the perfect fit as it is ranked by the number of links that point to that domain (as well as taking into account diversity of the origin domains as well).
- MatthiasPortzel 4y agoAnd moreover, the author’s conclusion is that the dataset is bad. > While I had expected some cleanliness issues, I wasn’t expecting to see this level of quality problems from a dataset that I’ve seen referenced pretty extensively across the web
- the_biot 4y agoYeah, but they're still providing a dataset that's just plain bad. It's hardly relevant how many sites link to some other site, if it's dead.
- Brian_K_White 4y agoIt's only bad data if it does not include what it claims to include. If the dataset is defined as inlinks, and it is inlinks, then the data is good.
- winddude 4y agopart of the problem is it's not the number of links, it's referring subnets. Fairly certain this includes, script tags.
- gravitate 4y ago> Domain normalization is a bitch I’m a no-www advocate. All my sites can be accessed from the Apex domain. But some people for whatever reason like to prepend www to my domains, so I wrote a rule in Apache’s .HTACCESS to rewrite the www to the Apex. Here’s a tutorial for doing that: https://techstream.org/Web-Development/HTACCESS/WWW-to-Non-WWW-and-Non-WWW-to-WWW-redirect-with-HTACCESS https://techstream.org/Web-Development/HTACCESS/WWW-to-Non-W...
- macintux 4y ago25 years ago I added a rule to my employer’s firewall to allow the bare domain to work on our web server. Inbound email immediately broke. I was still very new, and didn’t want to prolong the downtime, so I reverted instead of troubleshooting. A few months after I left, I sent an email to a former co-worker, my replacement, and got the same bounce message. I rang him up and verified that he had just set up the same firewall rule. Been much too long to have any clue now what we did wrong.
- JackMcMack 4y agoYou probably created a cname from the apex to www? This problem still exists today. From https://en.wikipedia.org/wiki/CNAME_record https://en.wikipedia.org/wiki/CNAME_record: "If a CNAME record is present at a node, no other data should be present; this ensures that the data for a canonical name and its aliases cannot be different." So if you're looking up the MX record for domain, but happen to find a cname for domain to www.domain , it will follow that and won't find any MX records for www.domain. The correct approach is to create a cname record from www.domain to domain, and have the A record (and MX and other records) on the apex. Most DNS providers have a proprietary workaround to create dns-redirects on the apex (such as AWS Route53 Alias records) and serve them as A records, but those rarely play nice with external resources.
- tux2bsd 4y ago> You probably created a cname from the apex to You can't do that, period. A lot of "cloud" and other GUI interfaces deceive people into thinking it's possible, they just do A record fuckery behind the scenes (clever in it's own right but it causes misunderstanding).
- altdataseller 4y agoAll these top million lists are very good at telling you the top most 10K-50K sites on the web. After that, you're going into 'crapshoot' land, where the 500,000th most popular site is very likely to be a site that got some traffic a long time ago, but now isn't even up. So I would take this data with a grain of salt. You're better off just analyzing the top 100K sites on these lists.
- TuringNYC 4y agoHow are people determining the "top" sites? We do some of this at work and we pay SimilarWeb a giant sum of money, are people able to find site traffic in inexpensive ways which allow for these analyses?
- giantrobot 4y ago> where the 500,000th most popular site is very likely to be a site that got some traffic a long time ago, but now isn't even up. That's literally the phenomenon the article is describing.
- altdataseller 4y agoOk let me reword it differently: the 500,000th most popular site on these lists most likely isnt the 500,000th most visited and it might not even be in the top 5 million. These data sources are so bad at capturing popularity after 50k sites or so simply because they dont have enough data
- iruoy 4y agoI haven't tested this, but the "Cisco Umbrella 1 Million" is generated daily from DNS request made to the Cisco Umbrella DNS service. That seems to be a very good and recent dataset. It does count more than just visiting websites though. If all Windows computers query the IP of microsoft.com once a day that'll move them up quite a bit. And things in their top 10 like googleapis.com and doubleclick.net are obviously not visited directly. So while it is quite a reliable and recent dataset, it is not a good test of popularity.
- smugma 4y agoI downloaded the file and looked at the second 000 in his file, which refers to wixsite.com. It appears that wixsite.com isn't valid but www.wixsite.com is, and redirects to wix.com. It's misleading to say that the sites are dead. As noted elsewhere, his source data is crap (other sites I checked such as wixstatic.com don't appear to be valid) but his methodology is bad, or at least his describing the sites as dead is misleading.
- code123456789 4y agowixsite.com is a domain for free sites built on Wix, so if your username on Wix is smugma, and your site name is mysite, then you'll have a URL like smugma.wixsite.com/mysite for your Home page. That's why this domain is in the top
- smugma 4y agoCorrect, that's why it's in the top. Your example further confirms why the author's methodology is broken.
- zinekeller 4y ago> other sites I checked such as wixstatic.com don't appear to be valid But docs.wixstatic.com is valid.
- quickthrower2 4y agoHe takes this into account by generously considering any returned response code as “not dead”. > there’s a longtail of sites that had a variety of non-200 reponse codes but just to be conservative we’ll assume that they are all valid
- mort96 4y agoThat doesn't take this into account, no. `curl wixsite.com` returns a "Could not resolve host" error; it doesn't return a response code, so the author would consider it invalid, even though `curl www.wixsite.com` does return a response (a 301 redirect to www.wix.com).
- allknowingfrog 4y agoI don't have any particular opinions on the author's conclusions, but I learned a thing or two about the power of terminal commands by reading through the article. I had no idea that xargs had a parallel mode.
- thelamest 4y agoProbably not news to anyone who works with big data™, but I learned, after additional searches, that using (something like) duckdb as a CSV parser makes sense, especially if the alternative is loading the entire thing to memory with (something like) base R. This was informative for me: https://hbs-rcs.github.io/large_data_in_R/ https://hbs-rcs.github.io/large_data_in_R/.
- pahool 4y agozombo.com still kicking!
- softwaredoug 4y agoMy current beliefs about how people use and trust information on the Web. First, trust is _everything_ on the Web, it is the thing people first think of when arriving on some information. But how people evaluate trust has changed dramatically over the last 10 years. - Trust now comes almost exclusively from social proof. Searching reddit, youtube, etc and other extremely _moderated_ sources of information, where the most work is done to ensure content comes from actual human beings. How many of us now google `<topic> reddit` instead of just `<topic>`? - Of course a lot of this trust is misplaced. There's a very thin line between influencers and cult leaders / snake oil salesmen. Our last President used this hack really effectively. - Few trust Google's definition of trust anymore -- essentially page rank. This made more sense when the Web essentially was social, where inbound links were very organic. Now with the trust in general Web sites evaporated, the main 'inbound links' anyone cares about come from individuals or community they trust or identify with. They don't trust Googles algorithm (its too opaque, and too easily gamed). This of course means the fracturing of truth away from elites. Sometimes this could be a good thing, but in many cases cough Covid cough it might be pretty disastrous for misinformation
- mountainriver 4y ago> How many of us now google `<topic> reddit` instead of just `<topic>` I sure hope not, Reddit is horrible place for information
- romanhn 4y agoWhen I have a specific technical question, I append "stackoverflow" to my search queries. When I want to read a discussion, I add "reddit" (or "hacker news").
- failTide 4y agoI use the strategy for a few things - including when I want to get reviews of a product or service. There's still potential for manipulation there, but you can judge the replies based on the user history - and you know that businesses aren't able to delete or hide bad reviews there. But in general I agree with you - reddit is full of misinformation, propaganda and astroturfing
- gumby 4y agoHis 'www' logic is flawed: https://www.example.com https://www.example.com and https://example.com https://example.com need not return the same results, but his checking code sends the output straight to /dev/null so he has no way of knowing.
- gojomo 4y agoMany issues with this analysis, some others have already mentioned, including: • The 'domains' collected by the source, as those "with the most referring subnets", aren't necessarily 'websites' that now, or ever, respnded to HTTP • In many cases any responding web server will be on the `www.` subdomain, rather than the domain that was listed/probed – & not everyone sets up `www.` to respond/redirect. (Author misinterprets appearances of `www.domain` and `domain` in his source list as errant duplicates, when in fact that may be an indicator that those `www.domain` entries also have significant `subdomain.www.domain` extensions – depending on what Majestic means by 'subnets'.) • Many sites may block `curl` requests because they only want attended human browser traffic, and such blocking (while usually accompanied with some error response) can be a more aggressive drop-connection. • `curl` given a naked hostname likely attempts a plain HTTP connection, and given that even browsers now auto-prefix `https:` for a naked hostname, some active sites likely have nothing listening on plain-HTTP port anymore. • Author's burst of activity could've triggered other rate-limits/failures - either at shared hosts/inbound proxies servicing many of the target domains, or at local ISP egresses or DNS services. He'd need to drill-down into individual failures to get a beter idea to what extent this might be happening. If you want to probe if domains are still active: • confirm they're still registered via a `whois`-like lookup • examine their DNS records for evidence of current services • ping them, or any DNS-evident subdomains • if there are any MX records, check if the related SMTP server will confirm any likely email addresses (like postmaster@) as deliverable. (But: don't send an actual email message.) • (more at risk of being perceived as aggressive) scan any extant domains (from DNS) for open ports running any popular (not just HTTP) services If you want to probe if web sites are still active, start with an actual list of web site URLs that were known to have been active at some point.
- spc476 4y agoIt dawned on me when I hit the Majestic query page [1] and saw the link to "Commission a bespoke Majestic Analytics report." They run a bot that scans the web, and (my opinion, no real evidence) they probably don't include sites that block the MJ12bot. This could explain why my site isn't in the list, I had some issues with their bot [2] and they blocked themselves from crawling my site. So, is this a list of the actual top 1,000,000 sites? Or just the top 1,000,000 sites they crawl? [1] https://majestic.com/reports/majestic-million https://majestic.com/reports/majestic-million [2] http://boston.conman.org/2019/07/09-12 http://boston.conman.org/2019/07/09-12
- spaceman_2020 4y agoNot surprising. We're far away from the glory days of the vibrant, chaotic web. In countries like India that onboarded most users through smartphones instead of computers, websites are not even necessary. There's a huge dearth of local-focused web content as well since there just isn't enough demand.
- deleted 4y ago[deleted]
- winddude 4y agoNo they're not.
- nr2x 4y agoMajestic is a shit list. Mystery solved.
- ghostly_s 4y agoWow, I would not have suspected `tee` is able to handle multiple processes writing to the same file. Doesn't seem to be mentioned on the man-page, either.
- remram 4y agoAll tee does is write its standard input (a single file descriptor) to a file (a single one) and its own output. xargs is the thing running multiple processes (and they inherit the same standard output, your shell's). What you're seeing is Linux being able to handle multiple processes writing to the same file.
- ghostly_s 4y agoWell, then that's a Linux feature I was unaware of. I found this SO[1] question with two conflicting answers that have almost the same number of votes, and even the "yes you can do this" answer seems to have enough caveats that it doesn't sound like a great idea. 1. https://stackoverflow.com/questions/7842511/safe-to-have-multiple-processes-writing-to-the-same-file-at-the-same-time-cent https://stackoverflow.com/questions/7842511/safe-to-have-mul...
- zX41ZdbW 4y agoThis looks surprisingly similar to the unfinished research that I did: https://github.com/ClickHouse/ClickHouse/issues/18842 https://github.com/ClickHouse/ClickHouse/issues/18842
- noiv 4y agoHow does a dead site make it into the top million?
- dredmorbius 4y agoTypically, during its pre-death phase.
- kderbyma 4y agowouldn't this imply that either the ranking system is broken.....or there are less than 1 million active sites.....
- banana_giraffe 4y agoThe takeaway from this is slightly off. There aren't 107776 sites that are dead, there are 107776 sites that don't run a HTTP server, or are otherwise dead. If you try to connect via HTTP or HTTPS, then a quick run yields 91106 sites that are dead, or 9.11% (And I ran this test on an AWS EC2 node with a fairly aggressive timeout. No doubt some % of sites play dead to AWS, or didn't respond fast enough for me)
- flas9sd 4y agohaving the luxury of scrutinizing the method and retesting: to "normalize" domains and skip the www skewed results - not all websites do their redirects across apex to www (and schemas). Some servers weren't answering the request with the default curl accept header / and needed encouragement. I retested the 000 class of .de ccTLD (1227) and found more than a third (473) of them answering when prefixed with www. Lots of german universities were false negatives - if this is representative I cannot tell, just a hint to retest.
- indigodaddy 4y agoAre there more cycles/cpu/work involved to `cat verylargefile | awk` vs `awk verylargefile` ?
- terrycody 4y agoNice work. Just one thing, analyze sites by total referring domains is not accurate as your result showed. A backlink can be easily faked and you can literally spam 1 million links within 1 day for any domain. Thus, this data source is not much useful. For a more accurate result, try to use Ahrefs top 1 million domains, ranked by their traffics. Ahrefs rank sites by their ranking keywords, thus infer the traffic numbers, meaning, these websites are live, and ranking with some keywords. You will see the result is much more accurate then, maybe not even a single website will be offline, because they are earning good cash.
- baby 4y agoFree.fr, one of the biggest ISP in France a while back, and perhaps still today, still runs all the old-school websites it was hosting for people (for free) today. It's quite insane, but a lot of the French web 1.0 is still alive today thanks to them. Truly an ISP ran by passionate technical people.
- ssl232 4y agoGood on them. Last year I randomly discovered an ancient email to my old Hotmail address from free website host Tripod, owned at the time by Lycos, that old search engine. As an 11 year old I had a website with them and wanted to dig it out to see what I had put there. I managed to convince them I was the owner and got my access back, only to discover nothing there. I guess at some point in the ~20 years since I made one they nuked their dormant sites.