6 ms·
Why does HN run such a low resource site on a single bare metal server (plus identical backup) in the first place? 2x E5-2637 v4 is not that much CPU power thes
by Shakahs 4y ago
Why does HN run such a low resource site on a single bare metal server (plus identical backup) in the first place? 2x E5-2637 v4 is not that much CPU power these days.
An IaaS VM from the same vendor (M5 Hosting) would provide equivalent resources with much higher reliability. What benefit is 2 bare metal servers providing over a single properly sized VM in a managed resource pool with SAN backed storage?
I don't think it saves any money, and when those bare metal servers suffered identical hardware failures from a firmware bug HN was down for most of a day and temporarily migrated to an IaaS provider anyway (AWS).
- trasz 4y ago>much higher reliability [citation needed]. Sure, it solves some reliability problems, but introduces others, and I’m not sure if it’s a net win.
- voxadam 4y agoI wonder if HN's single threaded nature has anything to do with it. Eight cores or 16 threads don't do much good when your bespoke Arc Lisp stack is single core limited.
- JohnHaugeland 4y agoI kind of wonder if it's somehow not thread safe, or not talking to a database; or if not, why they don't just run parallel copies
- _Wintermute 4y agoI can't find an authoritative source, but I've read that HN uses text files stored in a filesystem rather than a database, so that might be why?
- Havoc 4y agoDon't fix it if it isn't broken
- orf 4y agoIt was and has been broken. A few times.
- candiodari 4y agoA 10x more complex redundant (or "redundant") system often breaks faster (and definitely stays down longer) than a simple direct system. Many people just don't consider failure scenarios. Offsite live database backups, for example, are a great idea. Say ... how does your site perform, in percent of normal QPS, when the database is now 150ms away instead of 1ns? 1% ... that's not redundancy, despite the site being up, let's just call that a failure. And people forget one thing about hosting on AWS. Say ... when AWS is slow/has problems/blocked on firewall/down ... when your competitor is down, would you like your site to be up? How about vice versa?
- orf 4y ago> Offsite database backups, for example, are a great idea “Off-site database backups” that mean you now have a 150ms round trip for user facing queries? … what on earth are you talking about. Is your argument that “simple” things are better because “you can do stupid things and blame it on complexity”?
- candiodari 4y agoThe database had a fallback that, according to good practice, was hosted in a database in a different city, another provider (different country actually, but this is Europe, it wasn't actually that far in km, it was, however >100ms away). Because they had a really fast local database essentially all the time, every pageview started requiring more and more database queries, some 50 for the front page alone, as the developers added features. Then the database needed to failover. And the complexity hadn't actually killed IT (yet): it actually worked ... But of course 50 * 2 * 150ms = 15000, or 15 seconds per page. I'm saying simple things can be better even when they don't provide redundancy because there's a bunch of problems that increase complexity so much that you can literally fix a simple problem faster than redundancy can take care of it.