7 ms·
for #1, it could be fast, but the problem they're running up against is that the data appears to be held up by a single bottleneck, similar to pre-sharded datab
by nsfmc 12y ago
for #1, it could be fast, but the problem they're running up against is that the data appears to be held up by a single bottleneck, similar to pre-sharded database setups. in this case, their filesystemdb is hitting the limits of too many connections saturating much of their disk's i/o.
a possible way of approaching that problem is a divide and conquer approach with a reverse proxy that assigned manageable chunks of their content across numerous machines each serving less than millions of pages. the nyt already has a /<yyyy>/<mm>/<dd>/<section>/<subsection>/<slug> url scheme which would make this less painful.
I'm not sure how inefficient this would be, though, certainly a time investment, but it ends up offloading your disk i/o/ issues by creating more and more s3 buckets (or what have you) and routing via a proxy. i'd be curious to see when s3+cloudfront-as-host becomes too slow simply because of disk i/o limitations, although s3 almost certainly has its own abstraction above the bucket i'm not aware of which mitigates that.
it still doesn't address the serious and complex frontend issues they were facing, which seemed much more onerous to be honest. their server rendering seems to be pretty lean though, it looks like dom processing and client rendering make up easily 80% of their 3.6s pageload time.