3 ms·
> I'd love to see the "little delay" when reconstructing a file qualified somewhat. Are we talking < 5s or < 10s? I asked the engineers that work on that code,
by brianwski 6y ago
> I'd love to see the "little delay" when reconstructing a file qualified somewhat. Are we talking < 5s or < 10s?
I asked the engineers that work on that code, and they pulled a random sample from the logs (we time all of this) and said for files less than 1 MByte, it averaged around 250 milliseconds to reconstruct the file from the Backblaze Vault and get it onto the cache servers where it is then served up. In 95% of requests completed within 900 milliseconds, but there were a few up over 1 second (1.2 seconds was the highest they found). Those are live production numbers so it includes all the load on those Vaults.
A couple other notes just to add color. Any one Backblaze account is bound for life to what we call a "cluster", for example there is one cluster in Europe so all files are stored in Europe for any account in Europe. There is a load balanced array of "cache servers" in front of all the vaults specific to that cluster (the caching servers are physically located close to the vaults for latency reasons), and our biggest cluster has something like 20 of these SSD based caching servers. Ok, so the cache layer is not "shared", meaning each cache server only pulls directly from the Backblaze Vault. So if you were serving a file, and 20 separate customers got amazingly unlucky, the file would get the 250 millisecond lag every time for those first 20 fetches. The cool parts of this architecture is that then you have 20 populated caches that are completely unrelated to each other so you have 20x the bandwidth available to serve it up (and a rack of really fast 20 servers to serve it). Plus they are all totally independent so they can crash or be brought offline to upgrade the software without any downtime.
We can add these cache machines as we need them, they are these 1U units and we have "warm spares" for a variety of things. When we have had spikes in load in the past we toss some hardware at it pretty fast.