3 ms·
Disclaimer: I work at Backblaze. > I really want to see lightning fast response times and TTFB (Time To First Byte Served) If a file is "cold" (nobody has req
by brianwski 6y ago
Disclaimer: I work at Backblaze.
> I really want to see lightning fast response times and TTFB (Time To First Byte Served)
If a file is "cold" (nobody has requested it in the last 24 hours) then it needs to be reconstructed from the Backblaze Vaults and there is a little delay. After that, it should serve pretty fast for the following requests (off of a caching layer with SSDs).
In the end, Backblaze B2 is a good solution for some customers, and not ideal for others. If your application requires blinding speed, like sub 1 millisecond serve times, Backblaze B2 may not be perfect for you. But how often is that the case? Certainly not when fetching a web page, or storing a backup for a year, right? In those cases a small delay is FINE. This is an example web page served by Backblaze B2 here, how does it load for you? https://f001.backblazeb2.com/file/ski-epic-c/full/2015_scotland_will_macdonald_birthday_in_duns_castle/index.html https://f001.backblazeb2.com/file/ski-epic-c/full/2015_scotl... Fast? Slow? How is it?
For comparison, my regular hosting provider serving the same web page here: https://www.ski-epic.com/2015_scotland_will_macdonald_birthday_in_duns_castle/index.html https://www.ski-epic.com/2015_scotland_will_macdonald_birthd...
Personally I can't tell any difference. I still look silly in a kilt in both versions. :-)
> Second pain point is the number of retries needed for uploading a large batch of small files.
It really shouldn't take any retries, or geez, at VERY MOST something like less than 1% - why is that an issue? Software should handle the tiny failure rate. I'm honestly curious, we want to know why people aren't choosing our solution!!
- heipei 6y agoI understand you're probably not in a position to say anything about it, but I'd love to see the "little delay" when reconstructing a file qualified somewhat. Are we talking < 5s or < 10s? What do the percentiles for restore latency look like? How does file size play into it? This, for me, is one of the biggest unknowns right now since it's not easy to create a test benchmark for this case (i.e. upload a bunch of stuff and let it sit idle for at least 24 hours, hoping it will be expired from the caching layer).
- brianwski 6y ago> I'd love to see the "little delay" when reconstructing a file qualified somewhat. Are we talking < 5s or < 10s? I asked the engineers that work on that code, and they pulled a random sample from the logs (we time all of this) and said for files less than 1 MByte, it averaged around 250 milliseconds to reconstruct the file from the Backblaze Vault and get it onto the cache servers where it is then served up. In 95% of requests completed within 900 milliseconds, but there were a few up over 1 second (1.2 seconds was the highest they found). Those are live production numbers so it includes all the load on those Vaults. A couple other notes just to add color. Any one Backblaze account is bound for life to what we call a "cluster", for example there is one cluster in Europe so all files are stored in Europe for any account in Europe. There is a load balanced array of "cache servers" in front of all the vaults specific to that cluster (the caching servers are physically located close to the vaults for latency reasons), and our biggest cluster has something like 20 of these SSD based caching servers. Ok, so the cache layer is not "shared", meaning each cache server only pulls directly from the Backblaze Vault. So if you were serving a file, and 20 separate customers got amazingly unlucky, the file would get the 250 millisecond lag every time for those first 20 fetches. The cool parts of this architecture is that then you have 20 populated caches that are completely unrelated to each other so you have 20x the bandwidth available to serve it up (and a rack of really fast 20 servers to serve it). Plus they are all totally independent so they can crash or be brought offline to upgrade the software without any downtime. We can add these cache machines as we need them, they are these 1U units and we have "warm spares" for a variety of things. When we have had spikes in load in the past we toss some hardware at it pretty fast.
- christophilus 6y agoWhen I tested B2 a month or so ago, I was seeing pretty frequent failure rates. Granted, I was on a free tier, and it was probably around the time that this whole COVID thing started spiking traffic everywhere...