11 ms·
Bandwidth needs halved by new compression written in Go
- calinet6 14y agoGo is used, sure, but the cool part about this is the binary Railgun protocol. Really smart. Send only file hashes and binary diffs back and forth, do a little extra computation to figure out the changes, but only send the absolute minimum data you need to the CDN. That's just smart, and frankly, I hope other CDNs have been doing this already, because at any high volume it seems to be an obvious solution. So that brings up the question—is this just something CloudFlare is announcing for the PR, or is it actually innovative?
- 0x0 14y agoIt sounds like they reinvented rsync to me?
- jgrahamc 14y agoNo, because we have more information than rsync does. We own both ends of the connection and can keep versions synchronized.
- 0x0 14y agoThat sounds interesting, could you elaborate on how it is different from rsync though? "Keep versions synchronized" is a bit vague
- jgrahamc 14y agoThe piece in the CloudFlare network and the piece in the customer network are able to keep track of which page versions they each have and so the part in the CloudFlare network sends a request saying "Please do GET /foo and compress it against version X". That means that at request time there's no back-and-forth between the components deciding what compression dictionary to use.
- 0x0 14y agoA bit like rsync's --fuzzy or --compare-dest then?
- jgrahamc 14y agoWell, fuzzy tries to find something to use as a 'destination' file so it can send across some hashes. Railgun has more complete information because it is keeping synchronized and thus the part making a request can specify the dictionary to compress with in a single hash.
- 0x0 14y agoThanks for the explanation, that does sound useful! :)
- IheartApplesDix 14y agoDon't you understand? We need you to accept that this is a new technology and a ground-breaking algorithm and a new innovative (and valuable, non-obvious) technology. CloudFlare was established in 2007 with the goal to develop a faster, safter, better internet. CloudFlare, the web performance and security company, set records this month hitting more than 100 million daily active users and more than 50 billion monthly page views!
- DannyBee 14y agoWell, no good binary delta algorithm uses compression dictionaries anyway (since they are binary deltas, not compression algorithms :P), except to compress the newly added strings, which you can't avoid. Note of course, that relying on the data not being corrupt on the client (which you must if you assume the compression dictionaries are sane) is dangerous. I assume you guys must store some checksum that you compare once to make sure when someone says "i have version 5, delta against this", that they really have a good copy of version 5? SVN used to what you are suggesting, btw. We only send clients deltas against the versions they already have, and precompute them in some cases :)
- 14y ago
- abcd_f 14y ago> we have more information than rsync does. To what end? Rsync too works off both copies.
- DannyBee 14y agorsync is going to perform checksums on blocks to see if the blocks are the same. It transmits these checksums, and where the checksums differ, it deltas the blocks. Note that insertion/deletion in a file can push block boundaries off between two files, causing a problem known as "stream alignment", which can cause your binary delta to be much larger because it doesn't realize the block really shifted 16384 bytes over (or whatever), and so it thinks the client really doesn't have any of the bytes of that block. In any case, if you know the files are related, you 1. Don't need to do any of this. You can simply send the binary delta that is is usually copy/add instructions (IE copy offset 16384, length 500 to offset 32768) 2. Can precompute the deltas. You can actually precompute in any case, it just makes no sense unless you know you will be diffed against something else.
- huhtenberg 14y agoI always thought rsync detected block moves and that's what made it a worthy PhD thesis.
- DannyBee 14y agoYes, I simplified and I shouldn't have. It does detect them, but it does have a minimum size of block move it can detect due to the signature matching method.
- zobzu 14y agoI thought that too. It'd be interesting to see a comparison of the two software designs with actual difference in resource usage (cpu, io, bandwidth). That would be really cool in fact.
- dubya 14y agoI think it's more like rsync + git. You have copies of previous versions, and just ask for their hash to figure out which previous version to diff against the current version, then send the diff.
- calinet6 14y agoI think that's a good analogy; it's rsync with versioning. Very cool.
- judofyr 14y agoNot sure why you're downvoted. This is basically what rsync is: a smart algorithm for doing rolling checksums and only sending diffs.
- coldtea 14y agoBecause every diff algorithm "reinvents rsync", right?
- rjknight 14y agoThe title suggests that there's something unique about Go, either the language or its standard library, that enables bandwidth savings. In fact, Cloudflare have written some software which they claim enables them to reduce their bandwidth, and this software happens to be written in Go. This might be an excellent choice (and I suspect it probably is), but it's not Go per se that is reducing the bandwidth usage.
- jgrahamc 14y agoI agree. The benefit of using Go is that it's fast to write and has good concurrency features. To give you an idea of the size, there are 7,329 lines of Go code in Railgun (including comments) and a 6,602 line test suite. In the process we've committed various things back to Go itself and at some point I'll write a blog on the whole experience, but one thing that made a big difference was to write a memory recycler so that for commonly created things (in our case []byte buffers) we don't force the garbage collector to keep reaping memory that we then go back and ask for. The concurrency through communication is trivial to work with once you get the hang of it and being able to write completely sequential code means that it's easy to grok what your own program is doing. We've hit some deficiencies in the standard library (around HTTP handling) but it's been fairly smooth. And, as the article says, we swapped native crypto for OpenSSL for speed. The Go tool chain is very nice. Stuff like go fmt, go tool pprof, go vet, go build make working with it smooth. PS We're hiring.
- shanelja 14y agohttp://www.jobscore.com/jobs/cloudflare/technical-customer-support/cNW2NomN0r4QnGeJe4efaV?ref=rss&sid=68 http://www.jobscore.com/jobs/cloudflare/technical-customer-s... Not going to lie - I'm heavily considering taking this as an entry level position to get my foot in the door.
- xxdesmus 14y agoPlease do consider applying if you're interested. We are actively looking for qualified technical folks.
- corresation 14y agoI was just looking into what SDCH is (an accept-encoding option from Chrome) and it sounds very, very similar: It generates a dictionary and then uses VCDIFF between requests. Is this related somehow?
- jgrahamc 14y agoVaguely. Both Railgun and SDHC work by compressing web pages against an external dictionary. In SDHC the dictionary must be generated (somehow), and it is intended for use between a web server and browser. Railgun is back-end for our network and automatically generates dictionaries. http://calendar.perfplanet.com/2012/efficiently-compressing-dynamically-generated-web-content/ http://calendar.perfplanet.com/2012/efficiently-compressing-...
- corresation 14y agoThat is superbly illuminating. Thank you.
- jgrahamc 14y agoIt seems like SDCH has been around for 4 years, I presume the lack of data means it hasn't worked out. A barrier to implementation of SDCH is deciding what dictionaries to create and when to update them.
- jws 14y agoIs anyone aware of a performance analysis between SDCH and one of the dynamic compressions like deflate? I google, but all I find is people complaining their proxy/filter/appliance/diagnostic is breaking because it doesn't understand SDCH. It seems like SDCH has been around for 4 years, I presume the lack of data means it hasn't worked out. (I imagine that you could drastically reduce the CPU load of compression by making simple hard coded state machines for each dictionary. For content like XML or json you could easily make your field names and surrounding punctuation minimal. For many very short messages sharing a dictionary that would beat deflate on compression ratio, and for long messages of non-repeating field values it wouldn't be much worse. CPU use of expansion is probably comparable, though you might get better memory access behavior out of SDCH.)
- silvertonia 14y agoCould be very cool. I couldn't get through the article because it read like a press release. Maybe if someone who hasn't been spoon-fed the story reports on it, I'll take notice.
- peterwwillis 14y agoI don't know why you're being downvoted, the article is written pretty shittily. The article is mostly just quotes from jgc and the CEO and some filler by the writer. Also the assertion that "It has already cut the bandwidth used by 4Chan and Imgur by half" sounds disingenuous and possibly not backed up by moot's quote “We've seen a ~50% reduction in backend transfer for our HTML pages (transfer between our servers and CloudFlare's),”. Is backend transfer for HTML pages the only bandwidth they're using? Is the rest of it halved, and if so, how and why? The title of the story also makes me gag.
- meh02 14y agoConsidering that imgur serves... images... no, their transfer isn't being cut by 50%. This article is just a press release about a feature that has been done forever (look up "WAN accelerator").
- bitcartel 14y agoThe bandwidth reduction is due to use of a binary protocol, not Go. It just so happens the server code is written in Go and C. From the article: “Go is very light,” he said, “and it has fundamental support for concurrent programming. And it’s surprisingly stable for a young language. The experience has been extremely good—there have been no problems with deadlocks or pointer exceptions.” But the code hit a bit of a performance bottleneck under CloudFlare’s heavy loads, particularly because of its cryptographic modules—all of Railgun’s traffic is encrypted from end to end. “We swapped some things out into C just from a performance perspective," Graham-Cumming said. “We want Go to be as fast as C for these things,” he explained, and in the long term he believes Go’s cryptographic modules will mature and get better. But in the meantime, “we swapped out Go’s native crypto for OpenSSL,” he said, using assembly language versions of the C libraries.
- jgrahamc 14y agoThe binary protocol means we don't add (much) overhead, the bandwidth reduction is because we are sending page diffs which themselves are encoded in a compact binary format.
- shanelja 14y agoOn another note, it's always nice to see such an influential part of the HN community giving quotes for sites like this - not only does it make me a little proud to be associated with any of you, it makes me more hopeful for the chances of my future that I can call myself one of us.
- lclarkmichalek 14y agoSuccess by association seems about as valid as guilt by association.
- DoubleCluster 14y agoThis is WAN optimization, right? This is already being done but usually for (VPN) connections to other branches of a company.
- peterwwillis 14y agoNo. This is basically binary diffing and compression. Edit: err, you are correct, I didn't realize WAN optimization included binary diffing and compression. Should google before I comment.
- meh02 14y agoPeople have been doing this for at least 10 years. CloudFlare is a YC thing, so that's why this is even news.
- songgao 14y agoI'm curious about the crypto part. Could anybody explain to me, if it's a HTTPS link, where does SSL encryption happen? Does Railgun listener talk with the origin server over HTTP or HTTPS? If it's HTTP, then how does CDN handle certificates? Does it use CDN's certificates? If it's HTTPS, then 1) Isn't hash gonna be a lot different if if the two versions are very alike? 2) Why does Railgun encrypt the encrypted data again?
- jgrahamc 14y agoThe link between CloudFlare and the customer network (i.e. between the two bits of Railgun) is TLS. We have an automated way of provisioning and distributing the certificates necessary for that part. For the connection from Railgun to the origin server it will depend on the protocol of the actual request being handled. If HTTPS Railgun makes an HTTPS connection to the origin.
- songgao 14y agoThanks! That makes sense now :-)
- cobrabyte 14y agoThis is the third time this week that I've read or heard about Communicating Sequential Processes (CSP), the formal programming language devised by Sir Tony Hoare. Third time's a charm. Definitely going to have to investigate.
- justinsb 14y agoI think this is just RFC 3229, with a binary protocol (?) http://www.ietf.org/rfc/rfc3229.txt http://www.ietf.org/rfc/rfc3229.txt I've always thought there were some potential attacks there around cache disclosure (which Google avoided by going with SDCH instead). CloudFlare controls the server and the client, so they don't need to worry about the attacks or about persuading everyone to adopt their RFC.
- glymor 14y agoHow large is the per site cache? Are cookies part of the hash (and if so how do you strip meaningless cookies)? Otherwise the this is more compelling for content sites like the referenced 4chan. But still very cool.
- jgrahamc 14y agoThere isn't a per-site cache in Railgun because it's part of our large shared in-memory cache in our infrastructure. Currently, cookies are not part of the hash. We have customers of all types using Railgun. As an example, there's a British luggage manufacturer who launched a US e-commerce site last month. They are using it to help alleviate the cross-Atlantic latency. At the same time they see high compression levels as the site boilerplate does not change from person to person viewing the site. What sort of sites do you think it doesn't apply to?
- justinsb 14y agoSurely there is a per-site cache on the origin server (in what you call the "Listener")?
- jgrahamc 14y agoYes. That's up to the particular configuration of the site. It varies from site to site, but for optimal results you want it big enough to keep the content of the common pages of your site.
- glymor 14y ago> What sort of sites do you think it doesn't apply to? Single page webapps. In those cases the html/js is normally static and already CDN'ed and the data is a JSON API which varies on a per user basis. There would be some gain as the dictionary would learn the JSON keys but I doubt it would be very dramatic vs deflate compared to the content sites referenced in the article.
- justinsb 14y agoPresuming this is RFC 3229, this is transport compression, not webserver offload. The response is generated by the origin webserver as normal. But rather than sending that response using the normal HTTP encoding, instead the proxy first does a binary diff against any versions that the (CloudFlare) client says it has and that the (CloudFlare) proxy also has in its cache. They use e.g. ETags or MD5 to uniquely identify the entire response content. You can still do cookie stripping etc to try to avoid the request to the webserver altogether, but that's a separate concern.
- sigil 14y agoQuestion for jgrahamc: how much more efficient is your binary delta algorithm than cperciva's bsdiff [1]? I assume since you've got the preimages of compression, as well as control over the compression format, that the diff and patch operations are much more efficient in space and time than they would be with arbitrary binary data. But...by how much? [1] http://www.daemonology.net/bsdiff/ http://www.daemonology.net/bsdiff/
- j_s 14y agoI am particularly interested in this aspect of the discussion (explaining the process leading to deciding to develop a new tech in-house instead of re-using any existing approach). In an ideal world there would be plenty of experimentation with real-world data to justify things, but I don't read about that happening too often.
- sigil 14y agoAgree with you there on "profile first." But knowing jgrahamc, he did -- and I'd love to know the results.
- jgrahamc 14y ago:-) Yes, that's very true. Initially, I wasn't actually planning to do deltas for the compression technique and it was in testing with a whole bunch of common sites that I stumbled upon the fact that they don't change very much. That lead me to wonder about the algorithms that might be used. I did test quite a lot of stuff (and at one point thought I'd come up with a truly cool new algorithm only to realize that I was mistaken :-) to decide what to do. Railgun has to trade off three things: compression efficiency, space and time. Because we are trying to do this for performance time is the most important thing to optimize for, followed by efficiency, followed by space. bsdiff is very, very good at delta compressing binary things; Railgun isn't as good, but it's very, very fast.
- thristian 14y agoOut of curiosity, can you say anything about the algorithm you are using? A year or two ago I got quite interested in delta compression, read all the papers I could find on the topic, and eventually came up with an algorithm that seems pretty competitive, although I've mostly focussed on efficient compression rather than speed. Someday I'll get around to porting the code from Python to C and find out what the performance is really like. For what it's worth, my code is here: https://gitorious.org/python-blip https://gitorious.org/python-blip
- jamieb 14y agoFTA: "If it was written in C++, it would be threaded code" Uh, why?
- wmf 14y agoBecause many people find threads easier to understand than callbacks?
- jussij 14y agoBecause that’s one approach to getting the most out of all of those multiple core CPU servers. For Go that came for free because its Communicating Sequential Processes design does that for you.
- abraininavat 14y agoCame for free? Go takes advantage of multiple cores by using threads. CSP doesn't magically multiplex your code onto your cores.
- jussij 14y ago> CSP doesn't magically multiplex your code onto your cores Take a look at this Rob Pike video: http://blog.golang.org/2013/01/concurrency-is-not-parallelism.html http://blog.golang.org/2013/01/concurrency-is-not-parallelis... Now that video might well be crap, I'll be the first to admit I'm not skill enough to know one way or the other. But based on that video, it does appear to me that Go does offer some form of multi-core magic and it does appear to come at a minimal cost.
- abraininavat 14y agoIt's not magic. It's threads. Go multiplexes your goroutines onto N OS threads. There are also abstractions in C/C++ (though of course as libs, not part of the language, like in Go) which hide the usage of threads. But there is no magic. If your code is running in parallel, your code is using OS threads.
- xanadohnt 14y agoThe change detection algorithm is clever. But this is a classic memory vs. processor problem. The real trick here is that the Railgun service instantly adds massive amounts of cache to your service; it just so happens - if their claims aren't inflated - adding these additional resources to your service is transparent. This has nothing to do with Railgun being developed on Go.
- tuxidomasx 14y agoOther than general traffic data compression, I've always been somewhat interested in html compression in particular. I know lots of webservers zip their response data, but I was always curious about the things in html that show up very often and if there's a way to optimize around that. For example, most web xml data contains a lot of common tags, like "div" and "span" and others that are specific to html. I think if you add them up, they might make up a considerable percent of traffic data. Is it possible for the web server to swap those out for a single character before it sends the data, and have the browser replace it when it arrives? Or does zip compression already do that somehow?
- wisty 14y agoYeees, no. Zip will replace the common tags (like "div") with a single "div" (in the compression dictionary), then a single character every time it appears (more or less - it might be less than a single byte if it's a really common tag). So there'll be a wasted overhead of a dictionary of common tags (which is kind of wasted). It would be more efficient if both the browsers and compression algorithms could agree (beforehand) had a dictionary of common terms which would be likely to appear in the document. If you're compressing a lot of data which is likely to be similar, you can do this with a common dictionary. See - http://stackoverflow.com/questions/479218/how-to-compress-small-strings http://stackoverflow.com/questions/479218/how-to-compress-sm... Of course, my answer on Stackoverflow is pretty crude. You could create a dictionary used to compress the compression dictionary. Google is probably going to do this any time soon (if they haven't already) since they control the client (Chrome), server (google web server) and protocol (SPDY).
- philiac 14y agoThe article mentions how this compression technique is similar to image compression. Would anyone care to explain, in detail if necessary, how this is so? Thanks.
- radd9er 14y agoI think its because a whole bitmap isnt streamed for every new frame, just a diff telling the player about the parts of the map that need updating.
- coolj 14y ago> Today, [cloud providers Amazon Web Services and Rackspace, and thirty of the world’s biggest Web hosting companies] announced that they will support Railgun... I can't find any such announcements; anybody have links? Based on comments further down, I wonder if the author is confused. > CloudFlare will provide software images for Amazon and RackSpace customers to install That is very different from the claim in the first paragraph.
- eastdakota 14y agoAmazon and Rackspace customers need to install the software themselves (for now). The other listed hosts have made it one-click simple without the customer having to install anything. A couple announcements from major hosts today: Dreamhost: http://dreamhost.com/dreamscape/2013/02/26/cloudflare-railgun/ http://dreamhost.com/dreamscape/2013/02/26/cloudflare-railgu... Media Temple: http://weblog.mediatemple.net/2013/02/26/the-web-just-got-faster-with-railgun/ http://weblog.mediatemple.net/2013/02/26/the-web-just-got-fa...
- shotgun 14y agoI see that the article is tagged "open source." Is CloudFront going to open source Railgun? Publish any papers? This isn't an announcement about companies supporting Railgun...it's about companies supporting CloudFlare by installing the Railgun Listener.
- zobzu 14y agoUho. Binary protocol. The problem being, it's actually bringing financial advantages over HTTP. HTTP has the advantage of being standard, simple, plain text and thus easy to work with. Hopefully http2.0 will attempt solving this.. erm...
- pjmlp 14y agoAnother Go PR story. Same thing could be easily achieved using futures or any of the asynchronous libraries available to C++, Ada, JVM and .NET languages.