5 ms·
Hey, developer on zipkin here. Happy to answer any questions you have.
by bmdhacks 13y ago
Hey, developer on zipkin here. Happy to answer any questions you have.
- groat 13y agoHow is this different from, say, sending straces or something similar to Splunk?
- bmdhacks 13y agoZipkin is designed to piece together operations that are distributed across a fleet of systems. IE: nginx calls your rails server which calls memcache then calls your login server which hits a mysql database. It's most useful when your serving architecture is distributed across many many machines and many many different services.
- dpratt 13y agoWe built something similar last year for our internal services, but ran into problems when dealing with asynchronous code and threading context. How do you propagate the Zipkin info from thread to thread inside your code? An example - a request comes in, we generate a new request ID and pass this down into the processing code. Part of this code executes an async call to an external service with essentially a callback (in reality, it's a scala.concurrent.Future executing on an arbitrary context or an akka actor) - how do you properly rehydrate the Zipkin info when the response comes back? The only way we could think of is some sort of fiendishly complex custom ExecutionContext that inspected the thread local state at creation time and recreates that in the thread running the callback, or just have pretty much every method take an implicit context parameter. Neither of those solutions worked well, so we've largely bailed on the concept for the bits of our code that don't execute in a linear/blocking fashion.
- bmdhacks 13y agoTwitter util to the rescue! https://github.com/twitter/util/blob/master/util-core/src/main/scala/com/twitter/util/Local.scala https://github.com/twitter/util/blob/master/util-core/src/ma...
- dpratt 13y agoWe actually built something quite similar to that, but how do you avoid having to artificially insert a save() and restore(context) call all over the place? It looks to me like you'd have to do something like val context = Local.save() val eventuallyFoo = someServiceClient.makeCall("data").map { result => Local.restore(context) //do work } for every single async interaction.
- bmdhacks 13y agoThat's what we do, but it's built into the RPC library we use, finagle: https://github.com/twitter/finagle/blob/master/finagle-core/src/main/scala/com/twitter/finagle/tracing/Trace.scala#L181-L184 https://github.com/twitter/finagle/blob/master/finagle-core/... BTW, to the readers at home, note how almost all our core infrastructure is open-sourced.
- ludwigvan 13y agoThat's intriguing. Could you please elaborate on that? Does this mean that Twitter sees no danger in exposing these? Which parts of the infrastructure would be considered secret sauce and not open source? Or does it not matter when the company is as big as Twitter, since the core strength lies in user base, not technical infrastructure? What does Twitter primarily seek to achieve when it open sources its stuff? Talent acquisition or brand image or other benefits of open source such as collaboration?
- mariusae 13y agoMore or less 100% of asynchronous composition is done via Futures[1] whose default implementation[2] does this for you. [1] https://github.com/twitter/util/blob/master/util-core/src/main/scala/com/twitter/util/Future.scala https://github.com/twitter/util/blob/master/util-core/src/ma... [2] https://github.com/twitter/util/blob/master/util-core/src/main/scala/com/twitter/util/Promise.scala#L128 https://github.com/twitter/util/blob/master/util-core/src/ma...
- mariusae 13y ago
- dum 13y agodoes the name have anything to do with turkish word "zıpkın"?
- Ologn 13y agoI'm happy to see the video of the Zipkin presentation at Strange Loop was put up a few months ago ( http://www.infoq.com/presentations/Zipkin http://www.infoq.com/presentations/Zipkin ). People working on distributed tracing systems all tend to eventually come up with a similar architecture. Even going back to the IPS research system made at University of Wisconsin in the late 1980s (that's the earliest distributed tracing system I know of). They all tend to do tracing via minimalistic, low overhead logging of RPC calls between machines. They tend to do tracing via low-level libraries which application developers can ignore. The trace systems seems to be good at uncovering latency bottlenecks. I am ignorant of what success systems like Zipkin or Google's Dapper may or may not have had in areas outside of latency checks.
- bmdhacks 13y agoOne of the projects I've been working on is doing more aggregate analysis of zipkin data: https://github.com/twitter/zipkin/pull/276 https://github.com/twitter/zipkin/pull/276 We're using this at Twitter to better understand usage patterns for services upstream and downstream. For example, Gizmoduck, the user store at twitter is backed by memcache, and some disk-based storage behind the cache. While we can view individual traces that hit by memcache, the aggregate info shows us both the proportion of traffic for services calling Gizmoduck, as well as the proportion of time Gizmoduck spends in memcache versus the backend store. Furthermore, it can be useful for notifying of unusual behavior. If a service's aggregate durations has changed since yesterday, perhaps that's something we want to look at. Or if the ratio of traffic from some upstream service doubles, that's interesting to know.
- dkuebric 13y agoLatency isn't the only thing--if that's all you wanted, you wouldn't need the tracing, just log the latency at each step. The real amazing thing is the structure of requests. Having a record of distibuted N+1 problems is pretty amazing, for instance. As for where you hook into the stack, it's definitely a tradeoff. The lower level you go, the wider your coverage, but it also becomes more generic and potentially less actionable data. I've worked on building a similar system at AppNeta (it's basically commercial Google Dapper). Here's some slides from a lightning talk I gave about distributed tracing at Surge last year: http://www.slideshare.net/dkuebrich/distributed-tracing-in-5-minutes http://www.slideshare.net/dkuebrich/distributed-tracing-in-5...
- ldng 13y agoCould you explain a bit more how it works because, if a see the big picture, I fail to see how it could be useful outside Twitter's stack. I mean it seems to heavily rely on Finagle, right ? Please correct and explain if I'm wrong.
- theatrus2 13y agoZipkin depends on the services in the stack to be able to capture trace identifiers, record and send to the collector timing information, and forward trace identifiers for outbound calls. This is easy in Finagle as all of this logic is already abstracted away for most of the protocols. Note that Finagle is open source so its ready for use today. If you wanted to record and forward traces inside your own services, this is relatively straightforward (though not super trivial) to do. There are other implementations beyond the Finagle one in development by the community - I would suggest hopping on to the zipkin-users Google Group.