5 ms·
Pinterest open-sources Terrapin, a tool for serving data from Hadoop
- kevinbowman 11y agoThe URL https://engineering.pinterest.com/blog/open-sourcing-terrapin-serving-system-batch-generated-data-0 https://engineering.pinterest.com/blog/open-sourcing-terrapi... gives more info, which this article links through to.
- arthurcolle 11y agoI wonder if the authors are Maryland alumni!
- optimusclimb 11y agoEither more tools like this are going to pop up, or the existing ones will mature, as more people adopt Lambda style architectures, I'd imagine. While building one out, we looked at VoldemortDB, SploutSQL, and ElephantDB to serve bulk data coming out of Hadoop in batches. Voldemort turned out to be much rougher around the edges than expected, ElephantDB looked very bleeding edge, and SploutSQL wasn't as general purpose. In the end we turned to Cassandra and this tool - https://github.com/spotify/hdfs2cass https://github.com/spotify/hdfs2cass. Good to see Pinterest open sourcing this.
- ameyamk 11y agoBulk Uploads into KV stores are slow - so Terrapin allows KV access over immutable HDFS files. Very typical use case for recommendation systems etc. We face similar problems with latencies on HBase (At Groupon). So this solution seems interesting. Would be good to have comparison of other solutions Pinterest tried before building this. eg. loading data into Cassandra instead of HBase etc. In nutshell - very specific use case - but the one which comes across very often
- sjg007 11y agoWhat do you think of Apache Drill onto of say HDFS served files?
- varunsharma 11y agoThis is Varun from Pinterest. We did look at a few options before building this. ElephantDB seemed a bit heavy handed, such as having to modify ring configuration every time we added/removed servers and also, modifying domain spec yaml files for newly added data sets. It did not allow us to easily change # of shards across different versions of the data - something that our developers do often to make their jobs run faster etc. Also, it does not GC out older versions and since our workflows write new versions every day, this was a problem. We did look at Cassandra but we also did not want to operate another data store. However, we definitely wanted to get the data loaded fast i.e. through simple file copy operations. We found that for this option Cassandra had similar issues as HBase i.e. having to do major compactions to get rid of older data versions. Tweaking the number of reduce shards was also harder. With Terrapin, we essentially tried to build serving system on top of HDFS given the recent improvements in HDFS performance when there is data locality. We felt that HDFS was rock solid and the best storage system (in terms of scalability & ease of operation) for immutable data sets. On top of that, we built versioning, cheap garbage collection, extensible serving formats etc. as mentioned in the blog As for Apache Drill, it is more suited to running analyst queries with latencies ranging upto seconds or 100s of milliseconds. This is not acceptable for webscale work loads where the latencies must be < 10ms for lower level serving systems like terrapin.
- ameyamk 11y agoHi Varun, Can you also elaborate, how you read HFiles and serve it out from Terrapin servers? Are you using similar functionality as HBase? (With block cache like design if yes how do you keep both in sync). Your blog is missing this interesting detail.
- varunsharma 11y agoThat is correct, we are using the functionality similar to HBase. We pull in the HBase BlockCache library with some tweaks to make it work for our scenario. Note that data is never overwritten and HFiles are immutable. So the cache automatically, gets evicted/populated as HFiles are opened and closed. That said, there is a possibility to use more performant formats like rocksdb etc. in the future (the format is pluggable). Or even still use HFiles and have them loaded into some kind of specialized in memory data structure etc.
- varunsharma 11y agoThis is Varun from Pinterest. We did look at a few options before building this. ElephantDB seemed a bit heavy handed, such as having to modify ring configuration every time we added/removed servers and also, modifying domain spec yaml files for newly added data sets. It did not allow us to easily change # of shards across different versions of the data - something that our developers do often to make their jobs run faster etc. Also, it does not GC out older versions and since our workflows write new versions every day, this was a problem. We did look at Cassandra but we also did not want to operate another data store. However, we definitely wanted to get the data loaded fast i.e. through simple file copy operations. We found that for this option Cassandra had similar issues as HBase i.e. having to do major compactions to get rid of older data versions. Tweaking the number of reduce shards was also harder. With Terrapin, we essentially tried to build serving system on top of HDFS given the recent improvements in HDFS performance when there is data locality. We felt that HDFS was rock solid and the best storage system (in terms of scalability & ease of operation) for immutable data sets. On top of that, we built versioning, cheap garbage collection, extensible serving formats etc. as mentioned in the blog As for Apache Drill, it is more suited to running analyst queries with latencies ranging upto seconds or 100s of milliseconds. This is not acceptable for webscale work loads where the latencies must be < 10ms for lower level serving systems like terrapin.
- varunsharma 11y agoThis is Varun from Pinterest. We did look at a few options before building this. ElephantDB seemed a bit heavy handed, such as having to modify ring configuration every time we added/removed servers and also, modifying domain spec yaml files for newly added data sets. It did not allow us to easily change # of shards across different versions of the data - something that our developers do often to make their jobs run faster etc. Also, it does not GC out older versions and since our workflows write new versions every day, this was a problem. We did look at Cassandra but we also did not want to operate another data store. However, we definitely wanted to get the data loaded fast i.e. through simple file copy operations. We found that for this option Cassandra had similar issues as HBase i.e. having to do major compactions to get rid of older data versions. Tweaking the number of reduce shards was also harder. With Terrapin, we essentially tried to build serving system on top of HDFS given the recent improvements in HDFS performance when there is data locality. We felt that HDFS was rock solid and the best storage system (in terms of scalability & ease of operation) for immutable data sets. On top of that, we built versioning, cheap garbage collection, extensible serving formats etc. as mentioned in the blog As for Apache Drill, it is more suited to running analyst queries with latencies ranging upto seconds or 100s of milliseconds. This is not acceptable for webscale work loads where the latencies must be < 10ms for lower level serving systems like terrapin.