13 ms·
Open-sourcing a 10x reduction in Apache Cassandra tail latency
- alsadi 9y agoCan we add lz4 to the blend to reduce disk IO?
- agnivade 9y ago> We also observed that the GC stalls on that cluster dropped from 2.5% to 0.3%, which was a 10X reduction! Umm .. shouldn't the stalls go to 0, because now you have moved to C++ ? Or is this the time it takes for the manual garbage collection to occur ?
- Thaxll 9y agoWeird, did they try to use https://www.scylladb.com/ https://www.scylladb.com/?
- HippoBaro 9y agoMy thought exactly. Would be interesting to know if they did and if yes, why they chose to develop something in-house anyway.
- en4bz 9y agoI was going to say the same thing. It seems pretty clear at this point that Java is not a good programming language to build a database on if you care about strong 99% latency guarantees. The engineers in the article came to this conclusion and so did the Scylla people years ago. Scylla is AGPL for the OSS version though so testing it out would not be an option without getting a commercial license first.
- HippoBaro 9y agoDoesn't AGPL allow commercial use?
- halestock 9y agoPer the AGPL preamble[0]: "The GNU General Public License permits making a modified version and letting the public access it on a server without ever releasing its source code to the public. The GNU Affero General Public License is designed specifically to ensure that, in such cases, the modified source code becomes available to the community. It requires the operator of a network server to provide the source code of the modified version running there to the users of that server. Therefore, public use of a modified version, on a publicly accessible server, gives the public access to the source code of the modified version." [0] https://www.gnu.org/licenses/agpl-3.0#preamble https://www.gnu.org/licenses/agpl-3.0#preamble
- glommer 9y agoIt does, but there are conditions that apply. Some companies don't like going that route - which in my opinion is not a thing, but I can't understand the concern. But for testing, I don't see any impediment.
- adrianN 9y agoYes it does, but AGPL licenced software is super banned at all major companies because it is very viral. You have to make derivative software available under AGPL even if the end user accesses it only over a network.
- glommer 9y ago(ScyllaDB employee here) I don't believe one would need a commercial license just to test a product in any way? They are not making that part of any product at that point, so no concerns here.
- snovv_crash 9y agoThey can't test on production servers. Fake data, non-userfacing servers, sure.
- glommer 9y agoEven if that is a problem, we provide anyone that is interested with a 30-day evaluation license of Scylla Enterprise.
- Sir_Substance 9y agoThat puts you about 2/3rds of the way down the list of things to try. You're before all the products that have no evaluation license at all, but after all the open source options that developers can test at their leisure. If I'm trying a products evaluation license, you can be sure I've tried literally every other option under the sun first, including investigating the possibility rolling my own if situationally appropriate. No form of development is slower than the kind where I have to wait for a company in another timezone to give me permission to use their software, so it's always last on my options list unless the company has frankly amazing reviews that pique my curiosity.
- geofft 9y agoInstagram doesn't operate any user-facing Cassandra servers, though. They run user-facing web servers that talk to Cassandra internally. I don't like the AGPL because it's unclear on this exact sort of thing, but it does seem to me like the obvious reading of "all users interacting with it remotely through a computer network" does not encompass the connection between Instagram end users and their internal Cassandra. And, in any case, they released sources for the thing they came up with - which is all that the AGPL requires. If they're okay with doing that, they can definitely use the AGPL for production commercial software.
- gnud 9y agoThe server is AGPL. The client is Apache licensed. So I don't see a problem with using the AGPL version in commercial product. Noone claims that a product using the MySql driver is a derivative work of the MySql server? Edit: Of course, IANAL...
- lukev 9y agoThat's specifically what the AGPL does (as opposed to the GPL.) The copyleft "infection" is deliberately transmitted via network clients, not just static linking. So you actually can't release a permissively licensed client for an AGPL server. I mean, they did, clearly, but the AGPL itself would seem to make that inconsistent. But then none of this has ever been litigated and both the AGPL and GPL themselves are very confusingly worded so shrug.
- gnud 9y agoAs I understand it, with the GPL, you must offer source code under the GPL to everyone you distribute the software to. With the AGPL, the same goes for those that use the software over the network. So you must offer the source of the database everyone who connects to the database over the network, under the AGPL. But if you deliver a web app, not a database-as-a-service, your users don't connect to the database. And since this database uses the Cassandra protocol, I'd say your web app isn't a derived work of the database in any way. Of course, that last part is the sticky bit. But if applications using database servers via a well defined protocol are judged to be derived works, we might have other problems - hence the reference to MySql in my first post.
- snuxoll 9y agoRequiring GPL’d software to function means your product is a derivative work, full stop as far as the spirit and letter of the license is concerned. Using it over a network instead of linking against it doesn’t change this, if you depend on MySQL or any of the forks and distribute your software it must be made available under a GPL-compatible license. The requirements of the AGPL become clear in this regard as well; network use is distribution with the AGPL - incorporating AGPL’d software into your application means you must consider your entire application as licensed under the AGPL or compatible license. MongoDB muddied the waters here by deciding to interpret the AGPL differently, but I wouldn’t risk your business on it.
- ghshephard 9y agoThere are those who've deployed on Java with tight latency requirements: https://martinfowler.com/articles/lmax.html?t=1319912579 https://martinfowler.com/articles/lmax.html?t=1319912579 - Benchmarked at around 6 million transactions/second. The issue isn't so much Java the language, as it is being aware of the GC, and developing with it in mind.
- en4bz 9y agoThat removes the value proposition of Java though which is that you don't need to worry about memory management. If you need to mentally track every implicit allocation and deallocation in Java then you are essentially writing code in a kneecapped version of C++.
- mschuster91 9y ago> If you need to mentally track every implicit allocation and deallocation in Java then you are essentially writing code in a kneecapped version of C++. Well, Java has the advantage of being platform (and to a certain degree, runtime) independent, plus a robust set of best practices and ecosystem when it comes to modules and library handling, which is pretty hard to get done right for C/C++ projects.
- en4bz 9y agoNode.js is also platform independent and has a package manager but I wouldn't use it for a High Performance / Low Latency application like a database.
- majidazimi 9y ago> Well, Java has the advantage of being platform (and to a certain degree, runtime) independent What is the benefit of that? Who on earth runs a DB written in Java on windows? Any useful server software will end up using platform native features, be it SQL server, MySQL, HBase, ...
- mschuster91 9y ago
- geofft 9y ago> Scylla is AGPL for the OSS version though so testing it out would not be an option without getting a commercial license first. Huh? The AGPL is not a non-commercial-use-only license. If you have proprietary software that you would like to combine with AGPL code (i.e., not interact with as a service) and is available to the general public over the Internet, and you want keep your code proprietary, sure, you may not want to use the AGPL. But you could say the same thing about proprietary software you want to combine with GPL code and sell to the general public. If you're either using the software through it's existing defined public interfaces, or you're okay releasing anything you modify or link into the software, the AGPL (and the GPL) are fine. Lots of people distribute proprietary products that include GPL code, like Chromebooks, Android phones, routers, GitHub Enterprise, etc. We figured out years ago that the Linux kernel is not just a non-commercial product. Why are we having the same misconceptions about the AGPL?
- en4bz 9y agoI think you answered the question yourself [1]. The wording is ambiguous and as far as I know there have been no court cases yet that have yet to define what constitutes a connection between the end user and whether transitive connections count. If it's ambiguous to a software developer then corporate lawyers are definitely going to say no. [1] https://news.ycombinator.com/item?id=16523858 https://news.ycombinator.com/item?id=16523858 EDIT: I realize that AGPL is valid for commercial use but since its terms are so onerous, especially once the lawyers get involved, it effectively makes the AGPL unusable in a larger corporation.
- kevindong 9y agoThe exact interpretation varies from company to company. Some companies take the strict stance of "if you use this library in any way in your application, you must open source your entire application." I've found that some libraries explicitly state that requirement within their FAQs for their community/free edition as opposed to their commercially (and paid) licensed equivalent. At the end of the day, it's not worth risking yourself (or your company) when the owners of the library claims a software license works a certain way and you disagree. Sure you might be right and you might even prevail in court, but the potential legal fees usually aren't worth the trouble. I ran into this issue when I was selecting a library to generate PDFs for my internship over the summer: https://itextpdf.com/AGPL https://itextpdf.com/AGPL
- jjirsa 9y agoWhy throw away something proven to run at massive scale, that you understand and trust for something that's new, has never been run at that scale, and you have no experience running? If you have a team of software engineers, and the latency problem is a software problem, fix the software problem. When you already know Cassandra, and you already know RocksDB, and you already have an engineering team, it makes far more sense to combine the two things you know how to use at scale than to try to use some new thing NOBODY has run at scale.
- welder 9y ago> some new thing NOBODY has run at scale Outbrain uses ScyllaDB in production at scale across multiple data centers. Not sure if it's Instagram scale, but still enough to prove it's reliability and performance. https://www.outbrain.com/techblog/2016/08/scylladb-poc-not-so-live-blogging-third-update/ https://www.outbrain.com/techblog/2016/08/scylladb-poc-not-s...
- jjirsa 9y ago7 hosts in that poc, that is not "at scale"
- manigandham 9y agoScylla can handle 10-100x the load of Cassandra on the same servers. Scale is more than just the number of hosts.
- cnlwsu 9y agoData density is a thing. If u putting 10tb on a c* host, switching to Scylla doesn’t fix the issues that putting 1pb of data on a host would involve (ie backing that up). Throughput of 100mb of data done in marketing benchmarks are rarely relevant.
- manigandham 9y ago
- jakelarkin 9y ago- has anyone run it FB scale? for how long? - how many experienced scylladb devops are there globally that we can hire? Those questions asked at BigTechCo before it adopts somebody elses tech. FB already operates RocksDb and Cassandra so there's way less technical, career, financial risk for just hacking the two together with some aggressive refactoring.
- polskibus 9y agoDoes FB still use Cassandra? I thought they abandoned them ages ago and then databricks picked it up?
- ADefenestrator 9y agoFB abandoned Cassandra (which was really only used for message inbox indexing) when they redid how messages work years ago, but the re-adopted a large C* infra when they bought Instagram.
- wenc 9y ago> then databricks picked it up I think it's DataStax. Databricks is the company behind Spark.
- jjirsa 9y agoThe article is literally written by Instagram, which is FB.
- sciurus 9y agoWell, it's a separate product that Facebook acquired. True or not, it's a common perception that Facebook abandoned Cassandra. https://www.wired.com/2014/08/datastax/ https://www.wired.com/2014/08/datastax/
- jjirsa 9y agoNicely done! Looking forward to the pluggable storage engine.
- pas 9y agoThe JIRA tickets don't really shine with much hope :/ https://issues.apache.org/jira/browse/CASSANDRA-13474 https://issues.apache.org/jira/browse/CASSANDRA-13474 [2 comments from 2017 Apr] https://issues.apache.org/jira/browse/CASSANDRA-13475 https://issues.apache.org/jira/browse/CASSANDRA-13475 [~100 comments, but the last one is from 2017 Nov, by the InstaG engineer] And the Rocksandra fork is already ~3500 commits behind master, so upstreaming this will be interesting. Oh, and the Rocksandra fork is already kind of abandoned - no commits since 2017 Dec. (which probably means this is not actually the code that runs under Instagram.)
- jjirsa 9y agoI'm a committer, I'm familiar with the JIRA ticket.
- pas 9y agoCould you share your thoughts on how likely and how soon will the RocksDB engine be available as part of normal Cassandra? Also, how big is the gap between 3.0.x and 3.x? Any improvements between 3.0.x and 3.x regarding tail latency/performance? Thanks!
- jjirsa 9y agoI think it's likely. It's decidedly nontrivial, and the hardest part will be the (very slow) design phase where we actually make sure the interfaces are defined properly, but I think there are enough interested people to make sure it happens. There are some meaningful changes between 3.0 and 3.11 (notably a compressed chunk cache for storing some intermediate data blocks and a significant change to the way the column index is deserialized) that do help tail latencies, and there's certainly quite a bit more low hanging fruit, but the biggest contributor to p99 latencies is the GC collections, and the read path still contributes the most JVM garbage, so this is still probably a meaningful improvement over 3.11.
- fdeliege 9y agoJoin our meetup to chat with some of the developers: https://www.meetup.com/Apache-Cassandra-Bay-Area/events/248376266/ https://www.meetup.com/Apache-Cassandra-Bay-Area/events/2483...
- jjirsa 9y agoSo sad I’m not in town that week
- adrianratnapala 9y agoAs a Java scoffer trying to be fair-minded, I resisted the urge to joke that "it's was the GC, stupid" and assume that a big project like Cassandra had somehow worked around the GC latency problems. But, what? It turns out the article is really about replacing Java with C++.
- cestith 9y agoIt's about using something in one language for its features and only porting the critical sections to C++ via a clean API. This is the sort of advice we've been giving people for decades. Choose the language for what you want to build, measure and profile performance if necessary, find the bottleneck on the hot path, decouple that from the bulk of the code, and reach to a lower level for performance only in that clearly defined section. They managed to generalize one application that meets their feature needs to be a front end to another existing application with fewer features but better performance as a back end. They're optimizing their hot path by decoupling it from the rest of the application and handing off to C++ code they didn't even have to write. Adding pluggable storage engines to Cassandra means that if they make the API smooth enough they can have engines in C, C++, Erlang, Go, Rust, ML, or whatever in the future without changing their front end. That's a big win even beyond this tail latency issue.
- majidazimi 9y agoWell, other than storage engine, the next big part of a database software is the query planner/optimizer which Cassandra doesn't have (due to simple KV nature of it). So there isn't much remaining. In a long term plan, rewrite them all and you have single code base and you'll benefit from mighty C++ in other components of the database. And there is still room for more optimizations: SIMD, ... The GC problem is not limited to C*. This shit(virtual machine) is hitting the whole Hadoop stack: HDFS, Hive, Spark, Flink, Pig... Immense number of tickets in any fairly large cluster is related somewhat to GC and JVM behavior.
- rbranson 9y agoDid you all find that there were changes to the Java heap/GC configuration that would make tuning this setup different? I imagine if most everything that "sticks" is moved off heap, the GC could be tuned more heavily for young gen throughput vs trying to balance it with long-lived objects.
- dikanggu 9y agoYeah, for Rocksandra, we are able to use much smaller heap size, and most of the objects are recycled during the young gen GC.
- en4bz 9y agoHas any tried running Casandra on Azul Zing[1]? The slowdown here is not surprisingly related to GC pauses which Azul has eliminated in Zing. [1] https://www.azul.com/products/zing/ https://www.azul.com/products/zing/
- rbranson 9y agoThe licensing cost of Zing generally makes this a bad trade-off. It's much cheaper to purchase more hardware. Zing is targeted at vertically scaling very large JVM heaps, where it's valuable to have massive amounts of data on a single, big machine.
- nitsanw 9y agoAs an ex-Azul employee I can say there's a good number of Azul clients using a Zing+Cassandra setup, so the price point is right for some people at the very least. Zing licence cost has also changed in recent years (3.5k per server last I looked, and that is before you haggle some bulk deal) so not sure if your impression is calibrated to that new price point.
- jjirsa 9y agoHave friends who have used it, they report that it works reasonably well. Especially in p99.
- truth_seeker 9y agoBy what factor/magnitude p99 was improved ? Any idea ?
- spockz 9y agoActually, it appears that is one of the premises[1] they sell Zing on. [1]: https://www.azul.com/solutions/cassandra/ https://www.azul.com/solutions/cassandra/
- StreamBright 9y agoIn a similar situation we just adjust the GC and started to use G1GC which resulted in similar numbers.
- coryfoo 9y agoI bet that didn't take N engineers 12 months to build out, either
- StreamBright 9y ago2 engineers, 2 weeks because we had to evaluate every change we made with production traffic.
- threeseed 9y agoCassandra uses G1GC by default. If it was as simple as tweaking a few GC settings to get 10x improvement pretty sure Datastax would've done it by now.
- StreamBright 9y agoNot sure you understand that there are different versions of Cassandra and I never mentioned that we used Datastax version. Using G1GC with default settings does not give you anything btw.
- jjirsa 9y agoThe problem with making a general purpose DB is you have to have general purpose defaults. The Cassandra defaults are "dont crash anywhere", not "be super fast and low latency". You can definitely do 5x better than the default with some basic jvm tuning. That said: the IG folks certainly know how to tune JVMs. There are IG (and former IG, I saw rbranson post) folks in this thread that know how to tune the collectors, so assume that the 10x they see is beyond what you'd get from simple tuning.
- openasocket 9y agoI'm not an expert on these things, but it seems to me if you're implementing a database in Java you wouldn't want to keep your data on the JVM Heap, as this seems to indicate. My understanding is that in most applications (like servers) the average object lives for a very short period of time, and most GC implementations are built from that idea. But, in a database, especially an in-memory database, the majority of the objects are going to live for a very long time. That makes the mark phase of GC a lot more expensive, puts more pressure on the generations, etc. Is my guess here correct, or are there things I'm missing or mistaken on?
- wonnage 9y agoThe purpose of separating into young and old generation is that it's easier to find dead objects in the young generation (as you said, average object lives for a short period of time). You only have to scan this subset for a minor GC. It doesn't really matter how many long-lived objects you have as long as you can avoid needing to do a major GC.
- openasocket 9y agoDon't you still need to scan the old generation during minor GC, in case a field in one of the older objects was modified to point to an object in the young generation? Or are there optimizations you can use to quickly and efficiently find references from the older generation to the younger?
- bzbarsky 9y ago> in case a field in one of the older objects was modified to point to an object in the young generation? This is typically handled by https://en.wikipedia.org/wiki/Write_barrier#In_Garbage_collection https://en.wikipedia.org/wiki/Write_barrier#In_Garbage_colle...
- jakewins 9y agoThis is correct; the standard approach here is to use regular c-style memory management for the data the system is managing, and the JVM heap only for the database "infrastructure". This hybrid approach gives the benefit of a managed runtime and safety of GC for most of your code, but allows the performance of raw pointers/malloc for key code paths. Some examples of this pattern on the JVM: - The Neo4j Page Cache, Muninn, https://github.com/neo4j/neo4j/blob/3.4/community/io/src/main/java/org/neo4j/io/pagecache/impl/muninn/MuninnPageCache.java#L58 https://github.com/neo4j/neo4j/blob/3.4/community/io/src/mai... - The Netty projects implementation of jemalloc for the JVM: https://github.com/netty/netty/blob/4.1/buffer/src/main/java/io/netty/buffer/PooledByteBufAllocator.java https://github.com/netty/netty/blob/4.1/buffer/src/main/java...
- deleted 9y ago[deleted]
- dikanggu 9y agoWe do want to contribute our work back to the Cassandra upstream, instead of keeping it as a fork. So that more users from C* community can benefit from the improvements. The pluggable storage engine is an ambitious project (https://issues.apache.org/jira/browse/CASSANDRA-13474 https://issues.apache.org/jira/browse/CASSANDRA-13474). Any help will be appreciated!
- russellspitzer 9y agoSaw you talking about this on the Distributed Data Show https://academy.datastax.com/content/distributed-data-show-episode-37-cassandra-instagram-dikang-gu https://academy.datastax.com/content/distributed-data-show-e...
- OhDagny 9y agoFundamentally, the "pluggable storage" is probably the driver protocol. As in Scylla, Cassandra, or Cass/Rocks
- gfosco 9y agoRocksDB is used all over Facebook, powers the entire social graph. Great storage engine that pairs well with multiple DBMS: MySQL, Mongo, Cassandra... We'll be at Percona Live 2018 in April, giving several talks, and are looking forward to hanging out and talking with users in our lounge area. We're working hard to support our open source community as well! https://github.com/facebook/rocksdb https://github.com/facebook/rocksdb
- cmrdporcupine 9y agoI remember using quite early versions of Cassandra back in an ad-tech startup I was at back in 2009 or 2010, spending unfortunate amounts of time fighting the JVM GC and trying to tune things so it behaved responsibly. It was a real problem then and I know a lot of work went into fixing GC behaviour. Then I stopped using Cassandra for work, but it's unfortunate this is still an issue? What I took out of that is that I really feel like something like Cassandra is better suited to implementation in a language like C++ or Rust. And I believe others have since come along and done this. I really liked the gossip-based federation in Cassandra though.
- ADefenestrator 9y agoIt's still an issue, but a lot less of one. 3.0 is a big improvement in terms of GC behavior. Haven't tested 3.11.x yet, but it look in theory like a decent improvement in terms of rounding off corner cases and adding instrumentation.
- estebank 9y agoIt sounds like you might be interested in TiKV. https://github.com/pingcap/tikv https://github.com/pingcap/tikv
- cmrdporcupine 9y agoThanks. Since coming to Google I don't get the opportunity to compare/evaluate/deploy tools like this anymore. Smarter people than me make choices like that :-)
- 3uclid 9y agoUnrelated: as a CS undergrad, I read this article and was immediately inspired. This is definitely the type of work I want to be doing when I graduate (infrastructure engineering). But my next thought was: where do I start?! Any advice?
- therealdrag0 9y agoI'd say no matter what kind of job you get, you can put 10% of your time into similar problems. Even simple CRUD apps can have interesting problems like this. In my experience every project has instances of engineers shooting themselves in the foot, or unforeseen problems cropping up. If you have a bit of self-motivation you can dig into them and learn a lot and improve things. I do this and find it very satisfying.
- en4bz 9y agoCMU Database Group Lectures: https://www.youtube.com/channel/UCHnBsf2rH-K7pn09rb3qvkA https://www.youtube.com/channel/UCHnBsf2rH-K7pn09rb3qvkA
- jjirsa 9y agoHappy to help you get started working on Cassandra. http://cassandra.apache.org/doc/latest/development/patches.html http://cassandra.apache.org/doc/latest/development/patches.h... Has some basic entry pointers. There’s also a dev mailing list that’s reasonable active.
- ddorian43 9y agoStill in school ? (don't understand different <type>grad). See: GSOC Seastar Framework https://summerofcode.withgoogle.com/organizations/6190282903650304/ https://summerofcode.withgoogle.com/organizations/6190282903...
- 3uclid 9y agoYeah, still in school (3rd year). I have intern experience, but it seems like these type of positions are way too advanced for me at the moment. Just unsure how to progress...
- haglin 9y ago"To reduce the GC impact from the storage engine, we considered different approaches and ultimately decided to develop a C++ storage engine to replace existing ones." I wonder how the numbers would have looked with the new low latency GC for Hotspot (ZGC). https://wiki.openjdk.java.net/display/zgc/Main https://wiki.openjdk.java.net/display/zgc/Main Early results from SPECjbb2015 are impressive. https://youtu.be/tShc0dyFtgw?t=5m1s https://youtu.be/tShc0dyFtgw?t=5m1s
- itronitron 9y agoyes, a comparison across multiple JVMs would be nice
- tibbetts 9y agoYes, also Azul Zing. Really anytime someone says they have a problem with GC and suggests spending a million dollars of engineer time building a new system, they should consider Zing first. It works and is a way more efficient way of spending money to fix GC latency problems.
- ADefenestrator 9y agoFor a small to medium sized shop, sure. For someplace with thousands or tens of thousands of nodes, the new system ends up cheaper in the long run.
- majidazimi 9y agoBecause GC related issues don't undergo from "problem" state to "solved" state. It is just a never ending stream of issues, that the team need to resolve, specially in a database realm in which metrics are hugely workload dependent.
- the8472 9y ago> The graph shows that a Cassandra server instance could spend 2.5% of runtime on garbage collections instead of serving client requests. The GC overhead obviously had a big impact on our P99 latency No, this is not obvious. If you have a fully concurrent GC then spending 25 out of 1000 CPU cycles on memory management does not "obviously" have an impact on your 99th percentile latency. It would primarily impact your throughput (by 2.5%), just like any other thing consuming CPU cycles. > We defined a metric called GC stall percentage to measure the percentage of time a Cassandra server was doing stop-the-world GC (Young Gen GC) and could not serve client requests. Again, this metric doesn't tell you anything if you don't know how long each of the pauses are. If they are at the limit infinitesimally small then you are again only measuring the impact on throughput, not latency. Certainly, GCs with long STW pauses do impact latency, but then you need to measure histograms of absolute pause times, not averages of ratios relative to application time. That's just a silly metric. And neither does the article mention which JVM or GC they're using. Absent further information they might have gotten their 10x improvement relative to some especially poor choice of JVM and GC.
- foolfoolz 9y agoclassic hacker news comment. this thing you built and open sourced, has gotten you real measurable results? allow me to list the many ways you’re probably wrong and doing it incorrectly
- discoursism 9y agoMeasurable results are all well and good, but it can be helpful to know how the baseline was established. Measurable results aren't "portable" without a well-established baseline.
- teacpde 9y ago> If you have a fully concurrent GC then spending 25 out of 1000 CPU cycles on memory management does not "obviously" have an impact on your 99th percentile latency. I try to understand the meaning. Is it saying the latency caused be GC is applied to all requests, not just the ones that observe 99th percentile latency?
- tschellenbach 9y agoFor Stream's feed tech we also moved from Cassandra to an in-house solution on top of RocksDB. It's been a massive performance and maintenance improvement. This StackShare explains how Stream's stack works. It's based on Go, RocksDB and Raft: https://stackshare.io/stream/stream-and-go-news-feeds-for-over-300-million-end-users https://stackshare.io/stream/stream-and-go-news-feeds-for-ov...
- yazr 9y agoOr just try and benchmark Azul VM with pause-less GCs ?! (I have used Azul in low-latency production environments. It has pros and cons but it certainly beats re-writing the storage layer... )
- deleted 9y ago[deleted]
- truth_seeker 9y agoCurious to know the cons of using it, except being commercial.
- yazr 9y agoNeeds a stronger machine to be effective (more cores & more memory) Minor configuration issues (we had a very complex environment, custom kernel, weird network stuff, JNIs)
- manigandham 9y ago> except being commercial That's the biggest, especially for when it's a Facebook company. Otherwise it works well but can be pricey. The JVM is getting a new fully concurrent collector though called Shenandoah: https://www.google.com/search?q=shenandoah+gc https://www.google.com/search?q=shenandoah+gc
- domsj 9y agoNot just shenandoah[1], it's also getting zgc[2], so the low latency future looks bright for the jvm. 1. https://wiki.openjdk.java.net/display/shenandoah/Main https://wiki.openjdk.java.net/display/shenandoah/Main 2. https://wiki.openjdk.java.net/display/zgc/Main https://wiki.openjdk.java.net/display/zgc/Main
- steeve 9y agoWhy not use ScyllaDB ? (Serious)
- cnlwsu 9y agoAnswered a bit before but they have a team that knows c* well. Cassandra is proven to handle petabytes at scale in production systems.
- bfrog 9y agoMeanwhile scylladb looks like a better option for numerous reasons
- welder 9y agoGreat, now can you fix the Python Cassandra Driver to work in a multi-threaded application environment without the connection pooling bugs and default synchronous app-blocking (vs lazy-init) connection setup? https://github.com/datastax/python-driver https://github.com/datastax/python-driver
- ismail 9y agoSo question: Any thoughts on replacing HDFS + Yarn + Hive + HBASE with GulsterFS + Kubernetes + Cassandra ??
- ddorian43 9y agoHbase is sync+globally sorted, while cassandra is not, so probably not.
- xuanyue 9y agoIs there any trade off after replacing LSM tree-based storage engine to RocksDB storage engine?
- irfansharif 9y agoRocksDB is also an LSM structured KV store.