6 ms·
How Do I Cassandra?
- rbranson 15y agoFYI -- the presentation says "temporal data" is a bad use case for Cassandra, but that's a misuse of the word temporal. "Ephemeral data" is a better way to say that.
- dpritchett 15y agoAttended Rick's presentation last night. It was so choice. If you have the means, I highly recommend taking one in. The md5 hashing to shards around a keyspace ring and the read/write quora were particularly interesting. Definitely going to be browsing the Reddit source (https://github.com/reddit/reddit https://github.com/reddit/reddit) for ideas on how to use Cassandra.
- pgr0ss 15y agoYou should read the Dynamo paper: http://www.allthingsdistributed.com/2007/10/amazons_dynamo.html http://www.allthingsdistributed.com/2007/10/amazons_dynamo.h... The ring is first explained in Section 4.2 Partitioning Algorithm.
- deweller 15y agoIs there an audio or video archive of this presentation available on the web perchance?
- dpritchett 15y agoI didn't see any recording equipment last night, sorry. It's a shame too - Rick's presentation style and the Q&A added a lot. You can try to pin him down in #cassandra on FreeNode if you like.
- thomaslangston 15y agoNo. Unfortunately audio/video recording equipment seems to be sparse in the Memphis user group communities. As far as I know we only have one group that does it regularly. http://www.justin.tv/launchmemphis http://www.justin.tv/launchmemphis If anyone has a good, cheap solution for recording/streaming user group presentations, please share.
- mikeryan 15y ago<< META COMMENT >> I'm always surprised when someone's raw slide decks, for an interesting presentation make it to the front page of HN. I get so little from just an 85 slide deck, especially decks with a slide like DATA MODEL I gotcha column family ratttt heeeeyyyaahhh!
- forgotAgain 15y agoWhat works as a nice intro with audio doesn't always work standalone. The first dozen slide don't match the title. Live, it can be a way to warm up a crowd but with just the deck it's a drag.
- rbranson 15y agoYup. Slides are best used as an illustrative tool.
- checker 15y agoI got a lot out of your slides. Sometimes it can be difficult to discern the point of the slide without audio, but I felt that your slides communicated the point pretty well. Thanks for sharing!
- slyall 15y agoI read a lot of slide decks. I go to very few conferences (live in NZ) and videos of talks are often not available or of poor quality (camera stays on presenter or slides obscured). Also I can go through a slide deck in a few minutes while a 1 hour presentation takes an hour to view the whole thing and a significant amount of time for me to decide if it is uninteresting. Obviously you don't get much out of presetations that are just 20 flickr photos and 10 words but for decks like this you can learn a bit. Since many presentations are "read" by more people than "watched" I'd hope that authors will pay more attention into improving the readability of their presentations (perhaps by releasing more wordy versions).
- nowarninglabel 15y agoI'm used to attending and giving presentations that are a bit outlandish, but this is gimmicky enough that I started losing track of actually learning anything out of this. I'm sure it's better in-person or with a recording though.
- bkmontgomery 15y agoI was at the talk, and it was very good in person... I really wish we'd recorded it :(
- rvenugopal 15y agoHow does the ring range work with nodes going down and new nodes being introduced into the cluster. As per slide 38, it appears to be a static range.
- rbranson 15y agoThe ranges are assigned by a per-node token. With 4 nodes like the cluster shown in the slides, you'd space the tokens apart appropriately for 25% of the range on each node. When you add a node, you need to decide if you want to rebalance the cluster by moving every node's token a bit so that every node's range consumes 20% of the keyspace, or if you just want to quickly alleviate load by splitting a single node's traffic in half by assigning the new node's token half way into an existing node's range. Decommissioning a node works the same way, but in reverse. More detailed info: http://www.datastax.com/docs/1.0/operations/cluster_management#adding-capacity-to-an-existing-cluster http://www.datastax.com/docs/1.0/operations/cluster_manageme...
- koobe 15y agoDoes Consistency in Cassandra require client clocks to be in perfect sync? Y u no vector clocks?
- rbranson 15y agoTimestamps were chosen over vector clocks intentionally. "...the primary use case for classic vector clocks of merging non-conflicting updates to different fields w/in a value is already handled by cassandra breaking a row into columns." See https://issues.apache.org/jira/browse/CASSANDRA-580 https://issues.apache.org/jira/browse/CASSANDRA-580
- thomaslangston 15y agoIf I remember correctly from the talk (I was one of the lucky attendees), Rick claimed basically it is not an issue. Standard time synchronization methods are sufficient for the problem domains Cassandra is meant to address.
- idan 15y agoY U NO SPEAKERDECK? Seriously, it's a much more pleasant slide-consuming experience for your great slides.
- grout 15y agoYou don't, if you're smart. http://chip.typepad.com/weblog/2011/08/why-cassandra-is-unfit-for-production.html http://chip.typepad.com/weblog/2011/08/why-cassandra-is-unfi...
- jbellis 15y agoThat boils down to Chip's opinion that master/slave is the One True Way to do replication, despite its many demonstrated disadvantages. I see your Chip Salzenberg and raise you a Werner Vogels: http://www.allthingsdistributed.com/2008/12/eventually_consistent.html http://www.allthingsdistributed.com/2008/12/eventually_consi...
- grout 15y agosigh And here your PR guy wanted to fly me to New York to talk to you. Obviously the effort would have been wasted. Replication events are too important to be designed to fall on the floor. I want at any given time to know approximately how far behind replication is; I want primary write events to block until replication is at least durably scheduled; and if I pause or slow operation and wait for replication to catch up, I want to know that it _is_ caught up without holes. Cassandra fails utterly to meet these minimal requirements. Master/slave queue is not the only way to meet these needs, but unless a replacement can fulfill the requirements, I can't responsibly switch.
- jbellis 15y agoWhen I'm learning about a new architecture, I like to take the position of, "let's assume the authors aren't idiots. If they're not, why would they have designed things this way?" With that in mind, let me pose this question to you. Is using a Dynamo architecture (http://www.allthingsdistributed.com/2007/10/amazons_dynamo.html http://www.allthingsdistributed.com/2007/10/amazons_dynamo.h...) for S3 "irresponsible" of Amazon? I submit that by this point, Amazon (among others) has convincingly demonstrated that this approach can indeed achieve a high degree of reliability. If you agree, then I suggest that you read through the Dynamo and Eventually Consistent papers again with the "let's assume these people aren't idiots" approach, and see if you can spot what this architecture offers to achieve a similar goal to your "wait for replication to catch up" design.
- ajtaylor 15y agoFinally, I get it! This go around, the ideas of ColumnFamilies stuck. Thanks for the slides. Any chance the video is available to go along with it?
- deleted 15y ago[deleted]