9 ms·
Pebble: A RocksDB Inspired Key-Value Store Written in Go
- chaosharmonic 6y agoWas anyone else deeply saddened about three words into the headline, on realizing this wasn't a watch? (RIP)
- asadlionpk 6y agoYes. I have moved on to Amazfit though.
- jsight 6y agoI definitely was. I still use one now, more than 3 years after their business failure. I really wish someone would make something like the Pebble Time 2.
- m-p-3 6y agoI have an Amazfit Bip, but the UI isn't as good as Pebble sadly. There is some work in making a similar OS called RebbleOS[1] currently ongoing. [1]: https://github.com/pebble-dev/RebbleOS https://github.com/pebble-dev/RebbleOS Hopefully it will be portable to other low-end smartwatches.
- Polylactic_acid 6y ago3 years later and it looks like they are just at the trying to get hardware features accessible. At this rate all pebble devices will have died / been discarded before its able to show a notification on your wrist.
- m-p-3 6y agoHopefully the OS could run on something newer than a Pebble, like the PineTime.
- jsight 6y agoI use rebble services, but I haven't seen much progress on the OS. I have slightly higher hopes for PineTime.
- Wowfunhappy 6y agoFrankly, I'm satisfied enough with my Pebble 2 that I'm not sure I care whether anyone makes a new one. I just hope I don't end up in a situation where it breaks and I can't get a replacement.
- jsight 6y agoUnfortunately, the battery will likely fail after 4-5 years.
- Jnr 6y agoI still use my Pebble Time Steel from the Kickstarter campaign. Thinking of switching to Apple watch because I have iPhone and I also want to use Apple Pay without the phone. If they had better battery life, I probably would have done it already.
- m-p-3 6y agoI'd still use my Time Steel if the battery didn't die, and I'm afraid of breaking it and its waterproofing if I attempt a replacement.
- jsight 6y agoTechnically you have nothing to lose.
- malisper 6y agoSo I understand the rationale for writing your own storage layer and think this is an awesome project, but there's something missing for me. One of the issues Peter brings up is they've come across a number of serious bugs in RocksDB. My question is, why would Pebble have less bugs. In fact, I would expect it to have significantly more bugs because Coackroach is the only company using Pebble. They mention briefly how they are going about randomized crash testing: > The random series of operations also includes a “restart” operation. When a “restart” operation is encountered, any data that has been written to the OS but not “synced” is discarded. Achieving this discard behavior was relatively straightforward because all filesystem operations in Pebble are performed through a filesystem interface. We merely had to add a new implementation of this interface which buffered unsynced data and discarded this buffered data when a “restart” occurred. but this seems to only scratch the surface of possibilities that can come up with a crash. For example, it's possible the filesystem had synced some of the buffered data to disk, but not all of it. There's no guarantee about what buffered data was synced to disk. All you know is that some, all, or none of it made it to disk. Bugs in this area are still regularly found in e.g. Postgres, so I'm having a hard time seeing how Coackroach is making sure Pebble doesn't have similar problems.
- d--b 6y agoWell, I think what they're saying is that they'd rather have bugs in code they've written than in code that is written by other people and in another language, and for which they don't control the patching pipeline. If RocksDB had had no bugs, they wouldn't have needed to write Pebble.
- danesparza 6y agoI'm sure 'not have to cross the cgo boundary' is significant when debugging, as well.
- pdpi 6y agoThat's an argument for them using it, but it's also basically arguing why nobody else should.
- dfee 6y agoAs a consumer, why would I want something like this written in Go vs. Rust? Is it just that Rust is really good with developer relations? Because it feels like to me that all new foundational technology is safer and faster in a language like Rust, and things written in Go should be higher up the food chain.
- hashamali 6y agoThey mention in the article that it's mostly due to familiarity with Go: > CockroachDB is primarily a Go code base, and the Cockroach Labs engineers have developed broad expertise in Go.
- marcrosoft 6y agoBecause you already have a huge Go code base and you want an embedded kv db for your project?
- psanford 6y agoThis is not a standalone DB. Its a key-value store implemented as a library. You would want this if you were a Go developer working on an application that needed a built in key-value store. If you were a Rust developer you'd want something similar written in Rust.
- ecnahc515 6y agoIn addition to what others mentioned, even if Cockroach is written in Go, they could have used Rust, but the trade-off is that they would need to use cgo which introduces extra complexity for building, debugging, and has performance trade-offs that a pure-Go based solution doesn't necessarily have.
- strken 6y agoReading this, I wonder how hard it would be to write or generate assembly bindings for libraries like rocksdb and sqlite, and whether you'd be able to do any better than cgo.
- 6y ago
- cube2222 6y agoHow does this compare to Badger[0], another similar in nature key-value store in Go? What were the trade-offs which made it necessary to create something new instead of adapting what exists? [0]: https://github.com/dgraph-io/badger https://github.com/dgraph-io/badger
- pella 6y agofrom the article: "A final alternative would be to use another storage engine, such as Badger or BoltDB (if we wanted to stick with Go). This alternative was not seriously considered for several reasons. These storage engines do not provide all the features we require, so we would have needed to make significant enhancements to them. ... Lastly, various RocksDB-isms have slipped into the CockroachDB code base, such as the use of the sstable format for sending snapshots of data between nodes. Removing these RocksDB-isms, or providing adapters, would either be a large engineering effort, or impose unacceptable performance overhead. " https://www.cockroachlabs.com/blog/pebble-rocksdb-kv-store/ https://www.cockroachlabs.com/blog/pebble-rocksdb-kv-store/
- ohnoesjmr 6y agoBadger is written by mad people from my point of view, who disabled issues on github, from my understanding declared it as "done" and "bug free", and any issue tracking is now done on the forum where the threads roll off to the void with no further trace.
- dilyevsky 6y agoWait really? That’s both hilarious and disappointing
- mrjn 6y agoWao. You describe us as “mad people” because we choose to not use GitHub issues? Is that all it takes to dismiss an open-source software and badmouth its authors? You have gone really low on this. All the issues have been ported over to Discourse. And no one has declared Badger, bug-free. I don’t know where you got that idea.
- 6y ago
- djhworld 6y agoReally enjoyed reading this, thanks. Would be interested to see if the garbage collector has presented any problems when running in production
- tyingq 6y agoThere's some notes on that here: https://github.com/cockroachdb/pebble/blob/c39589c8cb36d95df29b37e85ec2b4c3e20273dc/docs/memory.md https://github.com/cockroachdb/pebble/blob/c39589c8cb36d95df...
- petermattis 6y agoThe TLDR is that the GC did cause problems so we had to avoid it for the block cache. Luckily we were able to do so without exposing the complexity in the API. Not for the faint of heart. Don't try this at home kids.
- throwdbaaway 6y agoAny reason why the block cache needs to be 10s of GB in size? Cassandra, for example, usually has a rather small key cache on heap, and then just relies on the kernel page cache. I don't have experience with cockroach, it looks like the block cache is similar to cassandra row cache, which can be configured to be on heap or off heap, but usually not beneficial to enable.
- 2020-09-15-tmp 6y agoThought it would be worth mentioning Sled as an alternative to RocksDB for the Rust crowd: https://github.com/spacejam/sled https://github.com/spacejam/sled
- silasb 6y agoWorth mentioning that there is https://crates.io/crates/rocksdb https://crates.io/crates/rocksdb which Rust bindings to RocksDB.
- mholt 6y agoNot to be confused with Let's Encrypt's ACME client testing CA server project (the "scaled down" version of Boulder), Pebble: https://github.com/letsencrypt/pebble https://github.com/letsencrypt/pebble
- erichocean 6y agoThis makes total sense for Cockroach Labs, and I trust their engineering ability to get it right.
- willvarfar 6y agoI’ve run into serious house burning down problems with myrocks too. Simple recipe to crash MySQL in a way that is unrecoverable: do ALTER TABLE on a big table and it runs out of RAM, crashes, and refuses to restart, ever. Googling and people have been reporting the error on restarting several times on lists and things. What help is it to report to Maria dB or something? But do FB notice? Seems not. Here’s hoping someone at FB browses HN... I don’t get why FB don’t have some fuzzing and chaos monkey stress test to find easy stability bugs :(
- nitrobeast 6y agoLikely your db configuration is very different from what FB uses in production, so they have no incentive to investigate or fix.
- willvarfar 6y agoIts that the 'fix' is so unprofessional. The problem is that the program runs out of RAM. The challenge is to write the data and metadata in such a way that a program crashing at any point for any reason is recoverable. This is the basic promise of the Durability in ACID, and people using MyRocks expect it. Rather than actually making sure that MyRocks is durable, they simply slap on a 'max transaction rows' to make it unlikely you run out of RAM. Instead, you simply get an error, and can't do stuff like ALTER TABLE or UPDATE on large tables. Of course its easy to run out of RAM despite these thresholds, and its easy to find advice when you google the error messages you get that lead you to up the thresholds and to even set a 'bulk load' flag that disables various checks you probably haven't investigated. The whole approach is wrongheaded! A database that crashes should not be corrupt!!! Isn't this reliability 101? Why doesn't myrocks have chaos monkey stress testing etc etc? </exasperated ranting>
- nemothekid 6y ago>A database that crashes should not be corrupt!!! Isn't this reliability 101? Why doesn't myrocks have chaos monkey stress testing etc etc? Because Facebook has little incentive to ensure that RocksDB works well in your use case. MyRocks was built for Facebook and anything that Facebook doesn’t do probably isn’t particularly hardened. They aren’t going to invest time doing chaos monkey stress testing on codepaths they don’t use. Things like durability might not be super important to them because they will make it up in redundancy. I remember being burned by something similar during the early days of Cassandra. I’m sure Cockroach has hit the same bugs.
- AtlasBarfed 6y agoWhy would someone remove a non-GC database engine with a database engine with GC? Has Go evolved better low-GC features? As I understand Go GC vs JVM GC, Go avoids major GC by simply pushing it to the future and consuming memory more readily. But a database is a long-running program, so you have to pay the piper eventually.
- mateus_amin 6y agoI wonder the same thing. I have not programed a DB but.. I would think the worst thing to program in a gc'd language would be the cache. So, if you write that layer of the database separately in a non-gc'd you would avoid most of the headache. EDIT: I read the parent comment as why program a db in a gc language. I think the new wave 1st gen db's are doing alot of novel things. Minus the cache I imagine the velocity improvements form the 'simpler' language makes sense. Key value stores are become well trod ground however.
- FridgeSeal 6y agoI’d also be curious why they didn’t go with something like Foundation DB either.
- petermattis 6y agoPebble and FoundationDB are apples and oranges. Pebble is per-node KV storage engine. FoundationDB is a distributed multi-modal database. Internally, FoundationDB uses a library like Pebble for the per-node data storage. I think at one point it used SQLite. I'm not up to date on what it currently use. I seem to recall FoundationDB was writing their own btree-based node-level storage engine to replace the usage of SQLite. The equivalent of FoundationDB is present inside of CockroachDB: a distributed, replicated, transactional, KV layer. This is where a big chunk of CockroachDB value resides. This is where our use of Raft resides. Pebble lies underneath this.
- ryanworl 6y agoThe current production storage engine is an old-ish version of the SQLite btree. A new btree engine is being written now and is available but I don’t know if it is being used in production anywhere. RocksDB is shipping soon thanks to some work by members of the community.
- fis 6y agoThe name is a little close to this existing LevelDB fork, maybe consider a different name? https://github.com/utsaslab/pebblesdb https://github.com/utsaslab/pebblesdb
- petermattis 6y agoDamned for using a unique name (CockroachDB), and damned for using an innocuous one. PS PebblesDB was a research project and is dead as far as I know.
- jasonzemos 6y agoConcurrency and multithreading are a major focus of both Go and RocksDB. This introduction makes little mention of those areas, and I'm curious if there's any more to be said on this. The article lists several features being reimplemented, including: > Basic operations: Set, Get, Merge, Delete, Single Delete, Range Delete It makes no mention of RocksDB's MultiGet/MultiRead -- is CockroachDB/Pebble limited to query-at-a-time per thread? I'm genuinely curious how this all translates into Go's M:N coroutine model currently and moving forward with Pebble.
- petermattis 6y agoPebble does not currently implement MultiGet as CockroachDB did not use RocksDB's MultiGet operation. CockroachDB can use multiple nodes to process a query by decomposing SQL queries along data boundaries and shipping the query parts to be executed next to the data. CockroachDB can't directly use MultiGet because that API was not compatible with how CockroachDB reads keys. RocksDB MultiGet is interesting. Parallelism is achieved by using new IO interfaces (io_uring), not by using threads. That approach seems right to me. See https://github.com/facebook/rocksdb/wiki/MultiGet-Performance https://github.com/facebook/rocksdb/wiki/MultiGet-Performanc.... My understanding is that io_uring support is still a work in progress. We experimented at one point with using goroutines in Pebble to parallelize lookups, but doing so was strictly worse for performance. Experimenting with io_uring is something we'd like to do.
- jasonzemos 6y agoIndeed the conceptual fork point mentioned is RocksDB 6.2.1 which came before those features. The problem with RocksDB is that one thread only makes one request at a time. I should've phrased my question more succinctly: Is Pebble/CockroachDB capable of saturating the backplane with requests in parallel? Does it multiplex a single query by dispatching smaller requests to a thread-pool?
- petermattis 6y ago> Is Pebble/CockroachDB capable of saturating the backplane with requests in parallel? Yes. > Does it multiplex a single query by dispatching smaller requests to a thread-pool? Yes, though it depends on the query. Trivial queries (i.e. single-row lookups) are executed serially as that is the fastest way to execute them. Complex queries are decomposed along data boundaries and the query parts are executed in parallel next to where the data is located.
- nhumrich 6y ago> written in go Why does the implementation language matter for non-library tool? Is that its only selling point?
- johncolanduoni 6y agoIt’s not a network-connected key value store so you need to interact with it from Go. That makes a pretty big difference.
- deleted 6y ago[deleted]
- LaserToy 6y agoI hope you folks know what you are doing. If your screw it up you will have a lot of angry former customers, us including. Maybe less aggressive rollout strategy?
- StreamBright 6y agoFew question comes into mind reading this: - what is the plan to tackle Go's GC? It seems to me that above a certain scale people run into GC problems with Go.[1] - have they considered WickedDB?[2] It appears to be a good candidate for their need. https://github.com/Fullstop000/wickdb https://github.com/Fullstop000/wickdb https://blog.discord.com/why-discord-is-switching-from-go-to-rust-a190bbca2b1f https://blog.discord.com/why-discord-is-switching-from-go-to...
- 02020202 6y agohm, no word on performance comparison with badger, bolt, moss, pogreb, pudge...
- cristaloleg 6y agoSolving different problems, huh? but still similar.