11 ms·
Badger – A fast key-value store written natively in Go
- zie 9y agoHow do I use it? there is zero docs on how to even get started, what the API is like(is there even one?) This could be worth all the peanuts in the world, but if the only way to learn how to use it, is read the Go source, very few people will bother.
- wiiittttt 9y agoThere is a link to the docs on the github page. https://godoc.org/github.com/dgraph-io/badger https://godoc.org/github.com/dgraph-io/badger
- lumost 9y agoGo projects typically use godoc to autogenerate documentation from comments, this has the added benefit that a single service can maintain updated docs for every open source go package in existence ex. https://godoc.org/ https://godoc.org/
- eternalban 9y agohttps://godoc.org/github.com/dgraph-io/badger/badger#example-package https://godoc.org/github.com/dgraph-io/badger/badger#example...
- zie 9y agoAH! I was looking in the docs/ dir, and found them to be completely missing. Thanks :)
- dkarapetyan 9y agoThere's that word again, natively. Should just say written in Go. The native part is redundant.
- slimsag 9y agoI believe the author is using the term 'natively' here to describe the project being written purely in Go, rather than using CGO (i.e. being a wrapper to some database written in C). I agree 'written purely in Go' would have been a better choice of words, though.
- piokuc 9y agoDoes it mean anything "written natively"?
- marcrosoft 9y agoWhy not boltdb?
- renke1 9y agoI was wondering too, but Badger doesn't offer transactions and stuff like that (on purpose). It seems to be more low-level in some regards. https://github.com/boltdb/bolt https://github.com/boltdb/bolt
- libeclipse 9y agoThey're based on different technologies. Bolt uses B-Trees and Badger/Rocks/Leveldb use LSM trees. Bolt also uses insane amounts of RAM and writes get slower and slower as the size of the database increases. (Personal experience, don't have benchmarks. Take this at face value.)
- nemo1618 9y agoWe've experienced these downsides as well, although it seems that boltdb's memory usage can be deceptive due to mmap'ing the db file.
- zimbatm 9y agoInfluxDB went from RocksDB (LSM) -> BoltDB (B+Tree) - > Custom (LSM again). Here is a pretty good writeup: https://docs.influxdata.com/influxdb/v0.9/concepts/storage_engine/ https://docs.influxdata.com/influxdb/v0.9/concepts/storage_e...
- rakoo 9y agoThere are basically two usecases for LSM trees vs B+trees: - Either you have a lot of random new writes and not so many updates, in which case LSM trees will ingest new data as fast as the disk can store them - Or you care more about read performance, and LMDB and BoltDB will have more predictable (and arguably better) performance InfluxDB is a timeseries database, so they write a lot of stuff and don't even read all of it. As data gets older it can be pruned efficiently with an LSM-based design, not so easily with B+tree. Dgraph on the other hand seems to sit right in the middle as it wants to be a general purpose database so there's no easy winner here. Hopefully the choice was correct for most use cases.
- fortytw2 9y agooff topic: flipping through the dgraph code, I noticed their licensing switch from Apache 2 to AGPLv3, anyone involved around to comment? Adding a draconic open source license is an unwise decision for an early stage database product imo (https://open.dgraph.io/licensing https://open.dgraph.io/licensing is a dead link)
- AsyncAwait 9y ago> Adding a draconic open source license is an unwise decision AGPLv3 is just a license, it is not "draconic", nor "cancer", (despite what Ballmer wants you to believe), it simply represents an ethical/moral agreement with the ideology of software freedom and since it is their code, they are free to express what they believe via the appropriate license. Nobody is making them do it, nobody is forcing you to use it and nobody is demanding you agree with it, it was a free choice and is therefore as far from "draconic" as one can possibly be, unless you believe that everybody who doesn't subscribe to your worldview is "draconic" by definition.
- tpush 9y agoI think what fortytw2 was trying to say is that the AGPL is not a wise choice for software that wants to gain the most popularity and usage as possible since usage of AGPL licensed software is categorically banned (even more so than GPLV3) by a some of companies. As an aside, one can certainly describe something as draconic(or whatever else) if one views it as such; it's just an opinion.
- sangnoir 9y ago> I think what fortytw2 was trying to say is that the AGPL is not a wise choice for software that wants to gain the most popularity and usage as possible Why should that be a laudable goal? Not all projects are megalomaniac.
- codebeaker 9y agoHaving just gone through a lengthy review process identifying a suitable license for our soon to be open-source software which has a commercial aspect, and selected AGPLv3, I'm very curious to know what companies have it "categorically banned". We did some research and didn't find that anyone had an issue with it. Whilst AGPL does open up come grey-areas which aren't as well understood as GPL the general reason to use it seems to be as part of a dual licensing scheme where companies with AGPL issues can simply purchase a non-transferrable limited MIT license or similar. Can you give me any more info on your sources?
- usegolang 9y agoHow does this compare to boltdb?
- libeclipse 9y agoSee https://news.ycombinator.com/item?id=14336771 https://news.ycombinator.com/item?id=14336771
- blaisio 9y agoThere is so much negativity in these comments! This project is really cool! I really appreciate the trend to rewrite C/C++ libraries in Go. It has always been really frustrating that hacky library wrappers for other languages leave performance and features on the table because they're either incomplete or just too hard to implement in the host language. For most of the languages out there, it is always better to have a native implemention. There are now a number of great embedded key/value stores for Go, which make it really easy to create simple, high performance stateful services that don't need the features provided by SQL.
- dkarapetyan 9y agoWas with you until you said you don't need SQL.
- deleted 9y ago[deleted]
- ckaygusu 9y agoFor really simplistic applications, omitting SQL is a fine choice, especially in golang you don't have the comfort of a well established ORM. There exists libraries that do provide an ORM, but since the declarative side of golang is limited compared to languages, say Python, I find them unwieldy. Also, the current trend shuns the usage of an ORM in golang and encourages directly or indirectly writing SQL queries and interacting with the database through database/sql and its extending libraries, like jmoiron/sqlx.
- dkarapetyan 9y agoFor really simple application omitting a k/v store is even better. Why isn't a simple map good enough?
- ckaygusu 9y agoDepends on the requirements (persistency, concurrency etc.) No offense, but it's bit pointless to discuss it further without deeper context.
- skyde 9y agoRange iteration latency is very important and might be limited by concurrency. I think you can only get 100K IOPS on Amazon’s i3.large when the disk Request queue is full. fio [1] can easily do this because it spawn a number of threads While working with Rocksdb we also found that Range iteration latency was very bad compared to a B+-tree and that RocksDB get good read performance mostly from random read because it's using bloomfilters. Does anyone know if this got fixed somehow recently? [1] https://linux.die.net/man/1/fio https://linux.die.net/man/1/fio
- mrjn 9y ago(Badger author) We have tried huge prefetch size, using one Goroutine for each key; hence 100K concurrent goroutines doing value prefetching. But, in practice, throughput stabilizes after a very small number of goroutines (like 10). I suspect it's the SSD read latency that's causing range iteration to be slow; unless we're dealing with some slowness inherent to Go. A good way to test it out would be to write fio in Go, simulate async behavior using Goroutines, and see if you can achieve the same throughput. If one would like to contribute to Badger, happy to help someone dig deeper in this direction.
- skyde 9y agoTo fill the queue on Linux goroutine wont be enough you would need to use libaio directly. sudo apt-get install libaio1 libaio-dev.
- mrjn 9y agoGo has no native support for aio. Based on this thread, Goroutines seem to do the same thing, via epolls. https://groups.google.com/forum/#!topic/golang-nuts/AQ8JOHxm9jA https://groups.google.com/forum/#!topic/golang-nuts/AQ8JOHxm... I think the best bet is to build a fio equivalent in Go (shouldn't take more than a couple of hours), and see if it can achieve the same throughput as fio itself. That can help figure out how slow is Go compared to using libaio directly via C.
- 9y ago
- libeclipse 9y agoThis could not have come at a better time, I've been looking for a fast, simple, pure Go key-value store. Few questions: - Does this have a log file? If so, what does it log and can it be disabled? - How is the data stored? (Single file, multiple files, etc.) - How is the RAM usage?
- mtrn 9y ago> I've been looking for a fast, simple, pure Go key-value store. Shameless plug: If you have write-once (or seldom) and read-often kind of access pattern for JSON documents, I wrote a simple (619 LOC), pure Go key-value store, that supports this use case: microblob[1]. It logs in common format, uses a single file backend and scales up and down with RAM.. [1] https://github.com/miku/microblob https://github.com/miku/microblob
- libeclipse 9y agoNot for this particular project, but I'll keep it in mind for if/when I ever need it!
- deleted 9y ago[deleted]
- thedatamonger 9y agonobody has said it so I must. Badgers? We don't need no stinking badgers!
- lucasmullens 9y agoWorth noting Badger here is a reference to the Wisconsin Badgers, since it's based on a paper from UW-Madison.
- mrjn 9y agoActually, no relation to Wisconsin Badgers. Badger is (and has been) Dgraph's mascot. So, we thought it would be nice to name the key-value store "Badger."
- reacharavindh 9y agoProbably a naive question, but how does it compare to Redis? When would someone look for a K-V store written in X instead of the already mature Redis?
- KAdot 9y agoBadger is designed to store data on disk (like RocksDB or LevelDB), while Redis is an in-memory storage which can't store data sets larger than memory.
- laumars 9y agoYou can store Redis on disk too. Not that you gain anything from doing so as persistence can be better achieved by clustering and I've never ran into issues where memory was a limiting factor, even with millions of records in Redis. Frankly if memory was a limiting factor then you'd probably want your KV store separate from your application anyway rather than the embedded approach that Badger takes. I think the real advantage of Badger is that it's not as sophisticated as Redis. ie you can have Redis-like functionality compiled into your application so less faffing about setting up another daemon / cloud micro-service inside your NAT / VPC / whatever.
- deathanatos 9y ago> persistence can be better achieved by clustering If you want persistence, then I'd recommend persisting to disk. While I've not had this fun with Redis, I've written code that took out an entire Cassandra ring. Were stuff only in memory, it would have not been pretty. Just because something is distributed doesn't mean it's guaranteed to never go completely down. (That said, if you're using Redis as an in-memory cache, this is a potentially acceptable tradeoff.)
- LaFolle 9y agoredis is not an embedded database, while rocksdb/boltdb/badger are.
- 9y ago
- Asmod4n 9y agoHow does it compare to LMDB?
- mjaniczek 9y agoNow let somebody run a QuickCheck on it, like they did with LevelDB: http://htmlpreview.github.io/?https://raw.github.com/strangeloop/lambdajam2013/master/slides/Norton-QuickCheck.html http://htmlpreview.github.io/?https://raw.github.com/strange...
- dmix 9y agoAll OSS projects should be tested with QuickCheck! I really hope this is a testing system that gets copied by other languages. I'm curious, is it a requirement to have types in order for QuickCheck to make sense? So you know what type of data to hammer a function with for example.
- hyperpape 9y agoNo, types are not necessary: http://hypothesis.works/articles/what-is-property-based-testing/ http://hypothesis.works/articles/what-is-property-based-test....
- lobster_johnson 9y agoQuickCheck is awesome, but is there a good Go equivalent? Anyone tried Gopter? https://github.com/leanovate/gopter https://github.com/leanovate/gopter
- faragon 9y agoFigures per minute? Same-process key-value should give millions per second operations (using one thread). And accessing via network, e.g. a la Redis, should be hundreds of thousands per second [1]. [1] https://redis.io/topics/benchmarks https://redis.io/topics/benchmarks
- mrjn 9y agoYou are comparing a memory first store against a disk first store. Everything is much faster if it only has to be stored in RAM before the update is considered successful.
- didip 9y agoThis is pretty exciting. I would love to see comparisons between badger and goleveldb
- mrjn 9y agoLevelDB should be slower than RocksDB.
- mirekrusin 9y agoSilly question - if SSDs or even motherboard had persistent storage of (couple of) 4 KB blocks where you'd be able to fsync < 4 KB (unfinished/non-yet-full pages) of data fast (DRAM + battery) - would that setup speed up writes in databases? It seems that databases often want to persist/flush data in unfinished (4KB or 8KB) pages when they're being built, once they are full, they don't change much - once full they could be normally persisted. Another kind of pages are those that change very frequently - ie. single "root" page which keeps counters or other stuff in single, root page. It seems a bit wasteful that multiple "checkpoints" (flushes/fsyncs) on partial blocks are triggering on hardware whole block rewrites. Similarly with "root"/"meta" pages that keep track of just few bytes frequently changing are triggering similar whole page rewrites. To be honest even some kind of PCI card with little battery and slot for DDR4 would probably do the trick, no? The rest could be implemented in software - as long as you'd have access to fast flushing with battery backed memory that survives hard crash - it should be fine. Is this silly idea?
- monocasa 9y agoA lot of higher end HBAs have battery backed DRAM write buffers for this reason.
- lumost 9y agoThere used to be a lot more innovation and variety in hardware following similar approaches. However, commodity gear was an easier deployment target for software and economics ensured commodity gear provided better performance per dollar than proprietary solutions.
- dis-sys 9y agoI am excited about the fact that the design is focused on SSD from day one (not just optimised for SSD). Wondering whether the authors have plan to optimise it for Intel Optane which has lower latency and much higher random IOPS at low QD. Currently I am using a cgo based RocksDB wrapper with the WAL stored on Intel Optane, the rest of the data is on a traditional SSD. It will also be great to have some comparison on write performance with fsync enabled. Overall, very interesting project, bookmarked, will definitely follow the development!
- iainmerrick 9y agoIt's interesting that the main motivation for this was that the Cgo interface to RocksDB wasn't good enough. I hate to turn this into a language war, but it's a big difference between Rust/Nim/etc where you just call C code more or less directly, and Go/Java/etc, which need a shim layer to bridge between C code and their language runtime. And possibly Go has the right approach! Is it better to make C integration as simple and smooth as possible, to leverage existing libraries, or is it better to encourage people to ditch all that unsafe C code and write everything in Go?
- openasocket 9y agoI've thought about this a bit, because I use a lot of Go at work. The reason calling C code from Go is mostly because Go has a different ABI: it does weird things with the stack, calling conventions, etc. But that doesn't stop you from calling assembly written to that ABI directly from Go. In fact, that is done a lot in the standard library. It is possible to create a wonky, C-like, low-level language that compiles into machine code with the Go ABI. Let's call this language Go-- :). With something like that, it might be possible to translate existing C code into Go--. However, the most likely use for Go-- would be for use in performance-critical sections of your Go code.
- iainmerrick 9y agoI'm skeptical that that would enable reuse of existing C code. If the workflow requires you to modify or annotate the C code, that's a lot of work (and risky) and you might as well just rewrite it in Go. If the workflow is totally automatic, so you can just use off-the-shelf C code, that's great! But in that case it's effectively just a C compiler, the "Go--" bit seems like a red herring. It reminds me a bit of "C--", Simon Peyton Jones' suggestion for a low-level target for Haskell and similar languages. It would remove some C functionality that language runtimes don't really need, and add some extra low-level stuff like register globals and tail calls. I don't think that got much traction, but it may have influenced the design of LLVM.
- udev 9y agoThe world wold benefit from some sort of map (or catalogue) of database engines and systems currently available. Anyone aware of such a thing?
- pkroll 9y agoA quick search shows there are such things, there's one on Wikipedia for instance [1]. But it's hard to assess everything in a single catalog. [1] https://en.wikipedia.org/wiki/Comparison_of_relational_database_management_systems https://en.wikipedia.org/wiki/Comparison_of_relational_datab...
- swedrupe 9y agoI threw a simple smoke test on it: walk a directory tree and store each regular file in the tree in the data storage thing. The directory tree wasn't anything radical - just 112MB, 260 keys/values, biggest one was 10MB. Then close the storage thing, open it again and see if we can retrieve the contents. No concurrency, no crash recovery, nothing. Sure, maybe not exactly a typical workload, but let's see how things behave at the extremes first. They do talk about big keys in the blog post. First impression. Easy to write the code. Second impression. Fast as hell. Third impression. Loses data. The last file in the set is consistently corrupted no matter how many times I try. I also tried on different directories. Exact same result everywhere. Last key/value written is truncated. Hmm. A data storage thing that fails to store data is maybe not exactly invoking feelings of trust, which should be the primary feelings you have about your data storage thing. Let's just throw it on something bigger. My /usr/lib. It's just 1.4GB. Shouldn't be too hard. [ 6546.783474] Killed process 4895 (main) total-vm:497120kB, anon-rss:368384kB, file-rss:0kB, shmem-rss:0kB Hmm. Do they just store everything in memory forever? Sure, this VM has stupidly little memory. But that's why we have disk. So that we don't have to keep everything in memory. But good. Now we get to test the crash recovery mechanism. [ 6823.635967] Killed process 4919 (main) total-vm:677964kB, anon-rss:367980kB, file-rss:0kB, shmem-rss:0kB Right. I guess there are reasons why experience taught me to let others be the first to use new data storage technology for a few years. In case someone wonders the test code is here: https://gist.github.com/art4711/9d781b8cf1f9f36df73a9c5c0403f703 https://gist.github.com/art4711/9d781b8cf1f9f36df73a9c5c0403... It may be all wrong, but considering that the "docs" directory contains 6MB of stuff just to display "hello world" twice, I had to write it based on what godoc said.