5 ms·
BoltDB author here. Yes, it is a bad design. The project was never intended to go to production but rather it was a port of LMDB so I could understand the inter
by benbjohnson 5y ago
BoltDB author here. Yes, it is a bad design. The project was never intended to go to production but rather it was a port of LMDB so I could understand the internals. I simplified the freelist handling since it was a toy project. At Shopify, we had some serious issues at the time (~2014) with either LMDB or the Go driver that we couldn't resolve after several months so we swapped out for Bolt. And alas, my poor design stuck around.
LMDB uses a regular bucket for the freelist whereas Bolt simply saved the list as an array. It simplified the logic quite a bit and generally didn't cause a problem for most use cases. It only became an issue when someone wrote a ton of data and then deleted it and never used it again. Roblox reported having 4GB of free pages which translated into a giant array of 4-byte page numbers.
- otterley 5y agoI, for one, appreciate you owning this. It takes humility and strength of character to admit one's errors. And Heaven knows we all make them, large and small.
- klabb3 5y agoI also appreciate the honesty, but I don't see the error in the author, quite the opposite. Afaiu, Bolt is a personal OSS project, github repo is archived with last commit 4 years ago, and the first thing you see in the readme is the "author no longer has time nor energy to continue". Commercial cash cows like Roblox (a) shouldn't expect free labor and (b) should be wise enough to recognize tech debt or immaturity in their dependencies. Heck, even as a solo dev I review every direct dependency I take on, at least to a minimal level. I can't speak to the incident response as I'm not an sre, but as a dev this screams of fragile "ship fast" culture, despite all the back patting in the post. I'm all for blameless postmortems, but a culture of rigor is a collective property worthy of attention and criticism.
- travisd 5y agoThe onus is more on HashiCorp here by this logic. Consul itself is open source but HashiCorp sells an enterprise version.
- otterley 5y agoConsul is much older than 4 years old (public availability in 2014; 1.0 release in 2017, with a lot of sites using 0.x in production long before). And the fact that they didn't encounter this pathological case until Q4 2021 tells us that they got a lot of useful life out of BoltDB. They also were planning to switch over to bbolt back in 2020[1]. The developers at Hashicorp are top-tier, and this doesn't substantially change their reputation in my eyes. Hindsight is always 20/20. Let's end this thread; blaming doesn't help anyone. [1] https://github.com/hashicorp/consul/issues/8442 https://github.com/hashicorp/consul/issues/8442
- benbjohnson 5y agoI think the design choice is mine to own but, as with most OSS software, liability rests on the end user. It always sucks to see a bug cause so much grief to other folks. As for HashiCorp, they're an awesome group of folks. There are few developers I esteem higher than their CTO, Armond Dadger. Wicked smart guy. That all being said, there's a lot of moving parts and sometimes bugs get through. ¯\_(ツ)_/¯
- dtheodor 5y agoI share the sentiment, but not for Roblox. Hashicorp, with a recent IPO, 200 mil operating revenue, and supposedly a good engineering reputation has one of its flagship products critically depend on a "toy project".
- tacLog 5y ago> BoltDB author here. How does this happen so often? It's awesome to get the authors take on things. Also thank you for explaining and owning it. Where you part of this incident response?
- benbjohnson 5y agoIt's on the front page of HN so it's pretty visible. However, I also use f5bot to notify on terms like "boltdb" and my other project "litestream".
- probotect0r 5y agoYou also made litestream?! Awesome, I love that project.
- benbjohnson 5y agoYeah, that's me too. Hopefully I don't crash another multi-billion dollar public company in 8 years with it though... :)
- simonw 5y agoSounds like pretty good success criteria to me!
- sjg007 5y ago> The project was never intended to go to production :)
- coldcode 5y agoHaving written a commercial memory allocator a quarter century ago, I remember dealing with freelists, and decided they were too much of a pain to manage if fragmentation got out of control. I chose a different architecture that was less fragile under load. Interesting that this can still be an issue even on today's hardware. It's also interesting how much a tiny detail can derail a huge organization. My former employer lost all services worldwide because of a single incorrect routing in a DNS server.
- chrislusf 5y agoYour answer should be voted to the top! :) OSS contributors are rarely noticed or appreciated. Did HashiCorp ever sponsor you or share any revenue with you? The OSS ecosystem is broken.
- benbjohnson 5y agoI had a few folks offer to sponsor at individual levels but no corporate sponsorship except Shopify paying my salary in the early days. That being said, I gave away the software for free so I don’t have any expectation of payment. I agree the ecosystem is broken but I don’t know how to fix it or even if it can be fixed.
- dottedmag 5y agoBy any chance could you (or bbolt folks?) update README to include this information?
- buchanmilne 5y ago> At Shopify, we had some serious issues at the time (~2014) with either LMDB or the Go driver that we couldn't resolve after several months Is there an issue/bug for this somewhere?
- benbjohnson 5y agoI can’t remember off the top of my head. We had an issue where every couple of months the database would quickly grow to consume the entire disk. We checked that there were no long running read transactions and did some other debugging but couldn’t figure it out. Swapping out for Bolt seemed to fix that issue. I haven’t heard of anyone else with the same issue since then so I assume it’s probably fixed.
- erthink 5y agoThese issues partially solved in the libmdbx (a deeply revised and extended descendant of LMDB). So BoltDB and LMDB affected users may switch to libmdbx as the Erigon (Ethereum implementation) does year ago https://github.com/ledgerwatch/erigon/wiki/Criteria-for-transitioning-from-Alpha-to-Beta#switch-from-lmdb-to-mdbx https://github.com/ledgerwatch/erigon/wiki/Criteria-for-tran... For now this is (relatively) easy since bindings for GoLang, Rust NodeJS/Deno, etc are available and the API is mostly the same in general. --- The ideas that MDBX uses to solve these issues are simple: zero-cost micro-defragmentation, coalescing short GC/freelist records, chunking too long GC/freelist records, LIFO for GC/freelist reclaiming, etc. Many of the ideas mentioned seems simple to implement in BoldDB. However the complete solution is not documented and too complicated (in accordance with the traditions inherited from LMDB ;)