11 ms·
Someone at my old company basically did this and put it into production. The first problem he encountered was that multiple connections couldn't both be using
by bvinc 6y ago
Someone at my old company basically did this and put it into production.
The first problem he encountered was that multiple connections couldn't both be using the database at a time without clobbering each other. "No problem," he thought, this is a good use case for micro services. A service sitting on top would ensure that there was only one operation being performed at a time.
Next, his problem was that the database would get corrupt sometimes when something bad happened in the middle of writing the file. His solution was to put the entire JSON format inside of a JSON string. If it could be parsed successfully, then he knew the whole file was written. Then all he needed were "backup" files for each table, in case the current one was corrupt.
Next, his problem was that querying and iterating through a large table performed badly, since it required parsing the entire thing first. Querying several times required the whole file to be parsed every time. The solution was to move SOME of the tables over to JSON-inside-SQLite.
EDIT: Oh yeah, the next problem was how to structure the data inside of sqlite. He decided to make a single table called "kitchen_sink" that held every JSON value. There was a column that said which "collection" it belonged to. There was another column that represented the row's primary key. So you could quickly query for a collection name, and a primary key, and get the full JSON row.
So the next problem was that you couldn't query quickly for things that weren't the primary key. So new columns had to be added called "opt_key1" and "opt_key2" where certain rows could put key values, and indexes could be added on those columns, so you could quickly query by it's first optional key, or it's second optional key.
- pgt 6y agoThis was...painful to read, but hilarious.
- Groxx 6y agoAFAICT (not being Node-fluent) this doesn't even use atomic file writing strategies :| So yeah, all of these are pretty likely to happen with this lib. Just use SQLite, people. Even JSON-in-SQLite is still likely to be an improvement.
- cortesoft 6y agoBut.... why?
- felipellrocha 6y agoJob security, maybe? Can't think of anything else... This person would certainly never be fireable again after merging that monstrosity
- waltpad 6y agoMaybe because it was tempting: JSON is fairly easy to handle, very portable, and when you look at a JSON document, it's straightforward to think about querying it, and thus DB, although JSON is structured, and DBs are relational.
- nailer 6y ago> although JSON is structured, and DBs are relational. You can have structured relational databases (RethinkDB for one, but there are others)
- jamil7 6y agoFunny read, why did nobody stop him?
- bvinc 6y agoI tried. I held a meeting to talk about the code. I found the problems hard to predict and hard to describe. It was decided that after the meeting he would work more on making his code less hacky and more production ready. But the real answer is that our team was very siloed. No one knew what anyone else was doing. The other problem was that he was actually solving real world problems, and he was a very high performer. He got stuff done. Arguing to start over a project that's already working is a difficult position to hold when talking to management.
- musingsole 6y agoI can sympathize but it seems hard to argue with this developer's approach then. If it met the needs of the company, particularly to the desired level at the times these features were requested, I don't think there's a valid critique of the developer's architecture beyond iT's NoT DoNe CoRrEcTlY. And still, there's a lot to be said for keeping your developer's entertained so they stick around.
- airstrike 6y ago> there's a lot to be said for keeping your developer's entertained so they stick around Really? At the expense of everyone else who has to deal with this monstrosity for the foreseeable future, or worse yet replace it with an actual tool that can be reliably used. This JSON-inside-sqlite-inside-JSON-inside-a-JSON-string beast should never have seen the light of day. You're not paid to be entertained, sorry. You're paid to be productive. As productive as you can, and to put the needs of the client and the long-term success of the company hopefully first but certainly before any resemblance of entertainment if you're getting paid Did I mention you're getting paid to work?
- tegiddrone 6y agoOn the other hand, other sorts of devs would probably not be entertained having to maintain a in-house database some unchained dev decided to introduce into the stack one day for "reasons." That tech debt will compound until it becomes more of a liability.. hopefully the product brings in enough money so that the in-house database can continue to be supported or removed. This sort of stuff is what deters me from being a developer sometimes. Fuck the salary, get me out of here.
- thesid 6y agopure genius
- ak39 6y agoBut what color was it? Mauve?
- boynamedsue 6y agoIt's easy to get a little laugh from this, but congrats to the guy for exploring. Now he knows first-hand the inordinate challenge and can describe it in detail, but more importantly avoid these hard-learned patterns later.
- teej 6y agoGreat for the dev, terrible for everyone else who had to deal with it -in production-.
- xienze 6y ago> but more importantly avoid these hard-learned patterns later. Depends, I’ve known people who have gone through similar experiences and still poo-poo all those “unnecessarily bloated” solutions like a proper database.
- colecut 6y agoI make crappy thrown together frontends and will probably forever poo-poo the 'unnecessarily bloated' js frameworks
- burgerboy 6y agoThen you are a bad dev with no critical thoughts
- meddlepal 6y agoI worked with someone that was keen to use reductionist logic and arguments... we ended up with a lot of shitty solutions to problems that were hard to maintain, hard to extend, and hard to use because the more "complex" solution was really just a fancy version of a folder and some text files.
- theturtletalks 6y agoHe took "What I cannot create, I do not understand" to a whole new level.
- 6y ago
- benologist 6y agoI have built this into some software and the reason is when you are developing, and for very limited use cases outside of development it is extremely convenient. Clone my project and start messing around without provisioning a database. It reduced dependencies by hundreds of modules too because all of the connectivity libraries were abstracted into separate modules you'd only install if you wanted that particular type in production.
- hnlmorg 6y agoThis is why I love sqlite. No provisioning required and you have a flat portable file. But you also have the added bonus that it's highly performant and there isn't much work to refactor your SQL from sqlite3 to most other RDBMS.
- bifrost 6y agoI think this is my favorite comment today.
- treeman79 6y agoMany many years ago In college, SQL sounded hard. So I built my own database in PHP. Enough said
- time0ut 6y agoWhen I was in university, one of our major projects was to implement a rdbms. Super fun project that taught us a lot of respect.
- blackrock 6y agoHow did you implement the file block sectors? Did you build your own sub FAT protocol? Like, initially allocating a large file block, and then you subdivided that yourself to get individual sector access?
- imtringued 6y agoAs long as the stakes are low everything is fine.
- btown 6y ago> So the next problem was that you couldn't query quickly for things that weren't the primary key. So new columns had to be added called "opt_key1" and "opt_key2" where certain rows could put key values, and indexes could be added on those columns, so you could quickly query by it's first optional key, or it's second optional key. It's all fun and games until you realize DynamoDB works more or less the same way: https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/bp-modeling-nosql-B.html https://docs.aws.amazon.com/amazondynamodb/latest/developerg...
- waheoo 6y ago"webscale"
- btown 6y agoWhich reminds me: it's almost the 10 year anniversary of "Mongo DB Is Web Scale" https://www.youtube.com/watch?v=b2F-DItXtZs https://www.youtube.com/watch?v=b2F-DItXtZs
- methodin 6y agoA lot of things are hard to do in DynamoDB but for the scenarios it excels at there are not many peers.
- idclip 6y agoBless him. I love it
- kaiby 6y agoAs others have mentioned, there are a ton of off-the-shelf solutions that would have been more than adequate for this. My question is, why didn't he go for any of the existing solutions when setting them up would've still been faster than rolling his own DB-in-a-JSON-file solution?
- liuliu 6y agoIt all sounds funny but the final solution is not far away except the JSON bit. FriendFeed 11 years ago first popularized the concept of storing schema-less data in MySQL: https://news.ycombinator.com/item?id=496946 https://news.ycombinator.com/item?id=496946. Uber did the similar thing a few years ago: https://news.ycombinator.com/item?id=16251143 https://news.ycombinator.com/item?id=16251143 I've been building an open-source alternative on mobile that based on similar concept (SQLite + FlatBuffers): https://dflat.io/ https://dflat.io/ SQLite own schema is already awesome, but in this way, you can have sum-types, better schema upgrade guarantees, index building can be asynchronously etc.
- badsectoracula 6y ago> Next, his problem was that the database would get corrupt sometimes when something bad happened in the middle of writing the file. I'm not sure i understand how this can happen... unless you try to update JSON in-place (which is a very bad idea for any text-based format), what you do is encode/write the entire JSON from scratch. So either the file is written properly or it isn't written. Honestly from the entire message it doesn't sound like JSON was a bad idea but that your coworker didn't know what he was doing and if he was doing something else then he'd still be doing big mistakes.
- oxymoron 6y ago1. A cosmic ray storm turned all ASCII charactees into ECBDIC 2. Lightning struck the 12 V feed and upped the voltage to 10 MV, turning all 0 and 1’s into 6’s 3. Someone spilled a New England Pale Ale on the server 4. The process was assinated by the mysterious killer only known from his modus operandi of leaving OOM written in blood across the syslog 5. Birds nested within the server and fed all the SATA cables to their babies Seriously though, disk writes aren’t atomic.
- lalaland1125 6y agoFile renames are atomic. This is a solved problem: 1. Write your updates to a copy of the file. 2. Do an atomic rename of that copy to the original.
- CapsAdmin 6y agoHow did you go about porting the database code to something more sane? (just assuming you did) I imagine if this database system is contained well enough, it shouldn't be so difficult to swap its internals with something else. Especially if it's all just JSON-like.
- StavrosK 6y agoI wrote something that works more or less the same way, but exclusively over SQLite: https://github.com/skorokithakis/goatfish/ https://github.com/skorokithakis/goatfish/ It's actually quite good as a quick-and-dirty datastore. I do need to move everything to SQLite's JSON field, though.
- deleted 6y ago[deleted]
- louis8799 6y agoOh yeah, so basically had he pushed further, he would have realized to avoid corruption, he would need to implement "write ahead log (WAL)". An for the implementation of WAL and other performance concern, he would have realized that storing the JSON as string is not the way to go, he'd need to implement other binary data structure. Then he'd have realized that he had just invented another NoSQL DB. Had he pushed further..... Had he pushed further, he'd have raised funding for the newly invented NoSQL DB, and built a startup company on top of it.
- itsthecourier 6y agoIn every paragraph I'm expecting CTO/engineering lead/senior dev to appear and punch the guy. :(
- golergka 6y agoHow was he hired for that position?
- goatinaboat 6y agoSomeone at my old company basically did this and put it into production. The is the mentality that plagues the industry, that anything more than a few years old is obsolete, and therefore experience is worthless, and therefore the wheel must be reinvented every time because those old programmers must have been dumb, why would they use SQL otherwise. Why real engineers don't take "software engineers" very seriously (and in turn why software engineers don't take webdevs seriously).
- bullen 6y agoThe solution is to put every "entity" in it's own JSON file and use the file system indexing with paths. Only downside is you need to format ext4 with type small otherwise you run out of inodes before you run out of disk!
- inshadows 6y ago> Next, his problem was that the database would get corrupt sometimes when something bad happened in the middle of writing the file. His solution was to ... https://www.cs.ait.ac.th/~on/O/oreilly/perl/cookbook/ch07_09.htm https://www.cs.ait.ac.th/~on/O/oreilly/perl/cookbook/ch07_09...
- mpweiher 6y agoI used sets of flat JSON files as our "database" in the Wunderlist iOS and macOS clients. Worked like a charm, never had a problem with it. It was actually put in as a placeholder until we had time to think about a real storage solution, but it turned out we never needed anything more sophisticated, and were actually the fastest and most reliable clients we had. In fact, every time I encountered a performance problem I was hopeful that I would finally have a good reason to do that real implementation, but it invariably turned out to be a simple bug. - Cocoa has -writeToFile:atomically:, which writes a new file and then renames, so no write-corruption - We were lucky that lists had just the right granularity for a single file to be read/written atomically - We likely wrote (quite) a bit more data than absolutely necessary, but I/O tends to have large fixed overheads so medium files tend to take around the same time as small files - We did not do anything with the data on disk except read it, so not a DB - We really did use files, not JSON strings inside SQLite - We flushed to disk asynchronously, but as quickly as possible