6 ms·
ToroDB – Document-oriented JSON database on top of PostgreSQL
- arthursilva 12y agoTo the author: you should probably use VODKA index.
- jeltz 12y agoVODKA indexes are still experimental and exist only as a branch in Alexanders own repo (https://github.com/akorotkov/postgres/tree/vodka https://github.com/akorotkov/postgres/tree/vodka). I do not think they are anywhere near production ready, and even if they are they are not an extension. Alexander is working on allowing new index access methods (the type of index: e.g. hash, btree, gin, gist) to be added in extensions.
- arthursilva 12y agoYep. And that's why I believe it's a great way to test/improve VODKA. PS: now I'm not sure if it applies considering that the data is actually normalized inside Postgres.
- wh-uws 12y agoReally interesting pdf of slides from the PGCcon talk about this here http://www.pgcon.org/2014/schedule/events/696.en.html http://www.pgcon.org/2014/schedule/events/696.en.html
- jakozaur 12y agoSuper interesting. Though would be interested how it works. Is it using JSON/hstore underneath? Right now reading the code seems to be the only option to analyze it.
- ahachete 12y agoNo, it's not using json/jsonb/hstore (to store the true JSON data). It uses a bit of jsonb but for a side part of the information. And this is the true power of it, that data is stored in normal, relational tables. Please see a comment above explaining this in more detail :)
- drawkbox 12y agoDoes all data have to be stored relationally or can you pick and choose? Reason I ask is for example something like a game session blob from a multiplayer game, I might not want that to be merged with the relational data and just throw it away after a certain amount of time/days/weeks so they are nice as blobs but other data would be better relational. Yes there are app/cache level ways around this but might be cool if there was a way to choose auto relational or blobby. I guess ultimately you could just have multiple stores for live and archival data but something to think about.
- ahachete 12y agoData is laid out transparently to the user. Indeed, there should be no (visible) difference as to how it is stored. Data within a single "level" is stored on a single table. If you want to keep blobs, you can definitely do that, they will be stored as such (bytea in postgres).
- ocdnix 12y ago"JSON documents are stored relationally, not as a blob/jsonb." How is the transformation designed, to go from a structured document to a flat set of tables, akin to what an object-relational mapper would do?
- zrail 12y agoThis is my question. How is the data actually broken down into tables/columns? I can't find a schema anywhere in the source, but maybe I just don't know where to look. Also the code talks an awful lot about DB cursors, which indicates that this is not really taking advantage of either SQL or the relational model at all.
- zapov 12y agoYou might find Revenj (https://github.com/ngs-doo/revenj https://github.com/ngs-doo/revenj) interesting too. It maps DSL schema to various Postgres objects and offer 1:1 mapping between view in the DB and incoming/outgoing Json.
- ahachete 12y agozrail, it's difficult to find the schema and/or tables in the source code as all of them are 100% dynamic. See my coments above about the internal workings of ToroDB. Or give it a try, run it and check how data is laid out! Regarding the use of cursors, they are absolutely necessary, as there is -in MongoDB- the concept of a session, and queries may be asked in subsequent packets to return the next results. However, I don't see how this impedes to take advantages of the relational model. ToroDB definitely does that, if you look at the created tables schema.
- Demiurge 12y agoI would also be very interested in how this kind of a breakdown occurs. It is not uncommon I have to store arbitrary structure, and if there is a better way than a blob, which can be abstracted away by an app, that would be quite useful on its own, without the MongoDB protocol.
- ahachete 12y ago
- jmasonherr 12y agoSo can I use it with meteor?
- justinsb 12y agoI would be surprised if ToroDB implemented the oplog which Meteor uses for efficient live-updating, but I guess it _should_ work with the polling mode. SQL support (i.e. direct Postgres support) is on the Meteor roadmap for post-1.0.
- ahachete 12y agoToroDB is definitely going to have MongoDB's oplog, so it should work as is for live-updating. SQL support for ToroDB... there's surely room for it in the future ;)
- ahachete 12y agoIt hasn't been tested, but as soon as the full Mongo protocol is supported, it should work exactly the same as Mongo does. Give it a try and let us know! :)
- e1ven 12y agoAnother project worth mentioning is mongolike - It uses PLV8 to implement the Mongodb interface ontop of Postgres. https://github.com/JerrySievert/mongolike https://github.com/JerrySievert/mongolike
- jeffdavis 12y agoAnd don't forget Mongres: https://github.com/umitanuki/mongres https://github.com/umitanuki/mongres
- tracker1 12y agoThis is really cool, thanks for pointing it out... I'd, personally prefer that some of these implementations allowed for both a mongo wire protocol compatibility, as well as being able to use a similar interface within plv8 ... that would be awesome.
- sandGorgon 12y agothese projects are so cool ! I wish all of them join hands to build a single Mongo emulation layer on Postgres. And drop in compatibility with MongoMapper
- michaelchisari 12y agoAs someone who has been coding web apps for a very long time, it feels odd to say this, the PostgreSQL community is doing some very interesting things. I'm excited to see where things go.
- jsherer 12y agoGenuinely curious...why does it feel odd to say that?
- chaostheory 12y agoI second that. Postgres has been doing interesting things for years now.
- michaelchisari 12y agoWhen I first started web development (mid/late 90's), Postgres was the very conservative open source choice. MySQL was fast, good enough, and broke a few big rules. That hasn't been the case for a while, but with the recent speed improvements in the past few years, prompted by NoSQL, and MySQL's troubles with Oracle, Postgres has become the Open Source SQL frontrunner, and everybody's taking it in really interesting directions.
- tracker1 12y agoI'm hoping the JSON & plv8 support gets much more in the box, and that the replication setup becomes more baked in as well. Right now, I'm migrating from MS-SQL to a MongoDB replica set for most of our core data. Mainly because the failover works better than most, and the licensing costs for even MS-SQL with replication and failover are budget blowing. Most of the PostgreSQL options for replication are really bolted on with some serious drawbacks, and automatic failover to a new master is another issue. I've always liked PostgreSQL (and Firebird for what that's worth)... I do hope that some of these features become more of a checkbox item during installation, and less of a bolted on, have to compile and dive into the deep to get them, and even then only have it half baked.
- pvh 12y agoCool concept. I'm not a huge fan of normalizing out the underlying data into pseudorelational tables. Reconstructing trees in Postgres is a little gnarly on larger tables but it's definitely neat to see more people trying to combine the user experience of Mongo with the technology advantages of Postgres. What inspired the creation of ToroDB? Are you using it in production yourselves in a limited way?
- ahachete 12y agoThanks! Documents are split into chunks before hitting the database, so there is no need to reconstruct trees in PostgreSQL. Indeed, many queries don't need to reconstruct the (whole) tree, just a part of it or even just one level. However, in doing so, ToroDB is able to offer queries that only need so scan a small subset of the database (compared to the whole database) to query your data. ToroDB was inspired by the DRY principle: relational databases like PostgreSQL are already good enough that with some tweaking may perfectly well as a document-store.
- platform 12y agoas far as I understand it, shredding hiearchical documents into a relationa store is optional feature of Oracle XMLDB http://www.oracle.com/technetwork/database-features/xmldb/xmlchoosestorage-v1-132078.pdf http://www.oracle.com/technetwork/database-features/xmldb/xm... It is called XMLDB structured storage (vs binary storage, that one actual stores the hierarchical XML naitively) The XMLDB structured storage I think has been available since 2003 (but I might be wrong there). Oracle's XMLDB structure requires XSD declaration. Within that declaration XMLDB can use 'hints' to indicate if a given set of attributes across documents is related. And if yes, it will shred the docs in a way that the related attributes will be co-joined. The advantage, as the author of ToroDB noted below is a) space saving b) ability to use relational joins that are using disk-optimized access strategies. The disadvantage (at least in Oracle XMLDB structured option) -- is the need for declaration of the model ahead of time, and joins for deeply/complex structured documents. It looks like ToroDB 'senses' the model of each document on the way in. I think there are definetely use cases for this approach, that tolerate the trade off between ingestion speed and storage control. Plus using a relational engine underneath allows for ACID properties (eg multi-object rollback/commmits) -- which native mongo does not provide.
- ahachete 12y agoplatform, very interesting the XMLDB info. As you point out, ToroDB does not require any model declaration ahead. It's perfectly "schema-less", as it 'senses' the model of each document. Regarding ingestion speed, it is very high, even compared to MongoDB's. There are some benchmarks in this presentation: http://www.slideshare.net/8kdata/toro-db-pgconfeu2014 http://www.slideshare.net/8kdata/toro-db-pgconfeu2014
- cmkrnl 12y agoI don't really see why ToroDB can't be implemented on top of Oracle :) There could be some demand for it there. I needed this for a side project at work a while back, but didn't get the time to implement it. Oracle's cumbersome enough already; I for one have no desire to drag XSDs into it to use XMLDB.
- jchamberlain 12y agoVery interesting approach. One question I have is around the licensing. I see that it's licensed using AGPLv3. I would think that would require any code that connects to it to be open source. Am I reading that right? If so, then I am guessing that there will be a commercial offering w/o the restrictions? Thanks
- ahachete 12y agoYes and yes :)
- kofejnik 12y agoWait. Does AGPL really require my Python code to be released as open-source if I want to connect to your server? As I understand AGPL , if you modify the code, you must make it available to people connecting over the network, that's all.
- ahachete 12y agoErrr yes, you're right, my "yes and yes" was too quick. Sorry for that!
- ahachete 12y agoThere is now on the project's wiki (https://github.com/torodb/torodb/wiki/How-ToroDB-stores-json-documents-internally https://github.com/torodb/torodb/wiki/How-ToroDB-stores-json...) more detailed information about how ToroDB stores in tables the JSON documents. Hope this information helps :)
- kxo 12y agoI'd really like to write a similar, mongo-compatible layer for FoundationDB. This is great work.