3 ms·
One particular part of that slide deck stuck out, and b/c there's no accompanying audio to go with it I thought I'd mention it. Towards the beginning of the dec
by dolinsky 15y ago
One particular part of that slide deck stuck out, and b/c there's no accompanying audio to go with it I thought I'd mention it. Towards the beginning of the deck in one slide he says 'MongoDB is fast.... until your data / indexes no longer fit in memory'. Later on there's the '120GB of data in MongoDB became 70GB of data in Postgres'. The next slide talks about using short key names + a key-mapping layer to preserve space.
While using longer key names in your objects will lead to larger objects / larger storage requirements, indexes do not contain the names of the keys of the documents they are created on, so if I have 10M documents in a collection and an index on a field 'ts', that index has the same size as if the field were named 'timestamp'. It's only the storage of the object that grows in size, and by the time you are handling 120GB of data it's very reasonable to imagine you would be using Memcached in front of MongoDB to obtain the actual objects.
- schmichael 15y agoWe used verbose key names and paid for it. We even failed to override _id with a useful primary key. So... not great. That being said, our data is not completely normalized in PostgreSQL, so it's not a totally unfair comparison. It may be possible to simulate an apples to apples comparison (small keys in mongo + perfectly normalized pgsql tables), but I'm not sure that would be a useful indicator for real world use cases. Suffice it to say: repeating your schema in every row/document is going to take more space than not repeating it.
- dolinsky 15y agoOh no doubt, and I think naming conventions is something that takes practice over time in this situation - too verbose leads to 'wasted' space, too little leads to making it near impossible for someone to view the database. Since indexes don't store the key names that does alleviate part of the problem as well. I tend to leave the default _id in place as it also acts as a 'created_on' timestamp, saving the need to add an additional field. Thanks for the deck. It's always good to hear about situations that didn't work out well and what can be learned going forward (though I do think you'd have a better experience using the 1.8 branch than you did with 1.4).
- moe 15y agoThe real shame is that this is still something the programmer has to worry about in first place. When you're a database in 2011 then you should have a damn good reason for asking programmers to space-optimize their identifiers. I can't see such a reason in Mongo really.
- schmichael 15y agoEveryone should be upvoting this. I vaguely remember 10gen brainstorming some ideas for fixing this, but it's a pretty poor design decision for a database you're advertising for use in "Big Data" scenarios. (Not that our measly 120 GB db constitutes big data... but even at Medium Data something as silly as key sizes can have a significant performance impact).
- bdarfler 15y agoSome/M=most Mongodb client libraries can do this for you. I can see arguments both ways to keep it in the server or keep it in the client library. 10gen does a hell of a job supporting a wide variety of very good client libraries.
- dolinsky 15y agoHow is this related only to MongoDB?
- rbranson 15y ago"Mongo eliminates the need for a separate object caching layer." -10gen