4 ms·
Let's see if I can help here. A lot people like async map-reduce. If you need to perform aggregation on a lot of data, its constantly growing, and you need the
by skjhn 11y ago
Let's see if I can help here.
A lot people like async map-reduce. If you need to perform aggregation on a lot of data, its constantly growing, and you need the results to be current, async map-reduce is great. In the best case scenario, the results are precomputed. In a worst case scenario, they are a few seconds out of date. However, you have the option of forcing an update if need be. Either way, it's a hell of lot faster than running the full aggregation every time it's requested.
Redis is great, but a) the memcached protocol is well established and b) Redis is more than a simple cache.
BSON vs. JSON, what's the point here?
A query doesn't have to hit every index node. That doesn't make any sense. In fact, it's quite the opposite. With local indexes, you would, in fact, hit every single node. With global secondary indexes, you hit the index node with the right index.
Are you talking about partial updates? If so, yes, that will be available in the next developer preview. Stay tuned.
- ddorian43 11y agoHi, 1. For indexes ? But only couchdb has them. If more people would liked them it would be more popular ? 2. Yeah. I agree that for distributed-persistent-memcache it's good. 2.5 Json is inefficient. 3. Yeah, but you usually shard indexes, say by user_id. So when you're filtering where user_id=x and column_b=y you hit only 1 node. 4. Things that don't have partial-updates are key-value dbs, right ? If yes, why don't you call yourself that till you actually have partial-updates ?
- skjhn 11y agoThere are a handful of databases that implement map-reduce one way or another - CouchDB, Couchbase, and MongoDB off the top of my head. Views might be a CouchDB/Couchbase concept, but not incremental map-reduce. In what way is JSON inefficient? Are we talking about size? GSI indexes may or may not be partitioned. With GSI, depending on the index size and resources available, you would most likely NOT partition the index - that's the recommendation. You can create an index on user_id and column_b, place it on a specific node, and you'd only be hitting that node for a query. Especially if it's a covering index. Again, databases without GSI indexes have an index partition on every single node - that means hitting every single one for every single query. I'm still not sure what you're trying to get at. I'm guessing you are referring to MongoDB shards and routers. However, that example doesn't make sense. If user_id is the shard key, then yes, the router sends the query to the right node. The same thing happens with Couchbase. Given the key, you get the document straight from the node that has it. However, if you have user_id, why are you querying on column_b too? Now, if user_id is not the shard key, then no, the router does not send the query to the right node, its sends the query to every single node. I'm generalizing, but key-value databases are best for key-value operations on arbitrary data. Document databases understand JSON and, as such, can provide access via queries. With Couchbase, you can choose from views, N1QL (SQL), geospatial (built on views), or full-text search (preview). Pretty far off from a key-value store. That, and it already has support for partial updates via N1QL. However, my assumption was that you were talking about partial updates via key-value operations.