3 ms·
Thanks for your interest in this! It currently uses RocksDB as the storage engine. If your server has enough resources, I believe it can store TBs of data with
by zh217 4y ago
Thanks for your interest in this!
It currently uses RocksDB as the storage engine. If your server has enough resources, I believe it can store TBs of data with no problem.
Running queries on datasets this big is a complicated story. Point lookups should be nearly instant, whereas running complicated graph algorithms on the whole dataset is (currently) out of the question, since all the rows a query touches must reside in memory. Also, the algorithmic complexity of some of the graph algorithms is too high for big data and there's nothing we can do about it. We aim to provide a smooth way for big data to be distilled layer by layer, but we are not there yet.
- samuell 4y agoMany thanks for the detailed answer!
- bryanrasmussen 4y agowhen you say currently it implies it will change? Does that mean all rows will not be in memory? what if you had not so many nodes but each node had a lot of data would that improve it? Probably not but just normally I think of the number of nodes in your graph as the problem.
- zh217 4y agoYes. For example, in Postgres you can sort tables arbitrarily large, not constrained by main memory. Postgres uses external merge sort when the tables are really large. There are other situations where the working data are disk-based when they are too large in Postgres. We will eventually be able to do that in Cozo as well, but no timetable is available yet. For your second question, say you have a relation with lots of fields, one of them particularly large. As long as you don't use that field in your query, it will not impact memory usage. The query may be slower though since the RocksDB storage engine needs to read more pages from disk, but the fields that are loaded by RocksDB but not needed will be promptly evicted from memory.