4 ms·
I am curious about how the query performance compares to working with JSON files in Spark for ~100GB data.
by ankitrohatgi 9y ago
I am curious about how the query performance compares to working with JSON files in Spark for ~100GB data.
- maxpert 9y agoThey have some benchmarks in paper https://people.csail.mit.edu/stavrosp/papers/vldb2017/VLDB17_TileDB.pdf https://people.csail.mit.edu/stavrosp/papers/vldb2017/VLDB17...
- jakebol 9y agoJake from TileDB, Inc. here: Depending on the structure of the JSON files you are querying you maybe able to take advantage of columnar compression and massively reduce the dataset size (especially if the json files contain numeric data). Also, repeat queries will not have to re-parse the JSON files. This may speed up queries quite a lot, but it depends on the specifics of your problem.
- StanSeltser 9y agoStanislav Seltser, Petacube you are talking comparing structured workload(array-based TileDB) to unstructured one (JSON+Spark). Once you convert your JSON to sparce array structure (one time conversion) TileDB will beat Spark+JSON by several orders of magnitude. Caveat: assuming your spark+json workoad is a some heayy processing not a lightweight one.