3 ms·
Disclosure: I am a member of the TileDB team. TileDB is designed primarily for persistent storage, similar to HDF5, Zarr and Parquet. TileDB is a cloud-optimiz
by Shelnutt2 6y ago
Disclosure: I am a member of the TileDB team.
TileDB is designed primarily for persistent storage, similar to HDF5, Zarr and Parquet. TileDB is a cloud-optimized, dense and sparse multi-dimensional array storage engine, which is broader than Arrow's Dataframe. In addition TileDB also handles updates, time traveling and partitioning at the library level, which are not possible with the Arrow project's current choice of persistent format, Parquet. TileDB removes the need for using extra services like Delta Lake for updates or Hive for cataloging and partitioning, as you get all this functionality in a single, embeddable library.
- missosoup 6y agoI feel like you guys need to create some additional engineering-level documentation to convey all this. The landing page and few links I navigated to were all too fluffy.
- MiroF 6y agoFor things that start as academic work, often the original paper is quite helpful. I find it an enjoyable read if you have a background in databases (it actually was a reading for my grad systems class) https://people.csail.mit.edu/stavrosp/papers/vldb2017/VLDB17_TileDB.pdf https://people.csail.mit.edu/stavrosp/papers/vldb2017/VLDB17... The basic jist is: TileDB is substantially faster than alternatives for random writes/reads - this allows for more sophisticated parallel linear algebra algorithms as well as evaluation of predicates.