5 ms·
Seth from TileDB here. The foundational invention is the TileDB universal storage engine based on dense and sparse multi-dimensional arrays. Genomic variants a
by Shelnutt2 6y ago
Seth from TileDB here.
The foundational invention is the TileDB universal storage engine based on dense and sparse multi-dimensional arrays. Genomic variants are 2D sparse arrays. Images and video are dense arrays. KV are interestingly sparse arrays as well[1]. TileDB is the first solution that can efficiently model all data with a single powerful data structure. TileDB universal storage engine natively supports cloud object stores, as well as local filesystems and we natively handle the eventual consistency through our MVCC design.
On top of the TileDB storage engine we have built and worked on integrations with a range of existing computation tools and frameworks. Our goal is to make the computation pluggable and provide fast and efficient (zero-copy or apache arrow where possible) integrations to allow you to continue to use your favorite computation tools in your favorite language. A few examples are our integrations with numpy/Pandas, Spark, MariaDB, gdal and more.
In addition, we built TileDB Cloud on our storage engine, which has two more innovations: (1) easy data sharing and logging at global scale (beyond organizations), and (2) a complete serverless infrastructure for scaling out compute, very similar to the DAG approach of dask.delayed.
[1] https://docs.tiledb.com/main/handling-key-value-stores https://docs.tiledb.com/main/handling-key-value-stores
[2] https://docs.tiledb.com/cloud/client-api/task-graphs https://docs.tiledb.com/cloud/client-api/task-graphs
- ben509 6y ago> TileDB is the first solution that can efficiently model all data with a single powerful data structure. How do you implement constraints in that model? Something like a foreign key constraint, for instance.
- eloff 6y agoIt doesn't seem like a relational database, and eventual consistency is going to make relational constraints problematic.
- ben509 6y agoNo doubt, but people need constraints. So they'll cobble together ad hoc constraints in their applications to prevent their stuff from crashing or being wrong. A DBMS's raison d'etre is to take those ad hoc code patterns and formalize them.
- eloff 6y agoYes, but but you're talking about a different kind of database. This kind of big-data store is for specialty use cases with different kinds of constraints. Foreign-keys are pretty low on the list of important things here, and when they exist likely point to an external relational database anyway.
- simonebrunozzi 6y agoSeth, congrats on the funding. > The foundational invention is the TileDB universal storage engine based on dense and sparse multi-dimensional arrays. Want some advice? Find a way to explain TileDB not to sound super intelligent, but to help readers understand. You will be much more successful if you do just that. Added bonus: start with what problem TileDB solves, that current solutions don't solve. Edit: 1) scalability for complex data; and 2) deployment; seem to be the problems you solve. Is that correct?
- qeternity 6y agoI'm not even sure this is the problem. For the target customer, this language is fine (sparse arrays being notoriously difficult). It's the rest of the language that seems to allude to some sort of innovation/breakthrough but stops short of explaining that.
- stanfordkid 6y agoIf you don't understand the value of TileDB, you probably haven't dealt with the data that it is meant to model. Using something like MongoDB for genomic or sparse geospatial data is incredibly cumbersome. The access patterns for analytical applications making use of the above data types are spatially collocated ... the subsequent implication of this is that data needs to be stored in ways that are geometrically optimized on the file-system. Consider a dataset which assembles information regarding terrain ... you collect a measurement using a laser. Some areas you use a laser scanner that is very dense (high number of samples per square km) ... in another you sparsely sample it. This results in a huge multi-dimensional array... this is well understood by people collecting the data. TileDB is built for this type of use case... "complex" data is ambiguous.
- qeternity 6y agoThanks for the reply. > dense and sparse multi-dimensional arrays Ok, this I get. Sounds very interesting. But I'm not sure how we make the jump to this: > The foundational invention is the TileDB universal storage engine I still don't get what the invention is, or what makes it any more universal than any other alternatives.
- swivelmaster 6y ago> The foundational invention is the TileDB universal storage engine Obviously what he's saying is that the secret sauce is that they have invented a secret sauce ;)
- tinco 6y agoHe's saying they're uniquely capable of storing both dense and sparse data efficiently, so which means they can store literally anything, hence the universal storage engine. Not making a value judgement btw, not very at home in the field, I wouldn't know if any other database is capable of this, nor if storing dense and sparse information in the same database is even something anyone wants or needs.
- jakebol 6y agoMost every (analytic) RDMS database system can model sparse arrays. A sparse array is modeled by defining a clustered index on the table "array" dimensions and defining a uniqueness constraint on that clustered index. This works well with columnar storage because the data needs to have (and assumed to naturally have) a total sort order on the dimensions. Ex. Vertica, Clickhouse, Bigquery... all allow you to do this. TileDB allows for efficient range queries through an R-Tree like index on the specified dimensions. Most real world data though is messy and defining a uniqueness constraint upfront (upon ingestion) is often limiting, so for practical use cases this gets relaxed to a multi-set rather than sparse array model for storage, and uniqueness imposed in some way after the fact (if required).
- Shelnutt2 6y agoI agree that many use cases of sparse data, uniqueness of the dimensions can't be guaranteed or you might not want to enforce the uniqueness. With the recent TileDB 2.0 release we introduced support for duplicates in sparse arrays which adds the support for multi-sets[1]. [1] https://github.com/TileDB-Inc/TileDB/pull/1504 https://github.com/TileDB-Inc/TileDB/pull/1504
- munro 6y ago> Images and video are dense arrays. Wouldn't that take up a lot of storage space if there is no image or video compression, and you're storing it as raw pixel & frame arrays? Is the only option to apply file compression like GZIP on that data, or can I compress the data with image/video compression like JPEG/MPEG?
- Shelnutt2 6y agoTileDB is designed around supporting multiple compressors or filters. We have architected the code so that we can add different filters as we find use cases or customer requests. Currently we support compressors such as gzip, bzip2, zstd, lz4 and also filters like double-delta, byteshuffle or and RLE. For a complete list see our docs[1]. Adding support for specific image compressors such as JPEG or video codecs is on our roadmap. [1] https://docs.tiledb.com/main/basic-concepts/tile-filters https://docs.tiledb.com/main/basic-concepts/tile-filters
- atombender 6y agoNot sure I understand the "universal storage" aspect. Is TileDB suitable for real-time workloads — real-time queries, but also real-time updates — as opposed to batch? Would it be appropriate for storing, say, inverted indexes for text search? Can it be used for secondary indexes, similar to what you'd use B-trees for? Can it store large pieces of data, e.g. large, structured JSON documents? Looks like the open-source version of TileDB is the database engine only, no clustering?
- solidasparagus 6y agoIt really isn't clear to me from your description why I might find TileDB useful or how it differs from using Redis or s3 directly. And I say this as a DL practitioner who would be very interested in a tensor database. Between the TileDB splash page and your comment I don't understand the use cases where TileDB is compelling nor do I see a lightweight way to try it out and determine the use case for myself.
- Shelnutt2 6y agoSeth from TileDB here. There are several differences from Redis or S3. Redis is primarily a in-memory key-value store. There is an option for persistence but it's not primarily designed for persisting data. TileDB is designed first and foremost for persisting data to disk or to a cloud object store. Redis does not directly/natively support persisting to cloud object stores, as it's not really in the design goals. Another difference is TileDB is a column store designed around dimensions (~primary index) and attributes (non-indexes columns), where redis is a key-value with several data types but effectively its single value (that is one key equals one value, it might be a list type, but you can't have 1 key equal to a structure or multiple datatypes). You must serialize your data structures to fit into the key-value or you need to keep track of multiple keys and indexes if you want to approximate a columnar storage using lists. Depending on your application and use case Redis and TileDB are more likely to complement each other than you to select one over the other. Using S3 directly, either with parquet or flat files is an option and one that is widely used. The problem we see with this is the fact that neither parquet nor flat files are designed for cloud objects stores and they leave a lot up to the application to implement since there is no single storage engine designed around any of these formats. This is why we've seen in the community over the last few years additional frameworks come about such as Delta Lake, Iceberg, Hudi and others. These systems are built to help facilitate the eventual consistency of cloud object stores, the multi-writer/multi-read problem and handling things like updates and deletes. By contrast TileDB has many of the features that are needed built directly into the format design and into the storage engine. TileDB's MVCC design instead of a single file allows it to natively handle updates and insertions with cloud object stores. We have designed it with the eventual consistency in mind, and are safe from corrupt reads or writes without the need for a central locking or transaction system. In addition to the advantages we believe the TileDB storage engine has with S3, I also want to mention that one of the main issues we wanted to solve with TileDB Cloud was sharing and access control of S3 data sets. S3 access policies do exist and can be used to facilitate sharing of objects with other users/accounts/public. However anyone who has dealt with the S3 policies has seen that they can grow quickly and become unwieldy as you try to limit different prefixes. It seems like many times companies end up making a bucket public to share the data instead of managing the access policies, and we all see the various data leaks that happen as a result. With TileDB Cloud we offer very easy and simple sharing capabilities. From our web console we aim for it to be trivial to select an array, and share it with another user, organization or even to make it public. [1] For some quick examples, we do have some example jupyter notebooks for running python examples [2]. These are also available on TileDB Cloud, if you sign up we are giving $10 of free credit and you are able to launch a jupyterlab instance and see these examples preloaded. We also have some quickstart examples in different languages if you prefer something other than Python [3]. The quickstarts are not full use cases but they do give you an overview of basic API usage. I'd love to get some feedback from you on what type of lightweight examples you are looking for. We are always aiming to improve our documentation and make it easier to discover about TileDB. [1] https://docs.tiledb.com/cloud/console/arrays/sharing-arrays https://docs.tiledb.com/cloud/console/arrays/sharing-arrays [2] https://github.com/TileDB-Inc/TileDB-Cloud-Example-Notebooks/ https://github.com/TileDB-Inc/TileDB-Cloud-Example-Notebooks... [3] https://docs.tiledb.com/main/quickstart https://docs.tiledb.com/main/quickstart