9 ms·
TileDB closes $15M Series A for universal data engine
- deleted 6y ago[deleted]
- lifeisstillgood 6y agoIs this solving "where are my HDFS files" ?
- chmod775 6y agoCan you elaborate? I can't figure out what you're referring to with just that, but I wonder whether it's a silly nickname for a common limitation of databases?
- ben509 6y agoHDFS is the Hadoop distributed file system, but maybe they meant HDF5 files, which is a format the Pandas library saves in. That makes use of Numpy's n-dimensional arrays, so it's related to what these guys are working on.
- lifeisstillgood 6y agoYeah my typo - I was trying to understand with limited time budget TileDB and their marketing pitch - they mentioned something like "the problem of knowing where to keep data was solved by traditional databases, now its solved by us" On a dumb level I was just wondering if they are tracking the data sets people create - it does seem more like a data policy than a product. But I may not understsnd it well
- stavrospap 6y agoTileDB Embedded is a storage engine like HDF5, with the following differentiators: (1) it is cloud-native, (2) it supports also sparse arrays, (3) it offers rapid updates, (4) it supports data versioning and time traveling built into its format. TileDB Cloud (our cloud SaaS solution) further allows you to see which arrays you own in the cloud and which ones you share with others, along with full access logs. You can also attach arbitrary descriptions and metadata that can search on, even find and access public datasets posted by you or others.
- qeternity 6y agoIt’s difficult to tell from the landing page, but what exactly is this? It says a lot of things that it’s not, but that leaves it sounding like it’s an object store with a tightly coupled map reduce framework allowing the “pluggable computer layer”. How wrong is this assessment?
- aroch 6y agoIn genomics land its less an object store and more DB built to house giant but very sparse columnar data (think the position of genetic variants in hundreds of thousands of people) with support for interval operations and some other genomics things that traditional DBs don't support. That said, we evaluated TileDB for genomics recently and found it lacking for our use case.
- stavrospap 6y agoStavros from TileDB here. Great description of the genomics use case for TileDB. We'd be interested in learning what limitations you've found. Happy to discuss over email as well (stavros@tiledb.com).
- aroch 6y agoHey Stavros. We were looking for a data-store to integrate into a clinical genomics LIMS that supports in-system analysis. We deal with de novo sequenced clinical samples (and not genotyped samples, which seems to be what TileDB-vcf had in mind?). There are some edge cases that TileDB-vcf explicitly disallows (updates/reinserts to the same sample, overlapping variants) that are not edge cases for us but rather common occurrences.
- stavrospap 6y agoThis is an API issue with TileDB-VCF. The core TileDB library supports inserts/appends/overwrites without issues and we just need to expose those operations in the TileDB-VCF APIs. Added to our backlog, thanks!
- deleted 6y ago
- Bootvis 6y agoI couldn't find a nice example of using TileDB together with data frames. From the documentation, I understand this is one of the uses of TileDB. Is someone aware of a performance oriented blog post about TileDB and data frames?
- longtermmemory 6y agoReally, I would have thought to successfully pull off a scam like this, you would need to wait at least another decade for the practitioners burned by MongoDB and the rest to retire or forget. Yet here we are in 2020, and it's suddenly 2009 all over again. Hats off to the founders for netting $15MM in VC money for a "disruptive, innovative, game-changing" rehash of IBM's IMS from 1966.
- longtermmemory 6y agoAdmit it, you had to google IBM IMS, didn't you, HN drones? Only question is, did you do it before or after furiously downvoting me?
- khazhoux 6y agoI'm going to guess that this technology is not exactly the same as something IBM did in 1966 :-) Now, I understand a certain crankiness about MongoDB (which arguably was overhyped), but surely you don't think all database development or attempts at innovation should stop, right?
- sixdimensional 6y agoAlthough this comment is voted into oblivion, and I know the tone wasn't great, the poster is actually making an interesting point in one sense - the fact that technology often comes back around another time in a new form. For example, along the lines of IBMs IMS, I also recall the MUMPS system and language [1] and I'd be interested to know if the inventors of TileDB were familiar with the history of sparse array interfaces to databases or not, if it influenced their design in any way or it was rediscovered. Regarding MUMPS: "The MUMPS language provides a hierarchical database made up of persistent sparse arrays, which is implicitly "opened" for every MUMPS application. All variable names prefixed with the caret character ("^") use permanent (instead of RAM) storage, will maintain their values after the application exits, and will be visible to (and modifiable by) other running applications. Variables using this shared and permanent storage are called Globals in MUMPS, because the scoping of these variables is "globally available" to all jobs on the system. The more recent and more common use of the name "global variables" in other languages is a more limited scoping of names, coming from the fact that unscoped variables are "globally" available to any programs running in the same process, but not shared among multiple processes. The MUMPS Storage mode (i.e. Globals stored as persistent sparse arrays), gives the MUMPS database the characteristics of a document-oriented database." [1] https://en.wikipedia.org/wiki/MUMPS https://en.wikipedia.org/wiki/MUMPS
- carterklein13 6y agoI'm looking at some of the comments and still having a little bit of trouble understanding what makes the data engine "universal." I see reference to "universal storage" in some areas, but keep landing on a multi-dimensional array structure for data storage - and this seems kind of at odds. Maybe I'm missing something, but isn't specifying the structure of the data inherently not universal? I'm relatively shielded in my databases knowledge, though - having only worked with "traditional" tools. If I'm missing something definitely let me know!
- oxfordmale 6y agoYet Another Database.... As Shelnutt2 states below, its success will depend on how quickly TileDB can be integrated into other tools. However, I don't see any benefit in TileDB supporting fast and efficient updates (and duplicates) of time series data. Time series should be immutable, and only in rare occasions require updating.
- Shelnutt2 6y agoTime series data can vary and I'd agree that most time series is immutable. There are however use cases in which updates can happen. For instance, at my previous job before TileDB Inc, we had a case where 99% of our data was immutable and never updated but a very small amount of data could be updated if there were late arriving parts. In order to get near-realtime data we accepted that sometimes some columns might not be available within the window the datapoint represents. In that case that record might reappear at a later time with the complete and correct values. Of course there are trade offs, we could have forced a longer waiting period until we were confident the data was finalized. We also could have ignored the updates. In the end we used a system of staging tables for loading the last 24 hours of data before merging out into a more finalized table. This kept the load of updating records in the database down, and still allowed us to achieve our goals. At the time, several years ago, I was not aware of TileDB else we would have considered it instead of a more traditional database vendor.
- sdinsn 6y agoIf anyone from TileDB is reading this thread, the link for "Geospatial" under Applications in the main page's footer points to the wrong link.
- Shelnutt2 6y agoSeth from TileDB here. Thanks for reporting this, we've fixed the incorrect link.
- gk1 6y agoCongrats on the raise! Meta comment: The confusion we see in this thread is what happens when a startup tries to create a new category -- a strategy known as "category creation." Founders imagine everyone will jump onboard with this category and run to them as the de facto leader of that category. In reality, it just creates another point of confusion. Whereas before you had to explain just one thing, now you have to explain two things: What your product does, and what the category means. There are many existing subcategories within the "database" category of products. Pick the subcategory where you want to compete, and market your innovation to stand out and win over new or existing customers in that category. Timeseries, in-memory, data lake, RDBS, ... Lots to choose from. This goes for TileDB and any startup founder reading this who's about to launch something like a "Intra-Terrestrial Data Pipeline Miracle" or "Middle-Out Machine Learning Capacitor" or "The First Cloud Fog Edge Dew Platform" or whatever.
- broken_symlink 6y agoI've looked at tiledb a few times. I think it would make a lot of sense to use it as a serialization format for legion. https://legion.stanford.edu/ https://legion.stanford.edu/ Its been on my todo list to try it out for a while. Maybe my next weekend project.
- sjg007 6y agoSeems similar to pilosa... What would the differences between these two be conceptually?
- khazhoux 6y agoSeeing the number of customer testimonials on their announcement makes me once again wonder how brand-new products get traction with customers (who really serve as guinea pigs). I'd love to hear (and learn) how startups have successfully gotten their foot in the door with technologies like this. For context, I spent a bit shy of a year a while back developing a middleware idea, and severely struggled to get anyone to try it. Friends suggested open-sourcing it, but even that would have been a struggle, I'm sure.
- fra 6y agoThere's a whole book on the topic: Crossing the Chasm. Long story short you need to build this ladder to climb up the adoption curve: 1. Enthusiasts 2. Early Adopters 3. Pragmatists 4. Conservatives 5. Skeptics Each will need a different pitch to be convinced, and each will have different needs & risk tolerance. It's worth a read!
- k-rus 6y agoI believe TileDB had some customers when it was framed as the product. According to its website: "TileDB, Inc. is a data management company spun out of Intel Labs and MIT"