4 ms·
Good information that deals with tagging issues that I have been thinking about and implementing for several years. I have been working on a new kind of data m
by didgetmaster 3y ago
Good information that deals with tagging issues that I have been thinking about and implementing for several years.
I have been working on a new kind of data management system that is designed to replace traditional file systems while also doing a bunch of database operations. It is an object store that makes extensive use of tagging. Nearly all file metadata comes in the form of tags attached to each file. Even its name is a tag, so every file can have multiple names (each object's unique ID is a 64 bit number).
The system is currently in beta and still needs work, but it can do some amazing things so far. It can create 100M+ files and attach dozen of tags to each one (there is a limit of 255 tags per object). Searching for all photos with certain tags, for example, can give you the results extremely fast. Even with containers with 100 million objects can give you the 100,000 that have a specific tag in just a second or two.
Tags have context. Instead of 'John' as a tag, you attach Person.FirstName = 'John' to photos of John or documents written by John. The whole system is like a sparse relational database table where each row/column intersection can have multiple values.
We are looking for more beta testers and anyone can download and try the software for free. Here are a couple short videos showing a few things it can do.
https://www.youtube.com/watch?v=yLPNLHm9fIk https://www.youtube.com/watch?v=yLPNLHm9fIk
https://www.youtube.com/watch?v=dWIo6sia_hw https://www.youtube.com/watch?v=dWIo6sia_hw
- zvr 3y agoThis is wonderful, but unfortunately only for Windows. If you want beta testers for a Linux version... I'm more interested in the Manager with an API, not the GUI Browser.
- runlaszlorun 3y agoNice. I’ve wanted something like this for 15 years. I’m curious how you implemented them. I’ve been mulling using a tag based system for a completely unrelated project. I figure one might get a lot of way there using hash tables and sets from any of the high level languages. But it seems you maybe used a different approach?
- didgetmaster 3y agoEach logical container (a Pod) has a set of data objects called Didgets (short for Data Widgets) where each one is assigned a unique 64-bit number as its ID. Each tag is defined with a data type (e.g. Person.FirstName has type STRING) using a schema like how a column in a relational table is defined. All tags within the same context are stored together in a highly optimized Key-Value store (also stored within specialized Didgets) like how a columnar store database stores all the data for a single column separately. So if you attach a tag like Person.FirstName = 'John' to a photo of John that is stored in a Didgets with ID = 123; then within the tag the value of 'John' is mapped to the key 123 within the Person.FirstName tag store. This makes it extremely efficient to find all the Didgets where 'John' was attached as a person's first name. It can also find everything where the first name starts with 'J' or a vowel (using regex). The values within each tag store are de-duplicated and reference counted. In the example, the value 'John' is stored once. If 100 things have 'John' attached, then its reference count is 100. This makes it extremely fast to find everything with the 10 most popular names. The tags were so efficient and fast that I built relational tables using them that are often faster to query than traditional database systems. In this case the 'Key' each value is mapped to is the rowID within the table. Here is a video showing how that compares to SQLite: https://www.youtube.com/watch?v=Va5ZqfwQXWI https://www.youtube.com/watch?v=Va5ZqfwQXWI