4 ms·
For the subproblem of being able to unify and query various data sources in different formats, I would suggest to take a look at Datalog and specifically Mangle
by burakemir 2y ago
For the subproblem of being able to unify and query various data sources in different formats, I would suggest to take a look at Datalog and specifically Mangle, my implementation of it. I don't want to plug the project here but more describe the approach.
Usually your data will comfortably fit in a file. Your data getter emits these files in facts (essentially relations). If you want structures data, it can also be a single column that is of some struct type (similar to protobuf).
With all data available the problem becomes one of querying. With a good enough query language and system, you write these can data transformations via Datalog rules which roughly correspond to database views.
It is always possible to write queries in code in a general purpose language, but is a bit clumsy and hard to get an overview or reuse. It may also be possible to do SQL but it SQL is not very compositional and you ask yourself whether the base data representation should be adapted refactored. Essentially you do not want to think about the optimal schema or set of structs but just do transformations you need in the lowest friction way.
With Datalog you may benefit from a unified representation (everything is "facts") and the transformations to useful different formats (different kinds of facts) can be factored and reused. It may mean duplication and denormalization but usually that does not matter.
Mangle supports aggregation and even calling some custom functions during query evaluation.
The repo is at https://github.com/google/mangle https://github.com/google/mangle and obviously there remains a lot to do, the API is unstable, there are bugs and the type checker is not finished... but a number of people and projects seem to use it. Even if you do not use it, it may give you how to use facts (relations) as a unified data structure for your project.
- joshcanhelp 2y agoThank you very much for writing all of this out! My abstracted YAML representation worked pretty well for this initial PoC but I'm seeing places where this is just not going to fly for more complicated queries. I'd much rather string together more competent technologies to accomplish my goal rather than spending a year reinventing the wheel. Along the lines of your comment ... trying these queries in a general language and in SQL brought me to the same conclusions you have here. One of the foundations of this is open source so if I write a recipe and contribute it, I want someone else to be able to take it and tinker with it or extend it but do that without having to deal with too much boilerplate. I will take a look at Datalog. I know I've seen it before but don't know much about it. Thank you again!