8 ms·
Logica – Declarative logic programming language for data
- usgroup 2y agoHas anyone used Datalog with Datomic in anger? If so, what are your thoughts about Logica, and how does the proposition differ in your experience?
- dang 2y agoRelated: Google is pushing the new language Logica to solve the major flaws in SQL - https://news.ycombinator.com/item?id=29715957 https://news.ycombinator.com/item?id=29715957 - Dec 2021 (1 comment) Logica, a novel open-source logic programming language - https://news.ycombinator.com/item?id=26805121 https://news.ycombinator.com/item?id=26805121 - April 2021 (98 comments)
- usgroup 2y agoI may be misremembering but I think that at the time, Logica was the work of one developer who happened to be at Google. I'm not sure that there was an institutional push to use this language, nor that it has significant adoption at Google itself.
- thenaturalist 2y agoThis seems supported by the fact that the repo is not under a Google org and it has a single maintainer.
- diggan 2y ago> that the repo is not under a Google org I don't think that matters? github.com/google has a bunch of projects with large warnings that "This is not a Google project", not sure why or how that is. From the outside it looks like if you work at Google, they take ownership of anything you write.
- azornathogron 2y ago> From the outside it looks like if you work at Google, they take ownership of anything you write. That is precisely how it works. Disclaimer: I am not a lawyer, and I'm sure the validity and enforceability of the relevant contract clauses varies by jurisdiction.
- Y_Y 2y agoIf, like me, your first reaction is that this looks suspiciously like Datalog then you may be interested to learn that they indeed consider Logical to be "in the the Datalog family".
- jp57 2y agoI think Datalog should be thought of as "in the logic programming family", so other data languages based on logic programming are likely to be similar. And, of course the relational model of data is based on first-order logic, so one could say that SQL is a declarative logic programming language for data.
- riku_iki 2y agoOnly one active committer on github..
- thenaturalist 2y agoI don't want to come off as too overconfident, but would be very hard pressed to see the value of this. At face value, I shudder at the syntax. Example from their tutorial: EmployeeName(name:) :- Employee(name:); Engineer(name:) :- Employee(name:, role: "Engineer"); EngineersAndProductManagers(name:) :- Employee(name:, role:), role == "Engineer" || role == "Product Manager"; vs. the equivalent SQL: SELECT Employee.name AS name FROM t_0_Employee AS Employee WHERE (Employee.role = "Engineer" OR Employee.role = "Product Manager"); SQL is much more concise, extremely easy to follow. No weird OOP-style class instantiation for something as simple as just getting the name. As already noted in the 2021 discussion, what's actually the killer though is adoption and, three years later, ecosystem. SQL for analytics has come an extremely long way with the ecosystem that was ignited by dbt. There is so much better tooling today when it comes to testing, modelling, running in memory with tools like DuckDB or Ibis, Apache Iceberg. There is value to abstracting on top of SQL, but it does very much seem to me like this is not it.
- Tomte 2y agoThe syntax is Prolog-like, so people in the field are familiar with it.
- thenaturalist 2y agoWhich field would that be? I.e. I understand now that it's seemingly about more than simple querying, so me coming very much from an analytics/ data crunching background am wondering what a use case would look like where this is arguably superior to SQL.
- tannhaeuser 2y ago> Which field would that be? Database theory papers and books have used Prolog/Datalog-like syntax throughout the years, such as those by Serge Abiteboul, just to give a single example of a researcher and prolific author over the decades.
- evgskv 2y ago
- cess11 2y agoIf this is how you want to compile to SQL, why not invent your own DCG with Prolog proper? It should be easy enough if you're somewhat fluent in both languages, and has the perk of not being some Python thing at a megacorp famous for killing its projects.
- taeric 2y agoI find the appeals to composition tough to agree with. For one, most queries begin as ad hoc questions. And can usually be tossed after. If they are needed for speed, it is the index structure that is more vital than the query structure. That and knowing what materialized views have been made with implications on propagation delays. Curious to hear battle stories from other teams using this.
- FridgeSeal 2y agoDepends who your users are and what the context is. Having been in quite a few data teams, and supported businesses using dashboards, a very large chunk of the time, the requests do align with the composable feature: people want “the data from that dashboard but with x/y/z constraints too” or “<some well defined customer segment> who did a|b in the last time, and then send that to me each week, and then break it down by something-else”. Scenarios that all benefit massively from being able to compose queries more easily, especially as things like “well defined customer segment” get evolved. Even ad-hoc queries would benefit because you’d be able to throw them together faster. There’s a number of tools that proclaim to solve this, but solving this at the language level strikes me as a far better solution.
- taeric 2y agoI supported a team at a large company looking at engagement metrics for emails. Materialized views (edit: manually done) and daily aggregate jobs over indexed ranges was really the only viable solution. You could tell the new members because they would invariably think to go to base data and build up aggregates they wanted, and not look directly for the aggregates. That is so say, you have to define the jobs that do the aggregations, as well. Knowing that you can't just add historical records and have them immediately on current reports. I welcome the idea that a support team could use better tools. I suspect polyglot to win. Ad hoc is hard to do better than SQL. DDL is different, but largely difficult to beat SQL, still. And job description is a frontier of mistakes.
- Agraillo 2y agoI think it is a good direction imho. Once being familiar with SQL I learned Prolog a little and similarities struck me. I wasn't the first one sure, and there are others who summarized it better than me [1] (2010-2012): Each can do the other, to a limited extent, but it becomes increasingly difficult with even small increases in complexity. For instance, you can do inferencing in SQL, but it is almost entirely manual in nature and not at all like the automatic forward-inferencing of Prolog. And yes, you can store data(facts) in Prolog, but it is not at all designed for the "storage, retrieval, projection and reduction of Trillions of rows with thousands of simultaneous users" that SQL is. I even wanted to implement something like Logica at the moment, primarily trying to build a bridge through a virtual table in SQLite that would allow storing rules as mostly Prolog statements and having adapters to SQL storage when inference needs facts. [1]: https://stackoverflow.com/a/2119003 https://stackoverflow.com/a/2119003
- cess11 2y agoPerhaps you already know this, but as a data store Prolog code is actually surprisingly convenient sometimes, similar to how you might create a throwaway SQLite3 or DuckDB for a one-off analysis or recurring batched jobs. It's trivial to convert stuff like web server access logs into Prolog facts by either hacking the logging module or running the log files through a bit of sed, and then you can formalise some patterns as rules and do rather nifty querying. A hundred megabytes of RAM can hold a lot of log data as Prolog facts. E.g. '2024-11-16 12:45:27 127.0.0.1 "GET /something" "Whatever User-Agent" "user_id_123"' could be trivially transformed into 'logrow("2024-11-16", "12:45:27", "127.0.0.1", "GET", "/something", "Whatever User-Agent", "user_id_123").', especially if you're acquainted with DCG:s. Then you could, for example, write a rule that defines a relation between rows where a user-agent and IP does GET /log_out and shortly after has activity with another user ID, and query out people that could be suspected to use several accounts.
- foobarqux 2y agoThere don't seem to be any examples of how to connect to an existing (say sqlite) database even though it says you should try logica if "you already have data in BigQuery, PostgreSQL or SQLite,". How do you connect to an existing sqlite database?
- kukkeliskuu 2y agoI was turned off by this at first, but then tried it out. These are mistakes in the documentation. The tools just work with PostgreSQL and SQLite without any extra work.
- foobarqux 2y agoHow do you connect to an existing database so that you can query it? There are examples of how you can specify an "engine" which will create a new database and use it as a backend for executing queries but I want to query existing data in an sqlite database.
- evgskv 2y agoTo connect to a database file use: @AttachDatabase("db_prefix", "your_file.db"); # Then you can query from it: Q(..r) :- db_prefix.YourTable(..r);
- foobarqux 2y agoThank you. You can't do Q(..r) in sqlite right? That's what I read in the tutorial.
- evgskv 2y agoAh, yes, you're right! Please do: # ... Q(your_column) :- example_db.YourTable(your_column:); You can query multiple columns of course. Feel free to start threads in Discussions of the repo with whatever questions you have!
- 2y ago
- avodonosov 2y ago> Composite(a * b) distinct :- ... Wait, does Logica factorize the number passed to this predicate when unifying the number with a * b? So when we call Composite (100) it automatically tries all a's and b's who give 100 when m7ltiplied I'd be curious to see the SQL it transpiles to.
- puzzledobserver 2y agoAs someone who is intimately familiar with Datalog, but have not read much about Logica: The way I read these rules is not from left-to-right but from right-to-left. In this case, it would say: Pick two numbers a > 1 and b > 1, their product a*b is a composite number. The solver starts with the facts that are immediately evident, and repeatedly apply these rules until no more conclusions are left to be drawn. "But there are infinitely many composite numbers," you'll object. To which I will point out the limit of numbers <= 30 in the line above. So the fixpoint is achieved in bounded time. Datalog is usually defined using what is called set semantics. In other words, tuples are either derivable or not. A cursory inspection of the page seems to indicate that Logica works over bags / multisets. The distinct keyword in the rule seems to have something to do with this, but I am not entirely sure. This reading of Datalog rules is commonly called bottom-up evaluation. Assuming a finite universe, bottom-up and top-down evaluation are equivalent, although one approach might be computationally more expensive, as you point out. In contrast to this, Prolog enforces a top-down evaluation approach, though the actual mechanics of evaluation are somewhat more complicated.
- avodonosov 2y agoBottom-up? Ok, I see. From the tutorial it seems Logica reifies every predicate into an actual table (except for what they call "infinite predicates") I found a way to look at the SQL it generates without installing anything: Execute the first two cells in the online tutorial collab (the Install and Import). Then replace the 3rd cell content with the following and execute it: %%logica Composite @Engine("sqlite"); # don't try to authorise and use BigQuery # Define numbers 1 to 30. Number(x + 1) :- x in Range(30); # Defining composite numbers. Composite(a * b) distinct :- Number(a), Number(b), a > 1, b > 1; # Defining primes as "not composite". Prime(n) distinct :- Number(n), n > 1, ~Composite(n); Look at the SQL tab in the results.
- anorak27 2y agoThere's also Malloy[0] from Google that compiles into SQL > Malloy is an experimental language for describing data relationships and transformations. [0]: https://github.com/malloydata/malloy https://github.com/malloydata/malloy
- pstoll 2y agoCame here to mention Malloy. Which is from the team that built Looker, which Google acquired. The Looker CTO founder then started (joined?) Malloy. And … he recently 6mo ago moved from Google to Meta. https://www.linkedin.com/posts/medriscoll_big-news-in-the-data-world-lloyd-tabb-activity-7192566372646707200-5eTk https://www.linkedin.com/posts/medriscoll_big-news-in-the-da... Also for those playing along at home - a few other related tools for “doing more with queries”. - AtScale - a semantic layer not dissimilar to LookML but with a good engine to optimize pre building the aggregates and routing queries among sql engines for perf. - SDF - a team that left Meta to make a commercial offering for a sql parser and related tools. Say to help make dbt better. (No affiliation other than having used / been involved with / know some of these people at work)
- transfire 2y agoNice idea, but the syntax seems hacky.
- cynicalsecurity 2y agoThis is going to be a hell in production. Someone is going to write queries in this new language and then wonder why the produced MySQL queries in production take 45 minutes to execute.
- evgskv 2y agoThere is a standard method of optimization - breaking predicate into smaller ones and saving intermediates into database. Typically program runs efficiently, but when optimization is needed - you can do it by breaking up the predicate.
- usgroup 2y agoIts nice to see Logica has come on a bit. A year or two ago I tried to use this in production and it was very buggy. The basic selling point is a compositional query language, so that over-time one may have a library of re-usable components. If anyone really has built such a library I'd love to know more about how it worked out in practice. It isn't obvious to me how those decorators are supposed to compose and abstract on first look. Its also not immediately obvious to me how complicated your library of SQL has to be for this approach to make sense. Say I had a collection of 100 moderately complex and correlated SQL queries, and I was to refactor them into Logica, in what circumstances would it yield a substantial benefit versus (1) doing nothing, (2) creating views or stored procedures, (3) using DBT / M4 or some other preprocessor for generic abstraction.
- thenaturalist 2y agoNever heard of M4 before and, lo and behold, of course HN has a discussion of it: https://news.ycombinator.com/item?id=34159699 https://news.ycombinator.com/item?id=34159699 The author discusses Logica vs. plain SQL vs POSIX. I’d always start with dbt/ Sqlmesh. The library you’re talking about exists: dbt packages. Check out hub.getdbt.com and you’ll find dozens of public packages for standardizing sources, data formatting or all kinds of data ops. You can use almost any query engine/ DB out there. Then go for dbt power user in VS Code or use Paradime and you have first class IDE support. I have no affiliation with any of the products, but from a practitioner perspective the gap between these technologies (and their ecosystems) is so large that the ranking of value for programming is as clear as they come.
- thom 2y agoM4 is absolutely ancient, one of those things you've probably only seen flashing by on your screen if you've found yourself running `make; make install`. I suppose it is a perfectly cromulent tool for SQL templating but you're right that you must be able to get more mileage out of something targeted like dbt/SQLMesh.
- pstoll 2y ago