6 ms·
(EdgeDB CTO here) In a classic relational model everything is a tuple containing scalar values. Graph-relational extends the relational data model in three wa
by RedCrowbar 5y ago
(EdgeDB CTO here)
In a classic relational model everything is a tuple containing scalar values. Graph-relational extends the relational data model in three ways:
- every relation always has a global immutable key independent of data (explicit autoincrement keys aren't needed)
- this enables us to add a "reference type", which is essentially a pointer to some other record (i.e. a foreign key)
- attributes can be set-valued, so you can have nested collections in queries and in your data model.
This is what lets us do `Movie.actors.name` instead of a bunch of `JOINs`, because `actors` is declared as a set-valued reference type in the `Movie` relation.
- miohtama 5y agoThis is a nice example. How the data is stored physically? Does the model work for large datasets and when it could break down? What are optimal workloads? Do we still need to fiddle with indexing and such?
- RedCrowbar 5y agoData is stored relationally in Postgres in 3NF. References are indexed automatically, but you still need to index type properties if you use them in `filter`.
- samhw 5y agoWait, it stores the data in Postgres? So this is essentially a data model on Postgres? FWIW, I do think there's a space in the market for a thin wrapper over Postgres (or MySQL) which would automate certain optimisations such as whether to index a particular table. It always struck me as perverse that that optimisation was delegated to the developer, when it's no more subjective or application-specific than a thousand other automated optimisations the engine makes. I'd be really interested if your project covered that.
- RedCrowbar 5y agoIt's built on Postgres, but it isn't a _thin_ wrapper. We lean hard into Postgres query machinery and type system in order to pull off EdgeQL and graph-relational efficiently.
- samhw 5y agoAh, OK, interesting! I don't have an immediate use case personally, but I wish you guys the very best. Honestly, database space needs way more competition than it has at present. There are countless permutations of the choices that database designers face, so it's a shame there aren't mature products for more of them. I hope this particular permutation turns out to be a good one for lots of people :)
- RedCrowbar 5y agoThank you!
- samhw 5y agoargh - the database space*
- Aeolun 5y agoDoes this mean that at the end it submits SQL queries to postgres? Or is the integration deeper?
- RedCrowbar 5y agoWe compile EdgeQL queries into SQL currently, because it makes the architecture simpler and less us run on unmodifed Postgres, but conceptually nothing stops us from targeting the query planner directly via an extension or an alternative frontend that consumes EdgeQL IR direclty.
- samhw 5y agoAh, this is interesting: so Postgres is effectively a 'backend' for you, in much the same way that e.g. InnoDB is a backend for MySQL[0]? And you - or hypothetically the end user - could change the backend, e.g. to Cockroach for better horizontal scalability, while trusting that EdgeDB will only rely on Postgres's public API at least in meeting its own public API/contract? [0] It's hard to make that analogy with Postgres b/c it only has one storage engine, but of course the separation still exists.
- gervwyk 5y agoI’ve always wished for MongoDB to have a “deepFind”, so that when I fetch a document, it will fetch the nested relations also instead of doing an aggression to do the lookup. Feel like if their objectID only included a collection name reference then somehow it should be possible. Perhaps a depth parameter would use be useful for more relational data. Congrats on the milestone! Will definitively have a look at edgeDB.
- jd_mongodb 5y agoMongoDB $graphLookup might do what you want. From the docs: Performs a recursive search on a collection, with options for restricting the search by recursion depth and query filter. https://docs.mongodb.com/manual/reference/operator/aggregation/graphLookup/ https://docs.mongodb.com/manual/reference/operator/aggregati...
- Too 5y agoThere used to be a DBLink data type containing both collection and id. But without any additional features using it. It was considered bad practice and eventually got deprecated. Since 99% of the time the collection link in a given attribute is fixed and known in advance so it’s just duplicate information.
- deleted 5y ago[deleted]
- swyx 5y agoput this straight onto your marketing page please!
- colinmcd 5y agoWill do.
- ComodoHacker 5y agoSo it's basically an ORM over Postgres (and only Postgres)?
- colinmcd 5y agoWe're working on a more comprehensive explanation of why EdgeDB isn't an ORM. Does EdgeDB do "object-relational mapping" under the hood — absolutely. The reason we try to distance ourselves from the category of ORMs is that the term "ORM" comes with a big bag of preconceptions that don't apply here. EdgeDB has: - Full schema model with indexing, constraints, defaults, computed properties, stored procedures - A query language that replaces SQL. If there's something you can do in SQL that isn't possible in EdgeQL, it's a bug. - The query language is backed by a full type system, grammar, set of functions and operators, etc. - A set of drivers for different languages that implement our binary protocol. By any definition, EdgeDB is a database. It's a new abstraction built on a lower-level abstraction: Postgres's query engine. Both abstractions indubitably fit any reasonable definition of "database". Basically: just because there's a declarative object-oriented schema doesn't mean this "is just an ORM" (unless your definition is quite pedantic).
- henryfjordan 5y agoHow exactly is EdgeDB run? Is it a separate process from Postgres, or some kind of plugin? Can I run it over an existing Postgres instance? If I build a DB Schema in EdgeDB, can I interact with the underlying Postgres instance using regular SQL?
- RedCrowbar 5y agoIt runs as a separate (stateless) process between the client and the PostgreSQL server. There was a talk about the details of the architecture on the live stream today: https://youtu.be/WRZ3o-NsU_4?t=5294 https://youtu.be/WRZ3o-NsU_4?t=5294
- nicoburns 5y agoThis sounds very similar to Hasura, which compiles graphQL down to SQL. Have you considered adding the subscription feature like they have?
- dudus 5y agoIt's the first time I hear about graph-relational DBs. I remember back in college learning about graph databases, but since I never touched one I don't remember much TBH. Is a graph-relational database something completely disjointed from a graph database? Or do they share some performance improvements to some use cases? Also does EdgeDB keep the advantages of a true graph database even being based on Postgres?
- RedCrowbar 5y ago> It's the first time I hear about graph-relational DBs. This is unsurprising, because we just invented the term :-) > Is a graph-relational database something completely disjointed from a graph database? Graph-relational is still relational, i.e. it's a relational model with extensions that make modeling and querying graph-like data easier. And in apps everything is graph-like (hence GraphQL etc). An important point is that graph-relational, like relational is storage-agnostic, i.e. it makes no assumptions on how data is actually arranged on disk. Pure graph databases, on the other hand, encode the assumption that data is actually _physically_ organized as a graph into their model and query languages. I guess the word "graph" is simply too overloaded in computing.
- larodi 5y agoThere’s at least one German-Bulgarian company called Plan-Vision that implemented such graph-relational approach like 15 years ago. their VSQL is similarly working on the E/R conceptual level and gets translated (or compiled into) to the underlying Postgresql or Oracle. You also get a neat EcmaScript like language that works with the collections in a graph like manner. Long before Arango, Orient etc. The company is absolutely nowhere near to you guys in terms of marketing, but their thing works with more than 40 enterprise clients so far. So you definitely did not ‘just’ invent the concept. A lot of companies approach the problem one way or another…
- ifdefdebug 5y agoHe said they invented the term, not the concept. I don't know if that is accurate either, but your missquote makes for a huge difference.
- BeefWellington 5y agoHi! Thanks for taking the time to engage on HN. I have a couple of questions around this. Firstly, what happens to the performance when I have a sizeable resultset of set-valued data? I've seen similar ideas implemented in the past that look fine for the Movies and Actors or Books and Authors examples but fall apart badly when you query a number of fields (20+) that have sets within them, which can happen on say, a sizeable reference database of marketing information. Another question: How deep in the graph can I go, and how much circular reference protection is there? E.g. if I query Movie.actors.movies.actors? I'm interested in graph databases and data modeling and while it offers some convenience I'm always skeptical but hopeful (mostly from having lost a lot of hours) that these problems have been solved sufficiently to keep performance good in practical use cases.
- RedCrowbar 5y agoEdgeDB is graph-relational, not a pure graph database, and so the performance characteristics of traversing links are that of a relational JOIN. Which, of course, depends wholly on the size of each relation being joined. So, if you want to select the list of actors for _every_ movie in your database and there are lots of movies, it'd be a pretty expensive operation. If, on the other hand, you want to select some relationships on a handful of objects (or even just one), then it doesn't really matter that much how deep your link traversal is, because all of the steps would be fast index scans. > How deep in the graph can I go, As much as you want, though the path must be explicit, EdgeQL currently doesn't have any way to say "traverse link foo recursively".
- deleted 5y ago[deleted]
- contravariant 5y agoThat's awesome. I think you've hit the nail on the head by trying to fix the SQL part of relational databases and not the relational part. It's been a pet peeve of mine for ages that relational databases have been described as inadequate for modelling relationships and graph databases have been described as the solution. You CAN'T fix the problem just by going from n-ary to binary relationships. How deeply is EdgeDB integrated into Posgresql? Any chance it could be used to query other databases eventually?