18 ms·
Joins 13 Ways
- ttfkam 3y agoExcellent! We need more articles like this that demonstrate the subtleties of the relational model to primarily app-level developers. The explanations and explorations in terms of functional programming are both concise and compelling.
- bob1029 3y agoOnce I started thinking about joins in terms of spatial dimensions, things got a lot easier to reason with. I like to think of the inner join like the scene from Stargate where they were describing how gate addressing works. Assume you have a contrived example with 3 tables describing the position of an off world probe in each dimension - X, Y and Z. Each table looks like: CREATE TABLE Dim_X ( int EntityId float Value ) Establishing a point in space for a given entity (regardless of how you derive that ID) is then a matter of: SELECT Dim_X.Value AS X, Dim_Y.Value as Y, Dim_Z.Value as Z FROM Dim_X, Dim_Y, Dim_Z WHERE Dim_X.EntityId = Dim_Y.EntityId --Join 1 AND Dim_Y.EntityId = Dim_Z.EntityId --Join 2 AND Dim_X.EntityId = @MyEntityId --The entity we want to find the 3D location of You will note that there are 2 inner joins used in this example. That is the bare minimum needed to construct a 3 dimensional space. Think about taking 3 cards and taping them together with 2 pieces of rigid metal tape. You can make a good corner of a cube, even if it's a bit floppy on one edge. Gravity doesn't apply inside the RDBMS, so this works out. This same reasoning can be extrapolated to higher, non-spatial dimensions. Think about adding time into that mix. In order to find an entity, you also now need to perform an inner join on that dimension and constrain on a specific time value. If you join but fail to constrain on a specific time, then you get a happy, comprehensive report of all the places an entity was over time. The other join types are really minor variations on these themes once you have a deep conceptual grasp (i.e. can "rotate" the schema in your mind). Playing around with some toy examples can do wonders for understanding. I sometimes find myself going back to the cartesian coordinate example when stuck trying to weld together 10+ dimensions in a real-world business situation.
- iamcreasy 3y agoI think you are saying that every join is a variation of cross join. Did I understand you correctly?
- michaelmior 3y agoThis reminds me a lot of HyperDex[0], which hashes values into a multidimensional hyperspace based on their attributes for indexing purposes. [0] https://dbdb.io/db/hyperdex https://dbdb.io/db/hyperdex
- 1970-01-01 3y agoOLAP and MOLAP https://www.ibm.com/topics/olap https://www.ibm.com/topics/olap https://en.wikipedia.org/wiki/Online_analytical_processing#Multidimensional_OLAP_.28MOLAP.29 https://en.wikipedia.org/wiki/Online_analytical_processing#M...
- toyg 3y agoThat Wikipedia article, and it's offshoot pages, are so awfully out of date... But yes, OLAP.
- rjbwork 3y agoThis is like, a hypernormalization of the data though, isn't it? I think in standard BCNF you'd just leave your table as CREATE TABLE EntityPosition ( int EntityId, float X, float y, float z ) It does remind me of data warehouse stuff though, given we're working with aggregates and piecing together bits of various dimensions.
- jagged-chisel 3y ago> …a hypernormalization of the data … Well, yeah - it’s an example to get the point across, not an exercise in finding the right level of normalization.
- farkanoid 3y ago
- benjiweber 3y agoJoins as a relational AND. https://benjiweber.co.uk/blog/2021/03/21/thinking-in-questions-with-sql/#!:~:text=relational%20AND.%20That%E2%80%99s%20where%20NATURAL%20JOIN%20comes%20in https://benjiweber.co.uk/blog/2021/03/21/thinking-in-questio....
- ss892714028 3y ago[dead]
- nickpeterson 3y agoI imagine the title ‘thirteen ways of looking at a join’ was taken?
- lbrindze 3y agoDunno why this was downvoted, I came here to make a similar comment about the possible Wallace Steven’s reference.
- lcnPylGDnU4H9OF 3y agoObviously I'd only be able to say for sure if I was the downvoter but I sometimes observe lighter text on comments which make a reference without calling out the reference. It's similar to using an acronym without defining it, though possibly more confounding if it's a niche enough reference. In this case, I did not get the reference and would have wondered why the parent commenter expected that title to have already been used. At best, such a comment is referencing something topical which most readers will get (and ostensibly be entertained by); at worst, it's a distracting non sequitur. It generally ties back to a community preference that comments are curiosity-satisfying before entertaining.
- lbrindze 3y agoWouldn’t an allusion to 20th century American poetry fall into the category of “curiosity satisfying”? Given they were not the only person who got the reference it feels kind of arbitrary to say this is frivolous entertainment when another person in the community (in this case, me) found it curious and also wondered if there was an allusion there. If it was an intentional allusion, then it may actually add depth/meaning to the conversation but we may not know since it was already downvoted…
- lcnPylGDnU4H9OF 3y agoIt could be if it was called out as such. Without the explicit callout, one runs the risk of it “going over the readers’ heads” so to speak. Anyway, I’m not intending to justify any behavior; just offering my interpretation based on past observations.
- lopatin 3y agoVery useful. Something like this but for Flink’s fancy temporal, lateral, interval, and window joins would be great too.
- mirekrusin 3y agoFor several days I'm having trouble finding good resources on _implementation_ for query execution/planning (predicates, existing indices <<especially composite ones - how to locate them during planning etc>>, joins etc). Google is spammed with _usage_. Anybody has some recommendations at hand? ps. the only one I found was CMU's Database Group resources, which are great
- malfist 3y agoIt seems content farms, farming keywords with just fluff has taken over google. Can't find how to do anything anymore, just pages and pages of people talking about doing something. You could try your query on a different search engine. I've had good luck with kagi.
- dattl 3y agoI find the Advanced Databases Course from CMU an excellent resource. https://15721.courses.cs.cmu.edu/spring2023/schedule.html https://15721.courses.cs.cmu.edu/spring2023/schedule.html You might want to look into academic papers, e.g., T. Neumann, Efficiently Compiling Efficient Query Plans for Modern Hardware, in VLDB, 2011 https://www.vldb.org/pvldb/vol4/p539-neumann.pdf https://www.vldb.org/pvldb/vol4/p539-neumann.pdf
- mrkeen 3y agoI just tried adding 'relational algebra' in front of my 'query planning' query, and at a glance it skews more towards implementation.
- aidos 3y agoMy go to pointer for this is to read the Postgres docs (and / or source - which is also super readable). https://www.postgresql.org/docs/current/planner-optimizer.html https://www.postgresql.org/docs/current/planner-optimizer.ht...
- skywhopper 3y agoIt’s out of date and probably has less detail than I remember but I got a lot out of “Inside Microsoft SQL Server 7.0” which does deep dives into storage architecture, indices, query planning etc from an MSSQL specific POV. The book was updated for SQL Server 2003 and 2008. Obviously the book also has a ton of stuff about MS specific features and admin functionality that’s not really relevant, and I’m sure there are better resources out there, but I’ve found the background from that book has helped me understand the innards of Oracle, MySQL, and Postgres in the years since.
- charles_f 3y agoVery nice explanation! > The correct way to do this is to normalize the table This is true for transactional dbs, but in data warehouses it's widely accepted that some degree of denormalization is the way to go
- maxdemarzi 3y agoThe 14th way is “multi way joins” also called “worst case optimal joins” which is a terrible name. It means instead of joining tables two at a time and dealing with the temporary results along the way (eating memory), you join 3 or more tables together without the temporary results. There is a blog post and short video of this on https://relational.ai/blog/dovetail-join https://relational.ai/blog/dovetail-join and the original paper is on https://dl.acm.org/doi/pdf/10.1145/3180143 https://dl.acm.org/doi/pdf/10.1145/3180143 I work for RelationalAI, we and about 4 other new database companies are bringing these new join algorithms to market after ten years in academia.
- namibj 3y agoNegating inputs (set complement) turns the join's `AND` into a `NOR`, as Tetris exploits. The worst case bounds don't tighten over (stateless/streaming) WCOJ's, but much real world data has far smaller box certificates. One thing I didn't see is whether Dovetail join allows recursive queries (i.e., arbitrary datalog with a designated output relation, and the user having no concern about what the engine does with all the intermediate relations mentioned in the bundle of horn clauses that make up this datalog query). Do you happen to know if it supports such queries?
- gavinray 3y agoJustin also has a post on WCOJ that's really solid: https://justinjaffray.com/a-gentle-ish-introduction-to-worst-case-optimal-joins https://justinjaffray.com/a-gentle-ish-introduction-to-worst...
- captaintobs 3y agoVery cool work!
- danbruc 3y ago1. cross join 2. natural join 3. equi join 4. theta join 5. inner join 6. left outer join 7. right outer join 8. full outer join 9. left semi join 10. right semi join 11. left anti semi join 12. right anti semi join 13. ???
- iaabtpbtpnn 3y ago13. lateral join :)
- roywiggins 3y agoThe 0th would be "it's an operator in relational algebra." https://en.m.wikipedia.org/wiki/Relational_algebra https://en.m.wikipedia.org/wiki/Relational_algebra ("The result of the natural join [R ⋈ S] is the set of all combinations of tuples in R and S that are equal on their common attribute names... it is the relational counterpart of the logical AND operator.") ⋈ amounts to a Cartesian product with a predicate that throws out the rows that shouldn't be in the result. Lots of SQL makes sense if you think of joins this way.
- gpderetta 3y agoThe cross-product plus predicate is referenced in "A join is a nested loop over rows".
- roywiggins 3y agoThat's true, but I think the benefit of treating a Cartesian product as fundamental is that it lets you stop thinking about loops at all, or which one is the inner or outer loop, or of it as an iterative process at all. Boxing all that up into the Cartesian product is a really useful concept, and the whole idea of relational algebra is to find convenient formalisms for relational operations, so it seems like it deserves a separate mention.
- cmrdporcupine 3y agoAbsolutely. One of the biggest dis-services that SQL does is getting people thinking of relations as "tables" with "rows" which ends up making them think in an iterative, sequential, tabular model for something that I think they'd be better off thinking more abstractly about. The whole relational model clicked for me a whole lot more once I started thinking of each tuple as factual propositions (customer a's name is X, phone number is Y), and then all the operations in the relational algebra start to look more like "how would it be best to ask questions about this subject?"... "I'm interested in facts about..."
- code_biologist 3y ago1,000%. I learned discrete/set math in college and then later SQL on the job. Helping a few analysts moving beyond Excel skills to learn SQL was interesting. They struggled with some things that clicked quickly for me because I immediately had "oh, these are set operators" intuition. Stuff like when to use an outer join vs a cross join, or how to control cardinality of output rows in fancy ways (vs slapping `DISTINCT ON` on everything). I passed on the set understanding where it explained an unintuitive SQL behavior and I hope it helped 'em.
- bot12345 3y ago[flagged]
- deleted 3y ago[deleted]
- smif 3y agoAn inner join is a Cartesian product with conditionals added.
- RobinL 3y agoThere's a big performance difference between creating the Cartesian product and then filtering for the conditional, and creating the conditionals directly. An inner join with equi join conditions creates the conditions directly; any non equi join conditions actually have to be evaluated
- smif 3y agoThat's true, but that is an implementation detail. In abstract terms, you can think of an inner join like that. Also, while what you say is true in general for modern DB's, there are some implementations like old Oracle versions where the only way to create the effect of an inner join was in terms of a Cartesian product.
- srcreigh 3y agoAnother missed chance to educate on the N+1 area. Join on unclustered index is still N+1, it’s just N+1 on disk instead of N+1 over the network and disk.
- tourist2d 3y ago"Another missed chance to talk about X problem I find interest in and would bloat the article"
- boredemployee 3y agoI always thought venn diagram was a good representation but I think I was wrong. Edit: Why did I get down voted? :)
- ericHosick 3y agoI would also like to know why you are getting downvoted. Even wikipedia uses a Venn diagram to explain JOIN https://en.wikipedia.org/wiki/Join_(SQL) https://en.wikipedia.org/wiki/Join_(SQL) . Not trying to use an argument from authority but just pointing out that this is not unheard of.
- radiospiel 3y agoVenn diagrams are a terrible way to describe join types (and, tbh, I don’t understand why Wikipedia has these) because it makes it look like applying an 1:1 relationship. In a M:N relationship „artefacts“ (for lack of better words) of both tables would appear multiple times, and the venn diagram obscures this fact
- ericHosick 3y ago> because it makes it look like applying an 1:1 relationship. People can form different mental models of the same abstraction so I see what you are saying I've never seen it that way because "Venn diagrams do not generally contain information on the relative or absolute sizes (cardinality) of sets." (see https://en.wikipedia.org/wiki/Venn_diagram https://en.wikipedia.org/wiki/Venn_diagram).
- somat 3y agoBecause a venn diagram does not describe join mechanics well. A venn diagram does however describes the UNION, INTERSECT, EXCEPT part of sql. https://blog.jooq.org/say-no-to-venn-diagrams-when-explaining-joins/ https://blog.jooq.org/say-no-to-venn-diagrams-when-explainin... And more meta, it is an innocent slightly incorrect statement, stuff like that should not be down voted, reply with a correction. Save down votes for outright malicious posts.
- 3y ago
- Anon4Now 3y agoA couple things: This is admittedly a bit pedantic, but in E.F. Codd's original paper, he defined "relation" as the relation between a tuple (table row) and an attribute (table column) in a single table - https://en.wikipedia.org/wiki/Relation_(database) https://en.wikipedia.org/wiki/Relation_(database). I'm not sure of the author's intent, but the user table example (user, country_id ) might imply the relationship between the user table and the country table. It's a common misconception about what "relational" means, but tbh I'm fine with that since it makes more sense to the average developer. If you ever need to join sets of data in code, don't use nested loops - O(n^2). Use a map / dictionary. It's one of the few things I picked up doing Leetcode problems that I've actually needed to apply irl.
- zzleeper 3y agoI think it depends on the data. If it's presorted by the join variable then rolling the loop is faster. Also, if the index is too big for memory, then it might be faster to loop.
- Anon4Now 3y agoYeah. Many database tables are too large for an in memory hash join. My comment was a not very well fleshed out tangential remark on the value of practicing DS&A problems. I know a lot of devs hate Leetcode style interviews. I get it. It's not fun. But contrary to what some people say, I have run into a fair number of situations where the practice helped me implement more efficient solutions.
- fuy 3y agoI think the blog author uses it the same way Codd does, that's why he's talking about two columns from the same table.
- abtinf 3y agoI think understanding that “relation” means the relationships of attributes to a primary key is a crucial, foundational concept. Without it, you can’t properly understand a concept like normalization, why it arises, and when it applies. You can’t even fully grasp the scope of “data corruption” (most developers have an extremely poor understanding of how broad the notion of corruption is in databases).
- waynecochran 3y agoCan join also be thought of as unification? Similar to type checking. https://www.cs.cornell.edu/courses/cs3110/2011sp/Lectures/lec26-type-inference/type-inference.htm https://www.cs.cornell.edu/courses/cs3110/2011sp/Lectures/le...
- kadenwolff 3y agoThe title of this post sounds like an advertisement for a cult
- BoppreH 3y agoFantastic post, I always enjoy reading about different computation mental models. And under "A join is a…join", there's a typo in the partial order properties. It currently reads: 1. Reflexivity: a≤b, And I'm pretty sure it should be 1. Reflexivity: a≤a, instead (i.e., every element is ≤ to itself).
- foldU 3y agoYou are correct! I will fix it when I get home, thank you for the correction!
- conor-23 3y agoThis dudes blog is fire. Very nice explanations of complex database topics.
- gavinray 3y agoJustin Jaffray is a gem
- TechBro8615 3y agoThis is a nice way of explaining a concept, and could probably be applied to any complex topic. As someone who learns best by analogy, concepts usually "click" for me once I've associated them with a few other concepts. So I appreciate the format, and would personally enjoy seeing more explainers like this (and not just about database topics).
- fifilura 3y agoI think the next article could be about group by. Like (only intuitively sofar...) A group by from A rows to B rows - is a map-reduce job - is an AxB linear transformation matrix from your linear algebra course - is...
- tmpfile 3y agoGreat article. Side note: his normalization example reminded me how I used to design tables using a numeric primary key thinking they were more performant than strings. But then I’d have a meaningless id which required a join to get the unique value I actually wanted. One day I realized I could use the same unique key in both tables and save a join. Simple realization. Big payoff
- roselan 3y agoI still like to have a unique id field per table. It helps logging and it doesn't care about multi fields "real" key. However I keep an unique index on the string value and more importantly point integrity constraints to it, mainly for readability. It's way easier to read a table full of meaningful strings rather than full of numerical id or uuids.