7 ms·
A SQL Heuristic: ORs Are Expensive
- deleted 1y ago[deleted]
- ntonozzi 1y agoI really like the extension pattern. I wish more of the tables at my company used it. Another huge benefit that we're realizing as we move some of our heaviest tables to this pattern is that it makes it really easy to index a core set of fields in Elasticsearch, along with a tag to associate it with a particular domain model. This has drastically cut down on our search-specific denormlization and let us avoid expensive index schema updates.
- prein 1y agoThis sort of thing is why looking at generated SQL while developing instead of just trusting the ORM to write good queries is so important. I find query planning (and databases in general) to be very difficult to reason about, basically magic. Does anyone have some recommended reading or advice?
- parshimers 1y agohttps://pages.cs.wisc.edu/~dbbook/ https://pages.cs.wisc.edu/~dbbook/ is a great overview, specifically chapters 13 and 14 on query optimization. it is difficult to reason about though, and every compiler is different. it takes time and enough examples to look at a query and have an intuition for what the plan should look like, and if there's something the compiler is not handling well.
- jalk 1y agoIt's a big help if you know how to retrieve and interpret execution plans for the database you use.
- hobs 1y agoYes, I was going to say, seeing the generated SQL can be almost useless depending on the execution plan. When you have a solid view of the schema and data sizes you can start to be more predictive about what your code will actually do, THEN you can layer on the complexity of the ORM hell code.
- ethanseal 1y agoHighly recommend https://use-the-index-luke.com/ https://use-the-index-luke.com/ It's very readable - I always ask new hires and interns to read it.
- progmetaldev 1y agoThis website taught me a ton, even after I thought I knew more than enough about performance. Just seeing how different databases generate and execute their SQL is a huge boon (and sometimes extremely surprising when looking at one DBMS to another).
- tehjoker 1y agoQuery Planners were considered "AI", at least among some folks, back in the day just FYI
- progmetaldev 1y agoIf you are looking to squeeze every ounce of performance from your entire application stack, I'd say you should be looking at everything your ORM produces. The ORM is basically to speed up your developers time to production, but most ORMs will have some cases where they generate terrible SQL, and you can usually run your own SQL in a stored procedure if the generated SQL is sub-optimal. I've done this quite a few times with Microsoft's Entity Framework, but as new versions come out, it's become less common for me to have to do this. Usually I need to drop to a stored procedure for code that allows searching a large number of columns, in addition to sorting on all the columns that display. I also use stored procedures for multi-table joins with a WHERE clause, when using Entity Framework. You still need to look at your generated queries, but the code is nothing like it used to be under Entity Framework under the .NET Framework (at least in my experience - YMMV - you should never just let your ORM create SQL without reviewing what it is coming up with).
- grandfugue 1y agoORM often produces horrible queries that are impossible for humans to digest. I think there are two factors. First, queries are constructed incrementally and mechanically. There is no an overview for the generator to understand what developers want to compute, or no channel for developers to specify the intention. I anticipant this will change w/ AI in the near future. Second, ORM models data following the dogmatic data normalization, on which queries are destined to be horrible. I believe that people should take a moment to view their data, think what computations they want to do on top, estimate how expensive they may be, and finally settle on a reasonable model overall. Ask ORM (or maybe AI) to help with constructing and sending queries and assembling results. But do not delegate data modeling out. With right data modeling that fits computations, queries cann't be that bad.
- jalk 1y agoCan we agree that this is only applies to queries where all the filter conditions use cols with indexes? If no indexes can be used, a single full table scan with OR surely is faster than multiple full table scans.
- ethanseal 1y agoAbsolutely. Though I don't recall seeing multiple sequential scans without a self-join or subquery. A basic filter within a sequential scan/loop is the most naive/simplest way of performing queries like these, so postgres falls back to that. Also, fwiw, BitmapOr is only used with indexes: https://pganalyze.com/docs/explain/other-nodes/bitmap-or https://pganalyze.com/docs/explain/other-nodes/bitmap-or.
- jalk 1y agoThat was the extreme case - the multi-scan would be gotten if a casual reader tried your (neat by all means) AND-only query on a non-indexed table (or partially indexed for that matter).
- deleted 1y ago[deleted]
- ethanseal 1y agoGotcha, I misunderstood your comment. The multiple counts is a definitely very contrived example to demonstrate the overhead of BitmapOr and general risk of sequential scans.
- cyanydeez 1y agoYeah, just watch community. Every OR splits the universe, duh.
- galkk 1y agoI strongly dislike the way the polroblem is presented and the “solution” is promoted. Author mentions merge join with count of top, and if th database supports index merges, it can be extremely efficient in described scenario. There are a lot of real optimizations that can be baked in such merges that author chooses to ignore. The generalized guidance without even mentioning database server as a baseline, without showing plans and/or checking if the database can be guided is just a bad teaching. It looks like author discovered a trick and tries to hammer everything into it by overgeneralizing.
- ethanseal 1y agoFor sure, there's definitely a lot of cool techniques (and I'm not aware of all of them)! And the first example is very much contrived to show a small example. I'm not super familiar with the term index merge - this seems to be the term for a BitmapOr/BitmapAnd? Is there another optimization I'm missing? The article links to my code for my timings here: https://github.com/ethan-seal/ors_expensive https://github.com/ethan-seal/ors_expensive There is an optimization that went in the new release of PostgreSQL I'm excited about that may affect this - I'm not sure. See https://git.postgresql.org/gitweb/?p=postgresql.git;a=commitdiff;h=ae4569161 https://git.postgresql.org/gitweb/?p=postgresql.git;a=commit...
- Sesse__ 1y ago> I'm not super familiar with the term index merge - this seems to be the term for a BitmapOr/BitmapAnd? Different databases will use similar terms for different operations, but I would guess that the comment refers to something similar to MySQL's index merge (which is essentially reading the row IDs of all the relevant ranges, then deduplicating them, then doing the final scan; it's similar to but less flexible than Postgres' BitmapOr).
- ethanseal 1y agoCool. I'll have to read up on that.
- to11mtm 1y agoIf ORs are expensive you've probably got a poor table design or the ORs are for some form of nasty query that is more UI/Dashboard based. A good table has the right indexes, and a good API to deal with the table is using the RIGHT indexes in it's criteria to get a good result.
- Animats 1y agoThe query optimizer knows how many items are in each index, but has no advance idea how many items will be in the result of a JOIN. An "a OR b" query on a table with millions of rows might have three hits on A, or millions of hits. The optimal query strategy for the two cases is very different. Has anyone put machine learning in an SQL query optimizer yet?
- throwaway81523 1y agoSome traditional planners try some different plans on random subsets of the data to see which plan works best. Don't need machine learning, just Bayes' rule.
- Sesse__ 1y ago> The query optimizer knows how many items are in each index, but has no advance idea how many items will be in the result of a JOIN. Query optimizers definitely try to estimate cardinalities of joins. It's a really, really hard problem, but the typical estimate is _much_ better than “eh, no idea”.
- ethanseal 1y agoI think having a way to build statistics on the join itself would be helpful for this. Similar to how extended statistics^1 can help when column distributions aren't independent of each other. But this may require some basic materialized views, which postgres doesn't really have. [1]: https://www.postgresql.org/docs/current/planner-stats.html#PLANNER-STATS-EXTENDED https://www.postgresql.org/docs/current/planner-stats.html#P...
- johnthescott 1y agocould you elaborate on pg not really having matviews?
- ethanseal 1y agoMaterialized views in Postgres don't update incrementally as the data in the relevant tables updates.^1 In order to keep it up to date, the developer has to tell postgres to refresh the data and postgres will do all the work from scratch. Incremental Materialized views are _hard_. This^2 article goes through how Materialize does it. MSSQL does it really well from what I understand. They only have a few restrictions, though I've never used a MSSQL materialized view in production.^3 [1]: https://www.postgresql.org/docs/current/rules-materializedviews.html https://www.postgresql.org/docs/current/rules-materializedvi... [2]: https://www.scattered-thoughts.net/writing/materialize-decorrelation https://www.scattered-thoughts.net/writing/materialize-decor... [3]: https://learn.microsoft.com/en-us/sql/t-sql/statements/create-materialized-view-as-select-transact-sql?view=azure-sqldw-latest#remarks https://learn.microsoft.com/en-us/sql/t-sql/statements/creat...
- Glyptodon 1y agoIf optimization is really as simple as applying De Morgan's laws, surely it could be done within the query planner if that really is the main optimization switch? Or am I misreading this somehow? Edit: I guess the main difference is that it's just calculating separate sets and then merging them, which isn't really DeMorgan's, but a calculation approach.
- Sesse__ 1y agoIn most cases, the end result is a lot messier than it looks in this minimal example. Consider a case where the query is a 12-way join with a large IN list somewhere (which is essentially an OR); would it still look as attractive to duplicate that 12-way join a thousand times? There _are_ optimizers that try to represent this kind of AND/OR expression as sets of distinct ranges in order to at least be able to consider whether a thousand different index scans are worth it or not, with rather limited success (it doesn't integrate all that well when you get more than one table, and getting cardinalities for that many index scans can be pretty expensive).
- renhanxue 1y agoThe core issue the article is pointing to is that most database indexes are B-trees, so if you have a predicate on the form (col_a = 'foo' OR col_b = 'foo'), then it is impossible to use a single B-tree lookup to find all rows that match the predicate. You'd have to do two lookups and then merge the sets. Some query optimizers can do that, or at least things that are similar in spirit (e.g. Postgres bitmap index scan), but it's much more expensive than a regular index lookup.
- jiggawatts 1y ago> but it's much more expensive than a regular index lookup. It doesn't have to be, it just "is" in some database engines for various historical reasons. I.e.: PostgreSQL 18 is the first version to support "B-tree Skip Scan" operators: https://neon.com/postgresql/postgresql-18/skip-scan-btree https://neon.com/postgresql/postgresql-18/skip-scan-btree Other database engines are capable of this kind of thing to various degrees.
- ape4 1y agoThe OR query certainly states user intent better then the rewrite. So its better at communicating with a future programmer.
- sgarland 1y agoSo add a comment. “X is easier to understand” shouldn’t mean taking a 100x performance hit.
- sgarland 1y agoAs an aside, MySQL will optimize WHERE IN better than OR, assuming that the predicates are all constants, and not JSON. Specifically, it sorts the IN array and uses binary search. That said, I’m not sure if it would have any impact on this specific query; I’d need to test.
- magicalhippo 1y agoWe had a case where a single OR was a massive performance problem on MSSQL, but not at all on Sybase SQLAnywhere we're migrating away from. Which one might consider slightly ironic given the origins of MSSQL... Anyway, the solution was to manually rewrite the query as a UNION ALL of the two cases which was fast on both. I'm still annoyed though by the fact that MSSQL couldn't just have done that for me.
- frollogaston 1y agoPostgres sometimes handles this for you, but I'm not sure exactly when it's able to do that, so I do UNION ALL.
- getnormality 1y agoI am very grateful for databases but I have so many stories of having to manhandle them into doing what seems like it should be obvious to a reasonable query optimizer. Writing one must be very hard.
- nasretdinov 1y agoIf an optimiser was as smart as a human it would take potentially minutes to come up with a (reasonably good) SQL execution plan for any non-trivial query :)
- dspillett 1y ago> it would take potentially minutes This is a key part of the problem, and something that people don't realise about query planners. The goal of the planner is not to find the best query plan no matter what, or even to find the best plan at all, it is instead to try to find a good enough plan quickly. The QPs two constraints (find something good enough, do so very quickly) are often diametrically opposed. It must be an interesting bit of code to work on, especially as new features are slowly added to the query language.
- nasretdinov 1y ago
- ulrikrasmussen 1y agoI am split on SQL. On one hand I love the declarative approach and the fact that I can improve the run-time complexity of my queries just by adding an index and leaving the queries as is. On the other hand, I hate how the run-time complexity of my queries can suddenly go from linear to quadratic if the statistics are not up to date and my query planner misjudges the amount of rows returned by a complex sub-query so it falls back to a sequential scan instead of an index lookup. I would be interested in a query execution language where you are more explicit about the use of indexes and joins, and where a static analysis can calculate the worst-case run time complexity of a given query. Does anyone know if something like that exists?
- atombender 1y agoI've had the same thought. I would be interested in a query language which allows me to write a query plan, including explicitly scanning indexes using specific methods. I like the "plumbing on the outside" approach of databases like ClickHouse where you have to know the performance characteristics of the query execution engine to design performant schemas and queries. The challenge here is that query planners are quite good at determining the best access path based on statistical information about the data. For example, "id = 1" should probably use a fast index scan because it's a point query with a single row, but "category = 42" might be have a million rows with the value 42 and would be faster with a sequential scan. But the problem doesn't surface until you have that much data. Query planners adapt to scale, while a hard-coded query plan wouldn't. That's not even getting started on things like joins (where join side order matters a lot) and index scans on multiple columns. I'm not aware of any systems that are like this. Some of the early database systems (ISAM/VSAM, hierarchical databases like CODASYL) were a bit like this, before the relational model allowed a single unified data paradigm to be expressed as a general-purpose query language.
- nijave 1y agoWhat gets me is when it's faster to build an index and rerun the query from scratch than run the original.
- nitwit005 1y ago
- boramalper 1y agoI find MongoDB's ESR (Equality, Sort, Range) Guideline[0] quite helpful in that regard, which applies to SQL databases as well (since nearly all of them too use B-trees). > Index keys correspond to document fields. In most cases, applying the ESR (Equality, Sort, Range) Guideline to arrange the index keys helps to create a more efficient compound index. > Ensure that equality fields always come first. Applying equality to the leading field(s) of the compound index allows you to take advantage of the rest of the field values being in sorted order. Choose whether to use a sort or range field next based on your index's specific needs: > * If avoiding in-memory sorts is critical, place sort fields before range fields (ESR) > * If your range predicate in the query is very selective, then put it before sort fields (ERS) [0] https://www.mongodb.com/docs/manual/tutorial/equality-sort-range-guideline/ https://www.mongodb.com/docs/manual/tutorial/equality-sort-r...
- VladVladikoff 1y agoI absolutely adore LLMs for SQL help. I’m no spring chicken with SQL but so many times now I’ve taken a poorly optimized query, run it with ‘explain’ in front of it, and dumped it into an LLM asking to improve performance, and the results are really great! Performance vastly improved and I have yet seen it make a single mistake.
- datadrivenangel 1y agoAnd the nice thing about this is that a SQL query can easily be tested to see if optimizations change the outputs!
- hans_castorp 1y agoNot sure on which Postgres version this was tested with, but the first example runs in about 2ms with my Postgres 17 installation ("cold cache"). It uses a BitmapOr on the two defined indexes. https://notebin.de/?5ff1d00b292e1cd5#AU4Gg8hnY6RAmS9LoZ18xWnGgfbk97iGLr4PkrF3WmBE https://notebin.de/?5ff1d00b292e1cd5#AU4Gg8hnY6RAmS9LoZ18xWn... This used the setup.sql from the linked GitHub repository.
- ethanseal 1y agoWhen you say cold cache, did you clear the os page cache as well as the postgres buffercaches? After setup.sql, the cache will be warmish - I get 4ms on the first run. I'm using postgres 17.5 See https://github.com/ethan-seal/ors_expensive/blob/main/benchmark.sh https://github.com/ethan-seal/ors_expensive/blob/main/benchm... where I use dd to clear the os page cache. This article by pganalyze talks about it: https://pganalyze.com/blog/5mins-postgres-17-pg-buffercache-evict https://pganalyze.com/blog/5mins-postgres-17-pg-buffercache-...
- hans_castorp 1y agoI did not explicitly evict the Postgres buffer cache, but using pg_buffercache to evict all buffers for the table, yields a runtime of 23ms for me (still going for the BitmapOr). https://notebin.de/?ac3fcf55e6850f47#ERXndRrqp3X4zEWX5EC3dZUwdXzYhJd7QssK4dnSH3Kh https://notebin.de/?ac3fcf55e6850f47#ERXndRrqp3X4zEWX5EC3dZU... Which plan does Postgres choose in your case that results 100ms?
- ethanseal 1y agoExactly the same one from what I see: https://github.com/ethan-seal/ors_expensive/blob/main/explain_plans/cold_or_explain_plan https://github.com/ethan-seal/ors_expensive/blob/main/explai... Given the buffer reads seem close to yours, I believe it's page cache.
- andersmurphy 1y agoSqlite does this automatically [1] https://www.sqlite.org/optoverview.html#or_optimizations https://www.sqlite.org/optoverview.html#or_optimizations