15 ms·
SQLite: Past, Present, and Future
- rafale 4y agoSQLite vs Postgres for a local database (on disk, not over the network): who wins? (Each in their most performance oriented configuration)
- RedShift1 4y agoSQLite is always going to win in that category just from the fact that there are less layers of code to be worked through to execute a query.
- remram 4y agoLatency-wise maybe, but throughput can be more important for a lot of applications or bigger databases. I say "maybe" because even there, SQLite is much more limited in terms of query-planning (very simple statistics) and the use of multiple indexes. That's assuming we're talking about reads, PostgreSQL will win for write-heavy workloads.
- electroly 4y agoAs long as you turn it into a throughput race instead of a latency race, PostgreSQL can definitely win. SQLite has a primitive query builder and a limited selection of query execution steps to choose from. For instance, all joins in SQLite are inner loop joins. It can't do hash or merge joins. It can't do GIN or columnstore indexes. If a query needs those things, PostgreSQL can provide them and can beat SQLite.
- ac2u 4y agoout of interest, what columnstore indexes are available to postgres? Would be happy to find out that I'm missing something. I know citus can provide columnar tables but I can't find columnar indexes for regular row-based tables in their docs. (use case of keeping an OLTP table but wanting to speed up a tiny subset of queries) Closest thing I could find was Swarm64 for columnar indexes but it doesn't seem to be available anymore.
- sophacles 4y ago> just from the fact that there are less layers of code to be worked through This is not an invariant. I've seen be true, and I've seen it be false. Sometimes that extra code is just cruft yes. Other times though it is worth it to set up your data (or whatever) to take advantage of mechanical sympathies in hot paths, or filter the data before the expensive processing step, etc.
- RedShift1 4y agoI'm not talking about extra code, I'm talking about _layers_ of code. With PostgreSQL you're still sending data over TCP/IP or a UNIX socket, and are copying things around in memory. Compare that to SQLite that runs in the memory space of the program, thus no need for copying and socket traffic. There's just less middlemen (middlepersons?) with SQLite that are unavoidable with PostgreSQL. So less layers = less interpreting/serialization/deserialization/copying/... = higher performance. I will even argue that even if the SQLite query engine is slightly less efficient than PostgreSQL, you're still winning because of less memory copying going around.
- fuckstick 4y ago> less interpreting/serialization/deserialization/copying/... = higher performance Unfortunately for many database workloads you are overestimating the relative cost of this factor. > even if the SQLite query engine is slightly less efficient than PostgreSQL And this is absurd - the postgresql query engine isn't just "slightly" more efficient. It is tremendously more sophisticated. People using a SQL datastore as a glorified key-value store are not going to notice - which seems to be a large percentage of the sqlite install base. It's not really a fair comparison.
- ok_dad 4y agoWith SQLite, though, you could reasonably just skip doing fancy joins and do everything in tiny queries in tight loops because SQLite is literally embedded in your app’s code. You can be careless with SQLite in ways you cannot with a monolithic database server because of that reason. I still agree there are use cases where a centralized database is better, but SQLite is a strange beast that needs a special diet to perform best.
- lvass 4y agoSQLite. The most performant configuration is unsuited to most usage, and may lead to database corruption on a system crash.
- rafale 4y agoShould have said the most performance oriented setting that's also safe from data corruption.
- lvass 4y agoThen it depends on the usage. You'd likely need to run with synchronous mode on, and even on WAL, multiple separate write transactions is a issue. If you don't have many writes or buffer them into not many transactions, SQLite is the most performant.
- thomascgalvin 4y agoThis is basically the exact use case SQLite was designed for; PostgreSQL is a marvel, and at the end of the day presents a much more robust RDBMS, but it's never going to beat SQLite at the thing SQLite was designed for.
- samatman 4y agoPostgres obviously. Sorry, just thought I'd buck the trend and assume a very write-heavy workload with like 64 cores. If you don't have significant write contention, SQLite every time.
- innocenat 4y agoWhere is write contention coming from if it's operated locally?
- Thaxll 4y agoSQLite is "single" threaded for writes.
- d3nj4l 4y ago... you can get tons of requests on a server?
- dinosaurdynasty 4y agoRedis has the same limitation (only one transaction at a time) and is used a lot for webapps. It solves this by requiring full transactions up front. The ideal case for sqlite for performance is to have only a single process/thread directly interacting with the database and having other process/threads send messages to and from the database process.
- innocenat 4y agoBut that isn't "locally"?
- ledgerdev 4y agoHere's sqlite doing 100 million inserts in 33 seconds which should fit into nearly every workload, though it is batched. https://avi.im/blag/2021/fast-sqlite-inserts/ https://avi.im/blag/2021/fast-sqlite-inserts/ So write contention from multiple connections is what you're talking about, versus a single process using sqlite?
- ergocoder 4y agoFunctionality-wise, SQLite's dialect is really lacking...
- simonw 4y agoIs it the SQL dialect there lacking or is it the built-in functions? I agree that SQLite default functionality is very thin compared to PostgreSQL - especially with respect to things like date manipulation - but you can extend it with more SQL functions (and table-valued functions) very easily.
- ergocoder 4y agoDepends on what easily means. Sqlite can't do custom format date parsing and regex extract. How do we extend something like this? If we go beyond a simple function to window function, I imagine it would be even harder. At this point, we nlmight as well use postgres.
- polyrand 4y agoAdding user-defined functions to SQLite is not difficult, and the mechanism is quite flexible. You can create extensions and load them when you create the SQLite connection to have the functions available in queries. I wrote a blog post explaining how to do that using Rust, and the example is precisely a `regex_extract` function [0]. If you need them, you also have a "stdlib" implemented for Go [1] and a pretty extensive collection of extensions [2] [0]: https://ricardoanderegg.com/posts/extending-sqlite-with-rust/ https://ricardoanderegg.com/posts/extending-sqlite-with-rust... [1]: https://github.com/multiprocessio/go-sqlite3-stdlib https://github.com/multiprocessio/go-sqlite3-stdlib [2]: https://github.com/nalgeon/sqlean https://github.com/nalgeon/sqlean
- ergocoder 4y agoWow this is helpful. I'm using sqlite for some of my projects and always bothered that some functions are missing. WITH RECURSIVE is too mind bending. This seems like I can add a lot more functions to it, not just regex extract. Came here to complain and learned something useful.
- nikeee 4y agoThe documentation offers some advice on this: https://www.sqlite.org/whentouse.html https://www.sqlite.org/whentouse.html
- bob1029 4y ago>most performance oriented configuration I am 99% sure SQLite is going to win unless you actually care about data durability at power loss time. Even if you do, I feel I could defeat Postgres on equal terms if you permit me access to certain ring-buffer-style, micro-batching, inter-thread communication primitives. Sqlite is not great at dealing with a gigantic wall of concurrent requests out of the box, but using a little bit of innovation in front of SQLite can solve this problem quite well. The key is resolve the write contention outside of the lock that is baked into the SQLite connection. Writing batches to SQLite on a single connection with WAL turned on and Sync set to normal is pretty much like operating at line speed with your IO subsystem.
- deleted 4y ago[deleted]
- prirun 4y ago> I am 99% sure SQLite is going to win unless you actually care about data durability at power loss time. SQLite will handle a power loss just fine. From https://www.sqlite.org/howtocorrupt.html https://www.sqlite.org/howtocorrupt.html: "An SQLite database is highly resistant to corruption. If an application crash, or an operating-system crash, or even a power failure occurs in the middle of a transaction, the partially written transaction should be automatically rolled back the next time the database file is accessed. The recovery process is fully automatic and does not require any action on the part of the user or the application." From https://www.sqlite.org/testing.html https://www.sqlite.org/testing.html: "Crash testing seeks to demonstrate that an SQLite database will not go corrupt if the application or operating system crashes or if there is a power failure in the middle of a database update. A separate white-paper titled Atomic Commit in SQLite describes the defensive measure SQLite takes to prevent database corruption following a crash. Crash tests strive to verify that those defensive measures are working correctly. It is impractical to do crash testing using real power failures, of course, and so crash testing is done in simulation. An alternative Virtual File System is inserted that allows the test harness to simulate the state of the database file following a crash."
- kpgaffney 4y agoI think the (unsatisfying) answer is "it depends". There's a huge amount of diversity in database workloads, even among the workloads served by SQLite as we mention in the paper. For read-mostly to read-only OLTP workloads, read latency is the most important factor, so I predict SQLite would have an edge over PostgreSQL due to SQLite's lower complexity and lack of interprocess communication. For write-heavy OLTP workloads, coordinating concurrent writes becomes important, so I predict PostgreSQL would provide higher throughput than SQLite because PostgreSQL allows more concurrency. For OLAP workloads, it's less clear. As a client-server database system, PostgreSQL can afford to be more aggressive with memory usage and parallelism. In contrast, SQLite uses memory sparingly and provides minimal intra-query parallelism. If you pressed me to make a prediction, I'd probably say SQLite would generally win for smaller databases. PostgreSQL might be faster for some workloads on larger databases. However, these are just guesses and the only way to be sure is to actually run some benchmarks.
- polyrand 4y agoRegarding hash joins, the SQLite documentation mentions the absence of real hash tables [0] SQLite constructs a transient index instead of a hash table in this instance because it already has a robust and high performance B-Tree implementation at hand, whereas a hash-table would need to be added. Adding a separate hash table implementation to handle this one case would increase the size of the library (which is designed for use on low-memory embedded devices) for minimal performance gain. It's already linked in the paper, but here's the link to the code used in the paper [1] The paper mentions implementing Bloom filters for analytical queries an explains how they're used. I wonder if this is related to the query planner enhancements that landed on SQLite 3.38.0 [2] Use a Bloom filter to speed up large analytic queries. [0]: https://www.sqlite.org/optoverview.html#hash_joins https://www.sqlite.org/optoverview.html#hash_joins [1]: https://github.com/UWHustle/sqlite-past-present-future https://github.com/UWHustle/sqlite-past-present-future [2]: https://www.sqlite.org/releaselog/3_38_0.html https://www.sqlite.org/releaselog/3_38_0.html
- kpgaffney 4y agoThat's correct, the optimizations from this paper became available in SQLite version 3.38.0. As we were writing the paper, we did consider implementing hash joins in SQLite. However, we ultimately went with the Bloom filter methods because they resulted in large performance gains for minimal added complexity (2 virtual instructions, a simple data structure, and a small change to the query planner). Hash joins may indeed provide some additional performance gains, but the question (as noted above) is whether they are worth the added complexity.
- Kalanos 4y agoi wish it had an optional server for more concurrent and networked transactions in the cloud
- jjtheblunt 4y agoyou could make one pretty easily, no?
- axelthegerman 4y agoI'd like to see that. I also think the single write situation is not great for web applications, but I don't see an easy way around it without sacrificing things like consistency
- dinosaurdynasty 4y agoIf you can treat sqlite transactions like redis transactions (send the entire transaction up front) it can work.
- bob1029 4y agoSee: LMAX Disruptor and friends. The magic spell that serializes many threads of events into one without relying on hard contention. You can even skip the hot busy waiting if you aren't trying to arbitrage the stock market. The way I would do it is a MPSC setup wherein the single consumer holds an exclusive connection to the SQLite database and writes out transactions in terms of the batch size. Basically: BEGIN -> iterate & process event batch -> END. This is very very fast due to how the CPU works at hardware level. It's also a good place to insert stuff like [a]synchronous replication logic. Completion is handled with busy/yield waiting on a status flag attached to the original event instance. You'd typically do a thing where all flags are acknowledged at the batch grain (i.e. after you committed the txn). This has some overhead, but the throughput & latency figures are really hard to argue with. It also has very compelling characteristics at the extremes in terms of system load. The harder you push, the faster it goes.
- bityard 4y ago
- manimino 4y agoTFA appears to be about adapting SQLite for OLAP workloads. I do not understand the rationale. Why try to adapt a row-based storage system for OLAP? Why not just use a column store?
- didgetmaster 4y agoIt is certainly possible to have a single system that can effectively process high volumes of OLTP traffic while at the same time performing OLAP operations. While there are systems that are designed to do one or the other type of operation well, very few are able to do both. https://www.youtube.com/watch?v=F6-O9v4mrCc https://www.youtube.com/watch?v=F6-O9v4mrCc
- Comevius 4y agoSQLite is significantly better at OLTP and being a blob strorage than DuckDB, and it doesn't want to sacrifice those advantages and compatibility if OLAP performance can be improved independently. In my experience for diverse workloads it is more practical to start with a row-based structure and incrementally transform it into a column-based one. Indeed in the paper there is a suggested approach that trades space for improved OLAP performance.
- deleted 4y ago[deleted]
- satyrnein 4y agoIt seems like one idea in there is to store it both ways automatically (the HE variant)! That might be better then manually continually copying between your row store and your column store.
- badgerdb 4y agoGreat discussion here. As one of the co-authors of the paper, here is some additional information. If you need both transactions and OLAP in the same system, the prevalent way to deliver high performance on this (HTAP) workload is to make two copies of the data. This is what we did in the SQLite3/HE work (paper: https://www.cidrdb.org/cidr2022/papers/p56-prammer.pdf https://www.cidrdb.org/cidr2022/papers/p56-prammer.pdf; talk: https://www.youtube.com/watch?v=c9bQyzm6JRU https://www.youtube.com/watch?v=c9bQyzm6JRU). That was quite clunky. This two copy approach not only wasted storage but makes the code complicated, and it would be very hard to maintain over time (we did not want to fork the SQLite code -- that is not nice). So, we approached it in a different way and started to look for how we could get higher performance on OLAP queries working as closely with SQLite's native query processing and storage framework. We went through a large number of options (many of them taken from the mechanisms we developed in an earlier Quickstep project (https://pages.cs.wisc.edu/~jignesh/publ/Quickstep.pdf https://pages.cs.wisc.edu/~jignesh/publ/Quickstep.pdf) and concluded that the Bloom filter method (inspired by a more general technique called Look-ahead Information Passing https://www.vldb.org/pvldb/vol10/p889-zhu.pdf https://www.vldb.org/pvldb/vol10/p889-zhu.pdf) gave us the biggest bang for the buck. There is a lot of room for improvement here, and getting high OLAP and transaction performance in a single-copy database system is IMO a holy grail that many in the community are working on. BTW - the SQLite team, namely Dr. Hipp (that is a cool name), Lawrence and Dan are amazing to work with. As an academic, I very much enjoyed how deeply academic they are in their thinking. No surprise that they have built an amazing data platform (I call it a data platform as it is much more than a database system, as it has many hooks for extensibility).
- gorjusborg 4y agoI came for SQLite, got sold DuckDB.
- naikrovek 4y agomy ipad won’t let me search through the PDF, but i couldn’t find where “SSB” was defined, if anywhere. i did not see it defined in the first paragraph, which is where it is first used. everyone: not all of your readers are domain experts. omissions like this are infuriating.
- ryanworl 4y agoStar Schema Benchmark https://www.cs.umb.edu/~poneil/StarSchemaB.PDF https://www.cs.umb.edu/~poneil/StarSchemaB.PDF
- jrochkind1 4y agocame here to ask this. I wondered if it was a typo for SSD!
- glhaynes 4y agoJust fyi: if you’re viewing the PDF in Safari on your iPad, you can search by typing into Safari’s Location Bar and then choosing “Find ‘xyz’” from the popup that appears.
- simonw 4y agoI shared some notes on this on my blog, because I'm guessing a lot of people aren't quite invested enough to read through the whole paper: https://simonwillison.net/2022/Sep/1/sqlite-duckdb-paper/ https://simonwillison.net/2022/Sep/1/sqlite-duckdb-paper/
- thunderbong 4y agoThat's a very comprehensive review. Thank you.
- airstrike 4y agoThank you for this. Big fan of your blog and all your contributions to Django
- sph 4y agoI waited for a tl;dr but this is even better. Much appreciated.
- badgerdb 4y agoIndeed, an excellent summary.
- stonemetal12 4y ago>SQLite is primarily designed for fast online transaction processing (OLTP), employing row-oriented execution and a B-tree storage format. I found that claim to be fairly surprising, SQLite is pretty bad when it comes to transactions per second. SQLite even owns up to it in the FAQ: >it will only do a few dozen transactions per second.
- tiffanyh 4y ago> SQLite is pretty bad when it comes to transactions per second. SQLite even owns up to it in the FAQ: "it will only do a few dozen transactions per second." That is an extremely poor quote taken way out of context. The full quote is: FAQ: "[Question] INSERT is really slow - I can only do few dozen INSERTs per second. [Answer] Actually, SQLite will easily do 50,000 or more INSERT statements per second on an average desktop computer. But it will only do a few dozen transactions per second. Transaction speed is limited by the rotational speed of your disk drive. A transaction normally requires two complete rotations of the disk platter, which on a 7200RPM disk drive limits you to about 60 transactions per second." https://www.sqlite.org/faq.html#q19 https://www.sqlite.org/faq.html#q19
- hnfong 4y agoGiven the prevalence of SSDs these days the figure might be out of date as well.
- hu3 4y agoYeah my GP got me confused. I remember doing 40k inserts/s in a trading strategy backtesting program with Go and SQLite. Reads were on the same magnitude, I want to say around 90k/s. My bottleneck was CPU.
- stonemetal12 4y agoWhat is your point? If I need transactions, not just bulk loading inserts, then SQLite isn't the bees knees. PG can handle at least an order of magnitude more transactions per second on the same hardware.
- kpgaffney 4y ago
- Thaxll 4y ago"While it continues to be the most widely used database engine in the world" It realy depends what do you mean by that, yes it's shipping in every phones and browser, but I don't consider that as a database. Is the windows registry a database? Oracle, MySQL, PG, MSSQL are the most widly used DB in the world, the web runs on those not SQLite.
- adamrezich 4y agothere are far, far more sqlite instances than Windows Registry instances in the world. "SQLite is likely used more than all other database engines combined. Billions and billions of copies of SQLite exist in the wild. [...] Since SQLite is used extensively in every smartphone, and there are more than 4.0 billion (4.0e9) smartphones in active use, each holding hundreds of SQLite database files, it is seems likely that there are over one trillion (1e12) SQLite databases in active use." https://www.sqlite.org/mostdeployed.html https://www.sqlite.org/mostdeployed.html
- dang 4y agoThe pdf: https://www.vldb.org/pvldb/vol15/p3535-gaffney.pdf https://www.vldb.org/pvldb/vol15/p3535-gaffney.pdf
- youngtaff 4y agoWhy do people have to publish papers in a weird two column academic format instead of something that's more easily readable?
- badgerdb 4y agoHa ha .. that is what the conference requires. Turns out that there is research that shows that when you are reading paper printed on paper this 2-column format is good for readability and not wasting paper. Conferences still insist on this format even though most people print papers. Now the good news is that these days, conferences have an accompanying video associated with the paper, and that may be a good place to start for many. That video will be published on the conference website (https://vldb.org/2022/ https://vldb.org/2022/) in about a week.
- youngtaff 4y agoThanks, will lookout for the video (I tend to read most things on a screen and find two columns of small text tiring)
- oxff 4y agoI read papers most of the time on phone and these two column papers are such a PITA to read lol
- mwish 4y agoI'm confused that why in Figure3, seems in Raspberry Pi, latency is slower than same queries' latency in cloud server. Did I missed something?
- oaiey 4y agoAfter seeing one diagram: Why the hack are we still talking to databases with SQL strings and not directly specifying the Query-AST? Admins, sure (a fancy UI could help there as well) but why in our code?
- AndrewDucker 4y agoCan you give an example of what you mean, and what we'd gain from it?
- dkjaudyeqooe 4y agoWhen you're accessing a SQLite database in code you have to generate a query string. Parameters ameliorate that somewhat, but in many cases you still have to regenerate a new string for each new query. It's inefficient to translate your query into a string only to have it parsed back into something structured by SQLite.
- ThemalSpan 4y agoIs that a fact or an intuition?
- AndrewDucker 4y agoAre you suggesting an alternate syntax for SQL-like queries, which is more compact? Or a specific one for SQLite? I'd be very surprised if generating the SQL query string and parsing it again was more than a trivial percentage of the query execution time, but happy to be proven wrong.
- rprospero 4y agoAs I said in another part of the thread, I was in a scenario where we were performed millions of inserts into a table of four integers, one row at a time. Generating the string and parsing it again wound up being enough to blow our 10µs time budget.
- spaniard89277 4y agoI've been learning SQL recently with PostgreSQL and MySQL in an online bootcamp here in Spain. So far very comprehensive. We've touched indexing and partitioning with EXPLAIN ANALYZE for optimizing performance, and I've implemented this strategies successfully onto an active forum I own. The SQL course has almost no love by the students but so far it has been the most useful and interesting to me. I was able to create some complex views (couldn't understand how to make materialized views in MySQL), but they were still very slow. I decided to copy most of this forum DB to DuckDB (with Knime now, until I know better), and optimization with DuckDB seems pointless. It's very, very fast. Less energy usage for my brain, and less time waiting. That's a win for me. My current dataset is about 40GB, so It's not HUGE, and sure people here in HN would laugh at my "complex" views, but so far I've reduced all my concerns from optimizing to how to download the data I need without causing problems to the server.
- wodenokoto 4y agoDon’t sell yourself short. I’m sure the minority here knows what a complex view is
- spaniard89277 4y agoI have no real world experience, I've seen things in Stack Overflow that I hardly manage to understand.
- aljgz 4y ago> The SQL course has almost no love by the students This is a big early career mistake. I've seen experienced developers use NoSql in a project where Sql is clearly a great fit, then waste lots of manpower to emulate things you get with Sql for free. Of course one's career can fall into a success path that never depends on SQL, but not learning SQL deeply is not a safe bet.
- spaniard89277 4y agoI've read this again and again in this forum and other dev communities, so I didn't hesitate. I can't say I love SQL, but it's not that bad, Databases look interesting to me.
- js8 4y agoWhy cannot SQLite have two different table storage engines for different tables, one row and the other column oriented?
- manigandham 4y agoThe same reasoning in the article applies: it's a lot of added complexity that isn't related to its core use as a general purpose in-process SQL database. Usually OLAP at these scales is fast enough with SQLite or you can use DuckDB if you need a portable format before graduating to a full on distributed OLAP system.
- ryanworl 4y agoStorage layout is not the primary issue here because IO throughput on commodity hardware has increased significantly in the last 10 years. DuckDB is significantly faster than SQLite because it has a vectorized execution engine which is more CPU efficient and uses algorithms which are better suited for analytical queries. If you implemented a scan operator for DuckDB to read SQLite files, it would still have better performance.
- 1egg0myegg0 4y agoWe have one of those! :-) And yes it is fast! https://github.com/duckdblabs/sqlite_scanner https://github.com/duckdblabs/sqlite_scanner