7 ms·
PostgreSQL for Everything
- TekMol 2mo agoSQLite has so many advantages over PostgreSQL. No deamon. Single file per DB. Less configuration overhead.
- pstuart 2mo agoIt's probably perfect for a majority of work (and DuckDB takes that even further). But for big, multi-writer work PG is the way to go.
- BowBun 2mo agoPostgres also has many advantages over SQLite. Supporting more than 1 writer per process. Strict typing. Access controls. Replication at scale is more effecient than copy-pasting files (seems SQLite has improved on this one).
- frdev1786855380 2mo ago[flagged]
- rwultsch 2mo ago"MySQL was also potentially faster as it did not implement all features of the SQL standard. " This is a not great start. I assume it refers to MyISAM which has not been relevant for over a decade at this point. InnoDB made different design than PG decisions and was (and perhaps still is) faster at point lookups.
- radiospiel 2mo agoWell, the poster explicitly talks about 2003 here: „ In 2003, MySQL was much more widely used than PostgreSQL. MySQL was also potentially faster as it did not implement all features of the SQL standard“
- browningstreet 2mo agoAt the time, MySQL was also the default for every PHP backed webhost provider. That's the market they lost. They're the Perl of DBs.
- actionfromafar 2mo agoDidn't very early mysql play fast and loose with the concept of actually syncing to disk? That was also fast. Web scale fast. :)
- deleted 1mo ago[deleted]
- hnrprtlpdb 2mo agoThe older I get the more I agree with this
- goosethe 2mo agoyes. my people. https://github.com/seanwevans/pg_gpt2 https://github.com/seanwevans/pg_gpt2 https://github.com/seanwevans/pg_shell https://github.com/seanwevans/pg_shell https://github.com/seanwevans/pg_os https://github.com/seanwevans/pg_os
- opengears 2mo agothere is also https://postgresforeverything.com/ https://postgresforeverything.com/
- launchforge 2mo ago[flagged]
- replwoacause 2mo agoI use SQLite for everything, and I'm perfectly happy with it. I'm aware of the concurrent writer issues, but at my scale it doesn't even matter.
- taoh 2mo ago[dead]
- bensyverson 2mo agoYes, especially for a web app where there’s realistically only a need for one VM/server. By the time you outgrow that approach, a very straightforward migration to Postgres is probably the least complex problem you face.
- joewils 2mo agoSame, I posted some corrections to Dr. Bauer's article: https://joecode.com/2026-08-19-sqlite3/ https://joecode.com/2026-08-19-sqlite3/
- zulux 2mo agoPerfectly reasonable: I'm a huge PG fan, so I start everything with it, but SQLite is sane, and it generally has a happy upgrade path to PG If you need it.
- thatwasunusual 2mo agoIt's the other way around for me: as 99% of the stuff I develop is .NET (and I use EF Core for database stuff), I can get away with SQLite for local development, prototyping (and even staging), and then just "flip a switch" for it to run on production PostgreSQL. Both are amazing technologies.
- bpavuk 2mo agoEF Core is so easy to turn into a disgrace for performance, developer experience, AND build times...
- Tsarp 2mo agosqlite for everything NVMe drives + Litestream + object storage(S3/R2..). sqlite simplifies things for the entire long tail of apps/services that aren't the Ubers and AirBNBs of the world.
- skybrian 2mo agoIt looks interesting, but deciding on where to put object storage is what keeps me from doing this. I don’t have an AWS or Cloudflare account and I’m not sure what to commit to. Also, apparently Litestream could use a filesystem instead of an object store?
- Ozzie_osman 2mo agoI love postgres and use it heavily, but I still don't fully understand how it overlook MySQL. Maybe because of Heroku adopting it. MySQL was generally faster, and while MyISAM was a bit limited Innodb was pretty powerful, and you had the choice. It was also simpler (imo) and avoided a lot of the xid/vacuum issues. That said, still love Postgres. But at the time it started eclipsing MySQL, MySQL felt better positioned.
- _joel 2mo agoMaria, MySQL, Oracle shenannigans, perhaps. Also postgres is a "proper" db, so I'm glad it generally won out.
- bingemaker 2mo agoMySQL was a proven solution back in the day, i.e late 2000s. Github/Twitter/Heroku etc were using it. In the past 10-15 years, Postgres has come a long way.
- jeremyjh 2mo agoMySQL was always behind in terms of features. In early 2000s it had very limited constraints. Most people were running in ISAM backend and did not even have transaction support. It was being used by people who did not understand how advanced relational DBs were being used. What changed is Postgres overtook its actual competition, which were Oracle, SQL Server and Sybase. MySQL caught up as well as far as I know, but it still may have some poor defaults that are widely used.
- roryirvine 2mo agoIn the early days, PostgreSQL was so much more awkward to deal with. Crash-prone at first, and then there was the whole business around having to drop the db during upgrades. It didn't really match MySQL operationally until around 2002. During the dot com era it was common to develop and launch on MySQL with the intention of migrating to something else if they became successful (though your typical LAMP stack developer regarded Oracle and SQL Server as being deeply 'weird', so many were willing to stick with MySQL despite the well-known limitations of MyISAM). From where I'm standing, it seems that PostgreSQL became clearly preferable for new projects from the mid 2000s onwards, but it was only the Oracle acquisition that began to push existing users off MySQL.
- _joel 2mo agoNo mention of https://postgis.net/ https://postgis.net/ - shameful
- spread2009 2mo ago[flagged]
- mvpboyy 2mo ago[dead]
- idoubtit 2mo agoWhy write a fanboy text with unfair comparisons that hide the Postgres limitations? For instance, for many simple needs MySQL is simpler than Postgres, with similar performance and consistency. * No need for a connection pool, while many use cases with Postgres require PgBouncer and Co. * Easy sort (and basic search) of multilingual text, because MySQL has case insensitive UTF8 collations. * No need to VACUUM, which can be a hard problem (it was, the last time I used Postgres). For full text search, I once worked on a project that considered several alternatives for this, including Postgres. Manticore Search was finally chosen because it was more performant, with better search results.
- fabian2k 2mo agoIf you run a single application, or a few instances of the same application, you don't need an external pool and most frameworks have an internal connection pool anyway. Not sure if I'm missing anything here, but if I want case-insensitive search I simply create an index on lower(column) and use that to query. VACUUM is something you need to pay attention to at scale. And at that point you need to know your DB anyway and tune it. For smaller applications (and I don't mean only toy applications) it usually isn't an issue.
- andriy_koval 2mo ago> * No need for a connection pool, while many use cases with Postgres require PgBouncer and Co. is there a strong evidence you even need client side connection pool at all? What is the purpose? The limitation is that you have many clients with connection pools, they hold internal PG connection without allowing it to be reused by other clients..
- b-man 2mo agoif you are searching for something similar but with more meat: https://ebellani.github.io/blog/2026/all-you-need-is-postgresql/ https://ebellani.github.io/blog/2026/all-you-need-is-postgre...
- CSMastermind 2mo agoOr an entire book covering it: https://www.manning.com/books/just-use-postgres https://www.manning.com/books/just-use-postgres
- cpursley 2mo agoOr an entire site: https://postgresisenough.dev/ https://postgresisenough.dev/
- jihadjihad 2mo agoFor the graph database idea in Postgres, PG 19 has native support for property graphs [0]. You can set up your tables and their relationships as nodes/edges, then query against them using Cypher-esque [1] syntax. 0: https://www.postgresql.org/docs/19/ddl-property-graphs.html https://www.postgresql.org/docs/19/ddl-property-graphs.html 1: https://www.postgresql.org/docs/19/queries-graph.html https://www.postgresql.org/docs/19/queries-graph.html
- piterrro 2mo agotrue to that - currently using psql (in a single monolithic codebase) as: sql db, json db, vector store, logs store, full-text search, queue, message bus. multiple processes connected to it.
- lightbendover 2mo ago[dead]
- erlich 2mo agoIt's more "what one tool can do everything", not that its ideal. Like why people use Microsoft Teams even though its terrible. The relational model and sql force us to simplify our data models too much by eliminating relationships or just not dealing with them. Think about a nested json blob from some web service api and storing it in SQL in normalized tables. No one is going to do that. Everything just becomes a denormalized mess and everything is hacked around it. Instead of modeling things in the proper way, most of the world's data is modeled in a way so that we don't have join explosions in sql queries because they look scary. Data pipelines become these scary batch transformations where data is dumped somewhere else without anyway to trace back where it came from. I encounter so many end-user applications and systems where you wonder: "why couldn't they allow a list of items here instead of a single box" or "why can't this reference this other thing".
- datadrivenangel 2mo agopeople do that with JSON way too often...
- orev 2mo agoI think you’re responding to the general idea of a relational database, not Postgres, and definitely not what’s in the article (DR;CA). Postgres has built in data types and functions that allows it to work with unstructured json documents, like you would use in MongoDB.
- aaaronic 2mo agoThat feature is definitely part of why it's still so relevant. The hstore approach wasn't nearly enough when it was all PG offered.
- sgarland 2mo ago"Work with" != "work well." GIN indices aren't the same as B+tree, and even then, you'll have to decide / know about jsonb_path_ops vs. the default operator class. Or you just accept sub-optimal performance, I suppose. The lack of a rigid schema makes it super fun as well. Does this attribute exist in this row? Who knows! Maybe there's a long-forgotten version lurking, waiting to be retrieved, that will utterly bork the calling app.
- ericpauley 2mo agoPostgres is great, but I certainly don't think it's great for everything. For instance, while you can in theory implement OLAP aggregation you're going to be hand-rolling a bunch of stuff that something like Clickhouse gives you for free declaratively.
- molf 2mo agoI don't think the point is that PostgreSQL is great for everything. But you may get by with a single piece of infrastructure instead of 7. In most of the applications we build or maintain we use PostgreSQL + cloud storage. That's it. And it works very well, also for: storing JSON, full text search, as a queue, as a vector database. Other software may be better at providing those features, but I'm extremely happy we only need to understand & manage PostgreSQL.
- ericpauley 2mo agoThe article says verbatim “PostgreSQL Replaces Clickhouse”. Coming from storing billions of rows in Clickhouse and performing dozens of materialized operations I shudder to think about what that would look like in a DB that doesn’t even support declarative IVM.
- tpetry 2mo agoThe article suggested using TimescaleDB which has its own concept of IVM: continuos aggregates. And compared to the approach by ClickHouse it can also update the materialized views when you update/delete old raw data https://sqlfordevs.com/books+courses/timescale/05-continuousaggregates https://sqlfordevs.com/books+courses/timescale/05-continuous...
- jeremyjh 1mo agoYes but that is just one use case. Columnar OLAP engines operating on object storage can do all kinds of stuff so much better than Postgres that it may as well be a completely different capability. That said - the point is that you can get a lot further with just Postgres than many people think, and now we also have options like pg_lake. But I wish I'd changed analytics platforms A LOT sooner than I did.
- devin 2mo agoThis kind of post (Postgres! It's all you need!) is getting pretty tiresome. Postgres does not even come close to a full replacement for Elastic, and that's just the first bullet. Looking down the list it is pretty easy to go: Yes, postgres can be used instead of that for extremely basic use cases, but it all goes out the window you actually need any of the power of these other tools.
- onesandofgrain 2mo agoif you need elastic youre doing something wrong
- switchbak 2mo agoOr something big. Which is often not wrong.
- deleted 1mo ago[deleted]
- jjordan 2mo agoI'm partial to Typesense, especially for smaller data sets, since it runs primarily in memory, is easy to use and is hella fast. For bigger data sets, I hear good things about Meilisearch.
- anarazel 2mo agoFwiw, I, as someone who has worked on Postgres for a long time, also find it quite tiresome. Like there's plenty stuff I wouldn't use Postgres for, and I can probably get get more out of it than most.
- switchbak 2mo agoExactly - these recommendations often come with no context or scale provisions. Yes Postgres can work in the small for a lot of things, it can even work at surprising scale if you use it according to its strengths. But if you use it for things it doesn't shine at, at inappropropriate scale - you'll almost certainly run into issues. And resolving those can often be a bigger challenge than choosing a more suitable solution in the first place. But often I think younger/less experienced engineers just have to burn themselves, thus why this never seems to die.
- Gluber 2mo agoI tend to agree with quite a few points in the article, but some topics warrant some careful scrutiny. * As a message queue: Only if your required features are very basic, like if you need cluster communication and run your own coordination protocol on top. * High Volume Time Series: TimeScale works, but composes badly with other workloads on the same DB server ( from an operational perspective at scale ) * Vector Database: The same issues as with TimeScale.. PgVector for example lives in its own seperate "world" and the query planner sees it as a very opaque thing. Forget about adding vector storage to an existing high volume db, that must server other complex queries.. PGVector will either trash your caches, or take over your cpu so that workloads that used to work fine stall. This is IMO not a pgvector problem itself ( Kudos to those guys ) but rather that postgresql extension apis are not very good at exposing custom costs and tradeoffs to the system as a whole. * Raw Data: Works for small files... why anyone would want to store large amounts of data in it would be a mystery, where it shines is accessing LOTS of small files where internal caching etc help a lot compared to raw filesystem access ( also a bit dependent on the filesystem and its tuning though ) * Microservice: If your service is ONLY exposing json data from some database model, then it should not exist at all IMO. Create a view and be done with it.
- jjice 2mo agoI like to consider Postgres the starting point for all of these things, that can be outgrown and replaced when appropriate. I do love just shoving everything in Postgres and seeing that I only end up needing a few additional dedicated services as the product groups. Redis is usually the next pickup for me.
- Gluber 2mo agoSure, thats a good way of working.. I just have the experience when handing over a project ( consulting ) anything i have put in place will never get replaced or kept for too long outgrowing its capacity by far, and offset with huge expenses in hardware or operations. Technically not my problem anymore ( except when it breaks on a maintenance contract ) but i still like to avoid it early if i can
- 2mo ago
- onesandofgrain 2mo agothe more ive coded the more this is true
- oreally 2mo agoisn't the process per connection restriction pretty heavyweight though?
- HighlandSpring 2mo agoThis isn't just theory either, for example: Revolut is a bank that does all its event persistence and streaming on top of postgres. No traditional message queues/brokers in their stack. https://medium.com/revolut/recording-more-events-but-where-will-we-store-them-4b1dad457cf5 https://medium.com/revolut/recording-more-events-but-where-w...
- stackskipton 2mo agoAs SRE dealing with this at current company, a benefit of using well known software like Kafka is a lot of problems you will run into have solutions/guidance already available vs you having to explore solutions which a lot of time end with “Kafka could easily do this. “
- cyh555 2mo ago100% except when Kafka goes wrong, who maintains it?
- striking 2mo agoNose goes
- ethbr1 2mo agoThere are two sizes of companies: those that can afford '1+ dedicated ____-person' and those that can't. Which should filter through to technology choices more than it does.
- erispoe 1mo agoIt's not 1+ person when a system needs to operate 24/7. To have a proper on-call rotation you need 5 to 8 people.
- CodesInChaos 1mo agoOne specialist and a group of generalists is often enough. Especially if you are allowed to contact the specialist outside normal working hours in rare emergencies.
- jppope 2mo agoI like Postgres. It is a good general purpose database. I like other databases too. Other databases can do some things that Postgres can't do as well.
- juancn 2mo agoAs usual, it depends on the scale, but it's a sane default for 99% of use cases. Different use cases have different scalability limits in PG, when you get to them you need to deal with them. It would be perfect if it had somewhat transparent sharding, I mean a way to add another instance and distribute load without having to stop everything. There are solutions, but they tend to be involved and when you get to that point in many cases it makes sense to just move that workload to something else that scales better.
- BenedictPorkins 2mo ago[dead]
- warpech 2mo agoHonorable mention - https://postgrest.org/ https://postgrest.org/
- lillecarl 2mo agoI find it odd that this isn't mentioned anywhere in an article about using postgres for everything.
- up2isomorphism 2mo agoHardware is so far nowadays to make people with little systems knowledge confident to make such claims at least from their use cases. However it is neither generally reasonable nor efficient.
- sreekanth850 2mo agoHow do you implement HA in postgres, i found MySQL HA stack pretty straight forward with Innodb cluster, MySQL router and Shell.
- macartain 2mo agohttps://patroni.readthedocs.io/ https://patroni.readthedocs.io/ But I sure wish it was 'core' and we didn't have to worry about it potentially going away, becoming de-supported..
- sreekanth850 2mo agoYes, especially something that touches DB. Router and shell are stupidly simple and you get auto failover in some 30 minutes. That is something make me stick to mysql.
- aleks_me2 2mo agoCan Partitioning be used to move data to S3 Storage, for long term archiving?
- ignaciovdk 2mo agoNo, and with timescaledb is a feature of their cloud platform. For on prem you can use something like Arc: https://github.com/Basekick-Labs/arc https://github.com/Basekick-Labs/arc
- aleks_me2 2mo agoThank you. I have decided to use clickhouse with that config because of missing S3 for logs and metrics for long term store. <clickhouse> <storage_configuration> <disks> <audit_s3> <type>s3</type> <endpoint>https://S3-EndPoint/{{ audit_bucket_name }}/clickhouse/</endpoint> <access_key_id>{{ clickhouse_audit_s3_access_key }}</access_key_id> <secret_access_key>{{ clickhouse_audit_s3_secret_key }}</secret_access_key> </audit_s3> </disks> <policies> <audit_tiered> <volumes> <default> <disk>default</disk> </default> <audit_s3> <disk>audit_s3</disk> </audit_s3> </volumes> </audit_tiered> </policies> </storage_configuration> </clickhouse>
- __s 2mo agoIn theory yes, just need s3 fdw I demonstrated this with ClickHouse: https://github.com/ClickHouse/pg_clickhouse/blob/main/doc/offload-partition.sql https://github.com/ClickHouse/pg_clickhouse/blob/main/doc/of... We're working on a chdb based mechanism to copy to/from s3, maybe with fdw on top we can back table in s3 You can try similar things with pg_duckdb & pg_lake
- efxhoy 2mo agoIve built data warehouses and job queues on postgres. The DW got replaced with bigquery when we started doing more tracking. We still run postgres as an app-facing cache of the aggregated data from bigquery though. The job queue runs on the cache db, scheduling jobs to move data from bigquery into postgres. It’s pretty neat. Now we’ve run into near-real-time requirements so clickhouse is getting thrown into the mix. It’s pretty funny the lengths we go to to implement user facing analytics that’s basically just “you are visitor number X” from 1995.
- silvestrov 2mo agoPostGIS is also another very useful addition for storing, indexing, and querying geospatial data. http://www.postgis.net http://www.postgis.net
- shayonj 2mo agoNice one re: flatbuffers in `blob` column. Have been working with flatbuffers a lot and that's a neat idea in general. Not 100% sure about using PG for file system at scale however. I'd love to hear more on the challenges (vacuum, toast, anything else?)
- vlindos 2mo agoHow many production system has serious queues using PostgreSQL?
- ethagnawl 2mo ago> Timescale lately released the pgvector extension, that turns your PostgreSQL into a vector database. I don't think this is accurate and smells like an LLM hallucination to me. From the Timescale/Tiger Data _pgvectorscale_ project's README: > pgvectorscale builds on pgvector with higher performance embedding search and cost-efficient storage for AI applications. I think this is where the confusion originates. I believe pgvector is primarily Andrew Kane (@ankane) and a cadre of OSS contributors. As an aside, I've used Timescale/Tiger Data products and was very happy with them and their support. Their team was very engaged and responsive to all of our questions. They also fixed a pretty gnarly indexing bug I uncovered in pgvectorscale in an impressively short amount of time.
- akulkarni 2mo agoThanks for the kind words. And yes, Andrew Kane (et al) are the people to thank for pgvector. We (Tiger Data) developed pgvectorscale and pg_textsearch (and timescaledb, and some others)
- vivzkestrel 2mo ago- mind coming and enligtening about PostgreSQL and XML? - https://www.reddit.com/r/PostgreSQL/comments/1vbo5j8/raw_xml_vs_normalized_tables_how_would_you_store/ https://www.reddit.com/r/PostgreSQL/comments/1vbo5j8/raw_xml... - your post did not have a single word on XML hence my comment
- florianherrengt 2mo ago> PostgreSQL Replacing Your Microservice I've done that before and the code was a mess. It works at the beginning but APIs do much more than piping data from the database. When you start dealing with ACL, external calls, code reuse, etc. It's just nice to have all the tools available to you from something like Python or Go.
- kavok 2mo agoSurprised it doesn't mention LISTEN / NOTIFY.
- lvncelot 2mo agoOne addition: Postgres for all your GPU-based machine learning training and inference needs: https://github.com/postgresml/postgresml https://github.com/postgresml/postgresml
- JaggerFoo 2mo agoUse case matters. I use a SQL databases as needed. I've used Postgres, Sqlite, Duckdb, Json files with AWS Athena, Oracle enterprise for ERP systems (a multitude of schemas and objects with interoperability), and others. I'm currently, deploying Duckdb with AWS S3 Tables (Iceberg) to see how it fits for a use case I have. IT is great and always changing. Keep trying new things. Cheers
- CodeWithLeo 2mo ago[dead]
- cauchyk 2mo agoas someone who loves postgres, this take is getting pretty old. yes we can do quite a bit with extensions but extensions often need to interface with external systems and even then managed providers don't consistently support all extensions. some gaps: bm25 indexes, olap support, also extensions also run into licensing restrictions.
- jankovicsandras 1mo agoYou can do BM25 and hybrid search in Postgres. Shameless plug: https://github.com/jankovicsandras/plpgsql_bm25 https://github.com/jankovicsandras/plpgsql_bm25 BM25 search implemented in PL/pgSQL ( Unlicense / Public domain ) The repo includes also plpgsql_bm25rrf.sql : PL/pgSQL function for hybrid search ( plpgsql_bm25 + pgvector ) with Reciprocal Rank Fusion; and Jupyter notebook examples.
- pelzatessa 2mo agoHey Raphael Bauer, if you're reading this, I suggest you change the color of hrefs on your website, all of them are purple and underlined, which usually is indicator for "Already visited URL". For me that's not that much of a problem, but it was something i kept noticing when reading the article. I wonder if anyone else also had this thought or am I alone as I didn't see anyone else mention this in the comments. But I wanted to signal that nevertheless :)
- theonewolf 2mo agoFor "replacing your microservice" you should checkout PostgREST. It basically turns PostgreSQL into a microservice.
- vantassell 2mo agoI tried PostgREST and regretted it. I ran into many situations where I wanted a thicker backend between my webapp and db.
- cyberax 2mo agoIn my experience, it still kinda sucks if you want to store blobs. Anything on this front?
- dzonga 2mo agoto risk sounding like a madman - if you're a solo individual- serving b2b small businesses. then Sqlite works as well too. can run the whole thing on Cloudflare. running Postgres isn't difficult. but dealing with a VPS for low traffic is a headache that's not necessary.
- vladshiyan 2mo ago[flagged]
- deleted 2mo ago[deleted]
- ChicagoDave 2mo agoThis is exactly how tightly coupled, unmaintainable software is constructed. By picking the tools before understanding the model and building bespoke architecture. You pick the tools that the business model requires. It might be a relational data store. It might not be. You might want an event store. You might want to reduce costs with lambdas and DynamoDB. You may need a pub/sub event broker. The OP clearly loves Postgres. Cool. They also have limited experience with complex systems architectures because if they had that experience, they would have never written this article.
- ballon_monkey 1mo ago> This is exactly how tightly coupled, unmaintainable software is constructed. No. If you're struggling to build software against a DB and then abstract parts to use Redis or ES or whatever in the future, that's kinda a skill issue you or your team have with building poor software to begin with. Nothing to do with using a DB for multiple things like a Queue/Search etc.
- ChicagoDave 1mo agoThe technical solution isn’t the skill issue I’m pointing towards. It’s the business modeling skill that most developers lack, so they skip it and believe an ERD will magically cover all invariants.
- groundzeros2015 1mo ago> You might want to reduce costs with lambdas and DynamoDB. I don't think that's ever saved money. > because if they had that experience, they would have never written this article. That's not true.
- throwawaythekey 1mo ago> I don't think that's ever saved money. I spent about a year as a consultant in the AWS space, visited about ~15 clients of varying sizes. More often than not there's a single pg aurora instance responsible for 50%+ of the bill. Even worse are the serverless aurora offenders. All the indexes and guarantees of PG don't come cheaply and dynamodb pricing is not cheap but comparatively reasonable. It really is a good product if you know how to use it.
- rgbrgb 2mo agoi love postgresql but once we added ai-generated dashboard to our homegrown analytics tool [0] some of the crazy (amazing) dashboards that the ops team was building began accumulating horrendously slow db queries. I considered dynamically adding indexes or alerting around postgres slow queries but also quickly prototyped mirroring the postgres data in clickhouse. At first could not believe how fast clickhouse was on arbitrary analytics queries - like 60s to 0.5s for some gnarly queries. truly amazing software that just works without any tuning for this kind of exploratory analytics workload. so yes, i'm still a postgres maximalist (worker queues still in pg [1]) but (especially in the age of quick LLM prototypes) it's always worth measuring the more purpose-built approach. [0]: https://setoku.com https://setoku.com [1]: https://worker.graphile.org https://worker.graphile.org
- psadauskas 2mo agoMy general rule of thumb is "Use Postgres until you've discovered why you can't use Postgres." Anything you introduce is another moving part you have to operate and maintain, and in the beginning, Postgres can probably handle it. Wait for load, see where its failing, and then you'll have a better idea if adding another tool is worth the cost.
- andai 2mo agoDoesn't the same argument apply even more to using SQLite instead?
- seki285 2mo agoIn a lot of cases using SQLite means you write queries incompatible with RDBMS. No need to worry about race conditions or the amount of queries you make, when 100 selects are uber fast.
- SoftTalker 2mo agoYes, for very small or embedded, single-purpose systems sqlite is usually a good choice. It's very well tested, and there's nothing extra to run or manage. But be careful if the system starts growing beyond that, you'll want a real RDBMS before you abuse sqlite too much.
- Lio 2mo agoDepends on what you mean by "small" and "growing". If you mean database size, SQLite can handle massive amounts of data. I've seen 281 TB quoted as theoretical max size.
- pojzon 1mo agoHow does it handle 100000 write transactions per second ? From multi-client architecture and with regional HA?
- 2mo ago
- mikkelam 2mo agoExcept horizontal scaling. But a lot of companies are trying to solve that, notably multigres, neki and even pgdog.
- cheesemayo 2mo ago> Contrary to popular belief - the answer to everything is NOT 42 42 is not the answer to everything. 42 is the Answer to the Ultimate Question about Life, the Universe, and Everything.
- cheesemayo 2mo agoI really want a daemonless PostgreSQL, in the style of SQLite.
- _joel 2mo agoSo https://pglite.dev/ https://pglite.dev/?
- sgt 2mo agoIntrigued by this > After some performance checks it became clear that PostgreSQL was even faster than reading from the file system for our use-case. PostgreSQL uses the file system very efficiently for its data - and it adds a lot of caching and efficient reading and writing strategies that can outperform writing and reading raw data on a file system. This goes against conventional knowledge. I've always heard (and followed best practice) to avoid storing binary data in BYTEA columns that should otherwise be put on a filesystem or an object storage like S3. I'd like to find out more about this, because in many cases it would be very convenient indeed to store it in the database itself.
- sgarland 2mo agoThe primary reason to avoid doing so is avoiding thrashing your buffers, along with increased size of backups, WAL bloat, etc. Can you? Yes. Should you? Not at anything beyond a toy scale, unless you want to pay for more RAM to ensure that your normal OLTP queries don’t take a performance hit.
- sgt 2mo agoAgreed. Even putting them on the filesystem and rsyncing in a cronjob would be better, which says a lot.
- Tostino 2mo agoListen to this advice. I had a system that has ~600gb of blob data in bytea that could have easily been an S3 bucket + db reference. It made backups way more of a pain than necessary. It was intentional in the design, because I wanted total consistency with a single backup for the system. It worked great for years. But as we got more and more clients, it really should have been migrated to the above design to make sure our backups could be taken / restored faster.
- sgt 1mo agoSo the original advice actually still stands. You can still start off with Posgres, store it in BYTEA columns, and then work on a plan to use an object storage as you grow. S3 may not be possible and you will be evaluating other options like Minio
- KronisLV 2mo ago> PostgreSQL allowed us to use a fulltext search plugin to do everything in one system. No need to sync any data. No need to maintain and run two systems. It just worked and made us smile (after some tweaks of course). Simplicity. I found MariaDB to be wonderfully simple to use for somewhat casual use cases: https://mariadb.com/docs/server/ha-and-performance/optimization-and-tuning/optimization-and-indexes/full-text-indexes/full-text-index-overview https://mariadb.com/docs/server/ha-and-performance/optimizat... and still reach for it in some personal projects, however the whole growing MySQL incompatibility is a big issue if the tech you use only officially supports MySQL and you can't (easily) get MariaDB specific DB drivers. Personally, one of the best things about PostgreSQL is transactional DDL, every DB should support it. Also they handle JSON pretty nicely (though I'd prefer not to store data like that unless necessary) alongside excellent plugins like pgvector and PostGIS. On the other hand, for things like queues, or even any sort of blob storage, I'd look at things like RabbitMQ or Garage (S3 compatible). Sometimes specialized software is nice for keeping things logically separated. I maintain that it's good to be able to divide your stack up by mechanisms/concerns (rather than business domain necessarily).
- throwatdem12311 2mo agoFunny. I was just this joking this morning with a colleague about using Postgres for everything. Considering adding mongo for unstructured data? Just use postgres jsonb. Building a search index? Postgres is fine too. Considering using redis for fragment caching? Just use an unlogged table in postgres with key value columns. Need pub/sub? Well just use postgres listen/notify. Using postgres for everything has served me very well.
- codegeek 2mo agoThese types of articles needed to be written because we have gone way too much in the other direction. The issue is that people use too many tools prematurely when they are not needed at their stage. So yea, in most cases, you are probably better off just with Postgres. I m a culprit of this myself so I wouldn't say that I know better. It is just too tempting to setup too many tools to feel cooler or feeling that "we must use elasticsearch as no one does search in db".
- FLeXMurphy 2mo agoI'm waiting for the followup contrarian shitpost: "Firebird for Everything".
- mrkeen 2mo ago> My tip: Start with PostgreSQL as a queueing system. Only when that does no longer perform well switch to other systems like Kafka, RabbitMQ or SQS. My tip: store your company's source code on a samba file server. Only when that no longer performs well, switch to other systems like Git.
- mrkeen 1mo agoHehe, not too popular an idea is it? Maybe there's some characteristic about a version control system that makes it qualitatively different from a file store. Maybe it has nothing to do with size or number of customers!
- jtwaleson 2mo agoAt Comper we have a very hot key-value store for annotating git data. We maintain a parallel git-blame data structure so we can do incremental "git blame -w -M -C -C". Typically a very expensive operation, but if you make it incremental, you can make it very cheap when new commits need to be analyzed. However, building the git blame tree is still pretty intensive for large repos. We currently use rocksdb with storage on the same node, and hit rocksdb 1000s of times per second during our analysis. About 20% writes, 80% reads. The issue is that we need to start scaling horizontally, for burstable workers and zero-downtime deployment. So we're thinking to offload to an external kv service instead of a local rocksdb. TiKV seems a good replacement, about 3-4x slower, but very scalable. Reading this article, I think a separate postgres cluster with unlogged tables might be a good idea. If anyone has some experience to share, let me know!
- CodesInChaos 1mo agoUnlogged tables still support MVCC and keep old versions of tuples around until they're garbage collected.
- reid_58 1mo ago[dead]
- robomartin 1mo agoYears ago I used PostgreSQL under Django to drive an industrial test and inspection robotic cell (which I also designed and built) at a major technology company. It worked very well. PostgreSQL maintained machine state, path planning, sensor readings, faults, operator input, etc. I wanted to see how far I could push that toolset. It worked surprisingly well. Django's capabilities meant such things as multi-user login pages, access controls and remote monitoring were very easy.
- socketcluster 1mo agoThis article seems like a reaction to DuckDB's surge in popularity. Having multiple DB engines to choose from is good and it often doesn't matter which one you use. One could make the same argument about DuckDB. Many database engines are multi-purpose. Though of course there are specific use cases where a different DB may be more appropriate... Anyway databases nowadays are a commodity. A sticky commodity but nonetheless they are replaceable; increasingly so in the age of AI where data migrations are easier than ever.
- gflo247 1mo ago[flagged]
- throwaway7783 1mo agoMy go-to has been replicas for each use case, with well defined semantics for replication lags. I'm working on something that does most of this seamlessly (transactional,search, columnar & time series, vectors and queues) without having to bother about extension management, replication setup or tuning. Hopefully there is some value in this - one click multipurpose postgres fleet.
- frollogaston 1mo agoI use Postgres for a lot of things where textbooks say not to, but not caching. I'm not going to do it with triggers. Maybe if it supported TTL properly, even then, probably don't want to think about whether caching will bog down the rest of the DB.
- TheCapeGreek 1mo agoAnecdotally: The main caveat as someone who works on mostly average web CRUD apps, is that "Use PG/SQLite for everything" usually falls flat when the tools I use day to day don't support that use case super well or have rougher edges. If your framework/ORM/whatever of choice doesn't support the full feature set of that driver compared to Redis/ES/Whatever you're replacing, you'll find yourself going down rabbit holes doing workarounds instead of staying with the "happy path" and just using separate tech for what it's specialised in. If you already are doing most of these sorts of features by yourself instead of with frameworks, maybe it's fine, but this does start to feel like a time-to-release hindrance if you don't want to fiddle with the minutia.
- ezekiel68 1mo agoBona Fides: I learned c on with the K&R book on an Amiga (and transitioned to enterprise software engineering from there). This seems like one more "When all you have is a hammer, everything looks like a nail" take. I agree with the other commenters who advocate for best-of-breed (e.g. Kafka, etc. for a message queue). PS I freakin love PostgreSQL as a relational (or even a time-series or OLAP) DB.
- micw 1mo agoTetris on postgres? Not doom? So not a candidate for everything! Just kidding. Of course there's doom for postgres: https://github.com/cedardb/DOOMQL https://github.com/cedardb/DOOMQL (pure SQL) and https://github.com/DreamNik/pg_doom https://github.com/DreamNik/pg_doom (extension). Oh and there's https://github.com/snaplet/postgres-wasm https://github.com/snaplet/postgres-wasm that allows to run everything else in postgres ^^
- alper 1mo agoFor full text search you could also use pg_search (Tantivy) which looks very cool. But at scale you probably don't want to manage a bunch of mission critical systems that were jacked into your database server. The database is slow? How do we monitor that? So I would definitely begin like this, but you need to have a plan to break all of these out sooner or later.
- yishaicohen 1mo agoThere's a missing essay: Postgres for nothing. A tiny store on SQLite/D1 with no admin UI is boring and it stays up. The day I need JSONB I'll know.
- oblio 1mo ago> A tiny store on SQLite/D1 with no admin UI is boring and it stays up. D1?
- erdaltoprak 1mo agohttps://www.cloudflare.com/products/d1/ https://www.cloudflare.com/products/d1/
- rco8786 1mo agoThe related posts at the end are...interesting. > Three Strike Dismissal in One-On-Ones > Keith Rabois recommends to dismiss your report when you feel bad about an upcoming one on one more than three times in a row. Interesting. …
- jroseattle 1mo ago> Events, queues and persistent logs are getting more and more important in today’s software systems. Systems like Kafka, RabbitMQ, SQS and others provide that functionality. But maintaining them is annoying, custom and you need the skillset. In tech stack choices, I prefer staying simple as long as feasible. That said, you also need to know and understand concepts at a thorough level. The above comment from the article suggests PG as a central server that simplifies event architecture. As if the "annoying, custom and needed skillset" around those specific alternatives are unnecessary baggage. If you know anything about queues, scaling, availability, access semantics, message formats and concepts such as delivery guarantees, you find out very quickly that the server which stores a queued message is not the high-order bit in that equation.
- sanderjd 1mo agoTo me, your concluding sentence cuts the opposite direction. The server is not the most important thing, and each new kind of server that exists in the system is an appreciable increase in maintenance burden. To me, taken together, this is an argument for waiting until you have a very concrete forcing function to introduce that new kind of server for this purpose. I think a good way to think of this is: What empirical metric will the introduction of kafka (or whatever you choose) move in a positive direction? Latency? Oncall burden? The amount of code you have to maintain? Do you have correctness or data integrity metrics that this change would register on? etc. This isn't at all intended as an unanswerable question. Many or most organizations will easily say "yes" that they expect some improvement on some metrics by adopting the "right" system for the job. But lots of other organizations are cargo culting "well we know this is the right way, so we should do it this way" long before any metrics they care about would demonstrate the improvement.
- jroseattle 1mo agoHeartily agree with the metrics for analysis, but this seems anchored around a focus on introduction of "something else" without regard firstly to correctness. > The server is not the most important thing, and each new kind of server that exists in the system is an appreciable increase in maintenance burden. Sure, but compared to what? It's right to consider complexity, but understanding tradeoffs requires depth. The OP's original premise was that learning all the things about specific queuing services was unnecessary chafe; that INSERTs, SELECTs and UPDATEs are all anyone needs. The maintenance burden sits with your producers & consumers (or publishers/subscribers, whatever your nomenclature...), whether you want it there or not. My learned experience is that as soon as you start moving messages that are beyond trivial and carry different operational characteristics, you're going to have to understand those deeper concepts anyway.
- spawrks 1mo agoOr... Just use files for everything until you actually need a database. ( I know I'm going to get heat for saying that )
- pojzon 1mo agoIts not a meme. Postgres does all of those things good enough to be always considered the first pick. I love Postgres. Im using it for majority of my career in IT. It never failed me, even on very big scale. Amazing piece of software.