12 ms·
For the record, I think this will be a disaster. There is far too much code that will get broken, largely silently, and much of it is not under our con
by eclipticplane 3y ago
For the record, I think this will be a disaster. There is far too much
code that will get broken, largely silently, and much of it is not
under our control.
regards, tom lane
(via https://lwn.net/ml/pgsql-hackers/4178104.1685978307@sss.pgh.pa.us/ https://lwn.net/ml/pgsql-hackers/4178104.1685978307@sss.pgh....)
If Tom Lane says it will be a disaster, I believe it will be a disaster.
- idiomaticrust 3y agoHe is right. Such rewrites cause a lot of problems if your compiler doesn't help you with avoiding data races. But there is another way.
- hnarn 3y ago> But there is another way. Ok?
- mycall 3y agoMicrosoft SQL Server has SQLOS which is another way [0]. [0] https://www.thegeekdiary.com/what-is-sql-server-operating-system-sqlos/ https://www.thegeekdiary.com/what-is-sql-server-operating-sy...
- carstenhag 3y agoThe person probably implied that Postgres should switch to another toolchain that guarantees more things at compile time, so probably Rust.
- blincoln 3y agoIf the existing code is old-school enough to use thousands of global variables in a thread-unsafe way, seems like changing it enough to compile as safe Rust code would push the "non-trivial" envelope pretty far.
- bb88 3y agoYou can take a chunk of code and just rewrite it in Rust. You'll learn a lot quickly by this.
- steve_adams_86 3y agoIt’s sort of like the inverse of the Matrix when Neo learns kung fu. You realize that you actually don’t know how to program :)
- tylerhou 3y agoThe boundaries within database code are not clear. There are too many interlocking parts to take a nontrivial chunk and rewrite it Rust.
- chc 3y agoI think it's meant to imply the solution given in their username ("idiomatic Rust").
- lelanthran 3y ago> I think it's meant to imply the solution given in their username ("idiomatic Rust"). I think "Idiom: a tic (Rust)" can also fit if I squint hard enough and decide it looks like a definition from an online dictionary :-)
- avgcorrection 3y agoDon’t mind the gimmick gallery (username).
- deleted 3y ago[deleted]
- zilti 3y agoIndeed, Zig is a nice language for this
- gremlinsinc 3y agoMaybe a better option would be finding a team to create nugres, aka a fork for this and other experiments. So that mainline remains stable.
- mattashii 3y agoThere are several forks of PostgreSQL, in various levels of license, additional features and activity. However, maintaining a fork in addition to a main project is inherently more expensive than maintaining just a single project, so adding features to new major releases of the main project is generally preferred over forking every release into its own, newly named, project. After all, that is what we have major (feature) releases and stabalization windows (beta releases) for.
- j16sdiz 3y agoThis won't work well for a multiyear project.. Either you have to stall the release process, divide it into smaller parts or fork.
- datavirtue 3y agoThis should be considered a research effort, assuming it will be a complete rewrite. In light of that, you should not draw down resources from the established code base to work on it. Ignoring the above, first state the explicit requirements driving this change and let people weigh in on those. This sounds like a geeky dev itch.
- duped 3y agoThat's an awful message with the only sensible reply.
- abhibeckert 3y agoReminds me of PHP 6... For those who don't follow PHP closely - that version was an attempted refactor of the string implementation which essentially shut down nearly all work on PHP for a decade, stagnating the language until it became pretty terrible compared to other options. They finally gave up and started work on PHP 7 which uses the (perfectly good) PHP 5 strings. Ten years of wasted time by the best internal PHP developers crippled the project - I'm amazed it survived at all.
- progmetaldev 3y agoI've used PHP in the past (PHP 4 and 5), as well as some simple templated projects in PHP 7. I try to keep up on news with what is happening in the PHP world, and it's difficult because of the hate for the language. Is the solution to Unicode strings still to just use the "mb_*" functions? I got my real professional start using PHP, and have built even financial systems in the language (since ported to .NET 6 for my ease of maintenance, and better number handling). I'm still very interested in the language itself, in case I ever have the need to freelance or provide a solution to a client that can't afford what I can build in .NET (although to be honest, at this point I'm roughly able to code at the same speed in .NET as in PHP, but with the added type-safety, although I know PHP has really stepped up in providing this).
- Hayvok 3y agoI believe so - most (all?) string functions have an mb_ equivalent, for working on multibyte strings. Regular PHP strings are actually pretty great, since you can treat them like byte arrays. Fun fact: PHPs streaming API has an “in-memory” option and it’s… just a string under the hood. Just don’t forget to use multibyte functions when you’re handling things like user input.
- robomc 3y agoI have the "Professional PHP6" book which I feel like should be a collectors item or something. Weird book IMO, because it has a lot of content that's just about general software development, rather than anything to do with PHP specifically, or the theoretical PHP6 APIs in particular.
- kccqzy 3y agoI don't expect you or others to buy into any particular code change at this point, or to contribute time into it. Just to accept that it's a worthwhile goal. If the implementation turns out to be a disaster, then it won't be accepted, of course. But I'm optimistic. The reply is much more reasonable than this blanket assertion of a disaster.
- giraffe_lady 3y agoAs an outsider it doesn't sound like something a few people could spin off in a branch in a couple months and see how code review goes. They're talking about doing it over multiple (yearly?) releases. It seems like it'll take a lot of expert attention, which won't be available for other work and the changes themselves will impact all other ongoing work. I'm not trying to naysay it per se, bc again I don't have technical knowledge of this codebase. But that's exactly the sort of scenario that can cause a large project to splinter or stall for years. Talking about "the implementation" absent the context that would be necessary to create that implementation seems naively optimistic, or at worst irresponsible.
- lbriner 3y agoYou are talking about implementation, the OP was talking about raising the concept with interested parties and seeing whether it is worth even starting to think about it. They could fork, they could add threading to some sub systems and roll it out over several versions. I don't know enough about the code but, of course, it is a hard problem but the solution might be to build it from the ground up as a threaded system, using the skills learned over 30 years and taking the hit on the rebuild instead of reworking what is there. I am most interested because I didn't realise there was a performance problem in the first place.
- axman6 3y agoAm I going crazy, or has the obvious implementation of such a change been missed on people? If they were proposing taking a multi-threaded app and splitting it into a multi-process one, I would predict they would find a hell of a lot of unexpected or unknown implicit communication between threads, which would be a nightmare to untangle. Going the other way, there is an extremely well understood interface between all the processes which run in isolation: shared memory. Nearly by definition this must be well coordinated between the processes. So the first step in moving to a multi-threaded implementation would be to change nearly nothing about each process, and then just run each process in its own pthread, keeping all the shared memory ‘n all. You would expect performance to be about the same, maybe a little better with the reduces TLB churn, but the architecture is basically unchanged. At that point, you can start to look at what are more appropriate communication/synchronisation mechanisms now you’re working in the same address space. I just don’t understand why so many people seem to think this requires an enormous rewrite - having developed as a multi-process system means you’ve had to make so much of the problematic things explicit and control for them, and none of these threads would know anything at all about each other’s internals.
- kristiandupont 3y agoYeah. Without being familiar with the Postgres source, this seems to be what I call a "somersault problem": hard to break down into sub-goals. I have heard that the Postgres codebase is solid which makes it easier but it's still mature and highly complex. It doesn't sound feasible to me. [link redacted]
- seedless-sensat 3y agoThe original post does describe several sub-problems. The group could first chip away at global state, signals, libraries. They can do this before changing the process model in any way.
- kristiandupont 3y agoGood point.
- lenkite 3y agoFeel like the PostgreSQL Core Team should just build a new database from scratch using what they have learned from experience instead of attempting such a fundamental architectural migration. It would give them more freedom to change things also. Call it "postgendb" and provide a data migrator.
- rdevsrex 3y agoThat's a great idea. I've been considering whether or not to use Cockroach Db at work, and I love the fact that it's distributed from the get go. Why not work on something like that instead of changing something that works? Especially since they the process model really only runs into trouble on large systems.
- osigurdson 3y agoHeikki Linnakangas has a good understanding of Postgres as well however. We all want Postgres to be competitive with numbers of connections, don't we?