Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
malisper
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
malisper
2mo ago
> A question on 20s postgresql time - It does not look like you are accounting for reading data from disk I choose the data size so that it would fit in memory on the machine I was testing on. fwiw, there's still a ton of overhead P
32.
▲
by
malisper
2mo ago
I want to build the best database possible. While Postgres is great, there are a lot of core issues that have been around for over a decade. We're working hard to get pgrust production-ready, and it will definitely be production-ready
33.
▲
by
malisper
2mo ago
I'll need to write up how the scheduler works at some point, but it's heavily based on these papers[0][1]. It solves two different problems. First, it lets us throttle resource-intensive queries. Second, it enables work stealing.
34.
▲
by
malisper
2mo ago
At least in terms of speed, we're much faster on clickbench: https://benchmark.clickhouse.com/#system=+_b|pnc|pgrs|gQ|saB...
35.
▲
by
malisper
2mo ago
pgrcolumnar is not the default storage method. Right now, it's exposed as a table access method. There's lots of design space for how to do this so I want to avoid pre-committing to anything
36.
▲
by
malisper
2mo ago
You can try it. We're happy to help you with it, but expect there to be issues to work through. You would want to do it for something non-critical
37.
▲
by
malisper
2mo ago
Author here. Let me know if you have any questions about the post or about pgrust. Let me take a shot at answering what I think will be the most common question: how can I trust pgrust? Our #1 priority right now is correctness. Over the pas
38.
▲
by
malisper
2mo ago
Author here. At least for databases, AGPL (or stricter) has become standard. The issue is it's so easy for megacorps (Amazon, Google, etc) to take a permissively licensed product and monetize it at the expense of the original standard.
39.
▲
by
malisper
2mo ago
Having tried all the coding harnesses, I find that using Pi is exactly like using Emacs. For anything you want to build you can ask your agent and it will build it. There's tons of existing code to help you configure it. At the same ti
40.
▲
Rebuilding Postgres for 300x faster analytics: batching, fusion, and SIMD
(malisper.me)
3 points
by
malisper
2mo ago
|
0 comments
41.
▲
by
malisper
2mo ago
Done! https://github.com/ClickHouse/ClickBench/pull/1163
42.
▲
by
malisper
2mo ago
> I think the postgresql maintainers don't claim to support moving a database from x86 to arm without a dump-and-reload I would be very surprised by that because that means replicating a database between the two platforms would lead
43.
▲
by
malisper
2mo ago
Right now I'm only doing very small simple functions. Kani[0] takes care of translating the code to an intermediate representation for me. It converts the Rust code and C code into a GOTO program[1] which verifiers can then run on top
44.
▲
by
malisper
2mo ago
Once it's stable, that's definitely something I'm going to look at doing
45.
▲
by
malisper
2mo ago
At least for ClickBench, the remaining bottleneck is memory bandwidth. I think that's mostly going to be solved by better data representations than SIMD. One of the biggest wins was our hash table implementation. Depending on the cardi
46.
▲
Pgrust v0.2: Now faster than Postgres and Clickhouse Latest
(github.com)
18 points
by
malisper
2mo ago
|
11 comments
47.
▲
by
malisper
2mo ago
Thank you for the kind words! My goal is to build the best database possible. I'm trying to imagine what Postgres would be like if it were built today. I've been able to make a bunch of big architectural changes that the Postgres
48.
▲
by
malisper
2mo ago
Yep, I did submit bug reports. This was on Tuesday. I don't see a public copy of the mailing list that has my bug reports yet. The four bugs were: 0) When parsing a macaddr[0], Postgres uses sscanf with %x. %x can wraparound. This mean
49.
▲
by
malisper
2mo ago
I recently came across a use case where formal methods were incredibly helpful. I've been rewriting Postgres in Rust and am currently focusing on correctness. The biggest challenge is that there's so much surface area to cover. Po
50.
▲
by
malisper
3mo ago
When I've dealt with this I've generally made sure the transactions are updating rows in a consistent order. You can do that by sorting the rows before you update them
51.
▲
by
malisper
3mo ago
> use set seqscan = off when testing your query plans esp when tables are empty or nearly so so you can see if indexes will be used when seq scans become less cheap How well does this work for you? I thought if you have _any_ index, Post
52.
▲
by
malisper
3mo ago
This is correct. We first used c2rust to translate the C into unsafe Rust. The generated Rust code had one crate per C compilation unit. We then took the unsafe unidiomatic Rust and one crate at a time converted it to safe Rust.
53.
▲
by
malisper
3mo ago
Actually the inverse. I initially gave claude an outline of what I wanted, had it do some research into how to write idiomatic rust, and then had it draft a series of skills to do the work. I would then try out the skills, audit the results
54.
▲
by
malisper
3mo ago
Nothing major yet. Once I wrap up the performance work I'm doing I'll start looking at the best way to go about testing. I suspect there's a lot of novel things you can do with agents.
55.
▲
by
malisper
3mo ago
Thank you! I'm a big fan of your writing. I wanted to make sure there was something people could try out so I could show pgrust was real and not vaporware.
56.
▲
by
malisper
3mo ago
Thank you!
57.
▲
by
malisper
3mo ago
> Threads does not offer any major performance advantage This is very not true. When it comes to parallel queries, a process model adds a ton of overhead. You can't pass pointers between processes because the address space is differ
58.
▲
by
malisper
3mo ago
I spent a couple years managing a Postgres cluster with a petabyte of data. I wrote a couple blog posts from my work then[0][1]. I also wrote dozens of posts on the Postgres internals[2]. I've also given talks on how to generate fracta
59.
▲
by
malisper
3mo ago
That's something I eventually want to fix. The challenge is the storage format is so integral to Postgres that it's going to be a huge PITA to come up with a novel design. Right now OrioleDB is in beta. Once that becomes productio
60.
▲
by
malisper
3mo ago
My approach has changed throughout the course of this project. Throughout most of the project, we were working off of a c2rust translation of Postgres to Rust. That gave us a bunch of Rust code that was unsafe but did pass the Postgres test
More ›