3 ms·
This article outlines well the paradox that JITs require to be truly more efficient: if more of the target language is available to optimize, it'll get waaay mo
by chucke 3y ago
This article outlines well the paradox that JITs require to be truly more efficient: if more of the target language is available to optimize, it'll get waaay more optimized, compared to dropping down to the layer below and try to hand-stitch it.
Of course, there is massive overhead in doing so. Just look at go, which had to rewrite practically everything already available in go, and must always require a native implementation (protobuf for example shares the underlying interface across ruby, python, php... but then has a full separate implementation in go, and java I think). And they have the budget for it at least, Google won't let go die under the overhead it created for itself.
So definitely, write more ruby, enough of those "fast-C gem - rewritten as C extension", but still keep using low level libraries like libpq.
- grncdr 3y agoIt’s funny you mention libpq, because my first thought upon reading the article was “I wonder if a pure Ruby implementation of the postgres wire protocol could possibly lead to performance improvements?
- danmur 3y agoThis pure Python library claims quite fabulous performance: https://github.com/MagicStack/asyncpg https://github.com/MagicStack/asyncpg I believe it because that team have done lots of great stuff but I haven't used it, I just remembered thinking it was interesting the performance was so good. Not sure how related it is to running on the asyncio loop (or which loop they used for benchmarks).
- chucke 3y agoI'd assume that the benchmark is showcasing the gains of the async model, more than making a point that native python is faster than C. The issue is ultimately: how much of the functionality available through libpq is not yet backported to asyncpg (i imagine it's nonzero, but Idon'tknow the answer). Which is why the pragmatic in me still prefer a middle layer like libpq in between.
- whizzter 3y agoIirc NodeJS started with a libpq binding but it was moved to pure-JS since as in the article, overhead of going in and out of JIT-land penalized it (as well as lacking any object optimizations).
- chucke 3y agoDo you mean this one: https://node-postgres.com/features/native https://node-postgres.com/features/native ? I don't know much about the node ecosystem to judge how many use this instead of libpq bindings, but judging by the number of features mentioned in that page as incompatible, it doesn't look like a free lunch.
- chucke 3y agoProbably, but you'd have to reimplement a lot of things you get for free out of libpq (prepared statements, pooler support, other things I'm forgetting about). But fwiw, there's already this: https://github.com/mneumann/postgres-pr https://github.com/mneumann/postgres-pr
- vidarh 3y agoI'm just toying with a Ruby terminal talking direct to X (no X lib) running a pure Ruby TrueType renderer (and running a pure-Ruby editor in it...). Ruby is still not "fast" (but I haven't tested it with yjit yet). I put up with it because I know I can make it significantly faster later (and the quick and dirty first approximation is memoize what I can, like glyph shapes) As it turns out for a whole lot of things Ruby is slow for people because one of the nice things about programming in Ruby a lot of the time is expressing the problem the more readable way first rather than paying attention to performance. When people are aware, it makes it easy to iterate fast and fix performance later. The problem is a lot of the time that leads to pathological cases of people not thinking this through at all. You'll notice most of his performance increase came from writing better Ruby. Clearly the original Graphql parser was nowhere near optimal. It's worth keeping in mind that writing a C extension ought to be a last resort after making the Ruby as fast as you can first, because often that means you don't need to. E.g. another of Aaron's recent articles (EDIT: [1]) is about speeding up a parser by cutting down on object creation, and one specific example he gave was to return just the token type from the lexer instead of [token_type, token_value]. The latter forced creation of an Array object to hold each token and a string object for the token value (though for fixed tokens you could avoid that by returning the same frozen string literal), and for a whole lot of tokens the parser had no interest in the token string (e.g. if you see :lparen, or :rparen, getting "(" and ")" is entirely uninteresting). When people run into slow Ruby code like that it's tempting to resort to C right away, rather than understand why their Ruby is slow. I love that yjit makes it easier to get to a point where you don't need to reach for C, though. EDIT: [1] https://tenderlovemaking.com/2023/09/02/fast-tokenizers-with-stringscanner.html https://tenderlovemaking.com/2023/09/02/fast-tokenizers-with... It's actually about the GraphQL parser mentioned in the original article.
- chucke 3y agoGreat point. Not just writing ruby, but writing more performant ruby. I found that we several "here's how to implemented this GoF pattern in ruby" to "thank" for a lot of the unoptimized code we see around.
- coldtea 3y ago>Of course, there is massive overhead in doing so. Just look at go, which had to rewrite practically everything already available in go, and must always require a native implementation (protobuf for example shares the underlying interface across ruby, python, php... but then has a full separate implementation in go, and java I think) Both Java and Go could just as well use FFI for those things "already available" as C. They don't because native (to the language) is better, and can be even faster (due to GC or lack of crossing the boundary). Python, PHP and co just take the easy way out. Plus, given their speed, if stuff like JSON and co were native-to-the-language they would be much slower than C extensions.
- chucke 3y agoThat's true, but I think that it is due to a few factors: JNI is considered a big penalty in the java, and treated as a "last resort" (libvips is an example); go community has been advocating against C bindings since the runtime was itself rewritten, and treated as an antipattern; protobuf and grpc teams are compromised of mostly googlers, which have more incentives to optimize java and go (it's known that a big chunk of their services are java and c++, and most new stuff is go). And the rest might as well just use the bindings, as it's probably easier to maintain for them, and low in ggwir priorities lidt. There's actually a pure ruby protobuf library (I believe square employees use and maintain it), but it probably won't ever be merged into google/protobuf for those reasons.
- Mawr 3y ago> Of course, there is massive overhead in doing so. Some concrete numbers for Go: • Calling C from Go: ~40-50ns [1] [2] • Calling Go from C: ~100-200ns [3] [1] "Cgo calls take about 40ns, about the same time encoding/json takes to parse a single digit integer. On my 20 core machine Cgo call performance scales with core count up to about 16 cores, after which some known contention issues slow things down." (https://shane.ai/posts/cgo-performance-in-go1.21 https://shane.ai/posts/cgo-performance-in-go1.21) [2] "In response to the "cgo is slow" questions, this shows that cgo calls are about 50ns on my four-year-old x86 MacBook Pro. Is that fast or slow? It depends on what the cgo call is doing. If it's executing a single add instruction, an extra 50ns is slow. If it's doing something more substantial, an extra 50ns may be nothing at all." – rsc (https://news.ycombinator.com/item?id=36006347 https://news.ycombinator.com/item?id=36006347) [3] "Calls from C to Go on threads created in C require some setup to prepare for Go execution. On Unix platforms, this setup is now preserved across multiple calls from the same thread. This significantly reduces the overhead of subsequent C to Go calls from ~1-3 microseconds per call to ~100-200 nanoseconds per call." (https://go.dev/doc/go1.21 https://go.dev/doc/go1.21)