8 ms·
Updates to the H2O.ai db-benchmark
- wenc 3y agoFantastic job by the DuckDB team. I’ve been using it for the past year to query 100s of GBs of Parquet files with complex analytic queries involving multiple levels of aggregations, joins and window functions and it all works and works fast. And I do all this from Jupyter Notebook. It’s actually faster than AWS Athena for me.
- ngrilly 3y agoThat’s great! Where are the Parquet files stored, and where is running DuckDB?
- marginalia_nu 3y agoYeah I've started using it with my search engine as well. It's fantastic how versatile it is for data all manner of data munging. Just the other day I used it to transform an unordered 60 GB CSV file with links and texts into a 3 GB parquet file that's so fast I can create a projection for the relevant data of each partition in like a minute (which then fits in memory). It has some minor stability issues so I'm not sure I'd build a full blown application on top of it, but for data transformation tasks it's amazing.
- LunaSea 3y agoAre you not getting OutOfMemory erros? We are in the same scenario (querying lots of Parquet files in S3) and we noticed that DuckDB quickly crashes with OOMs in environments with a few gigs of RAM. Setting the memory limit setting or the disk swap file has not worked.
- ayhanfuat 3y agoI encountered the same issue. Polars' memory usage was much lower for the datasets I tried.
- LunaSea 3y agoProblems I encountered with polar was that the immutability meant that copies of the dataframe were generated and not cleaned up fast enough.
- wenc 3y agoNot since 0.9.0. There’s been a lot of work on out of core stuff especially with big joins. It’s in the release notes. Also if you set a temp folder it will spill to disk (not set by default). I used to have to chunk my data to avoid OOMs but I haven’t had to do that. That said there are a few more out of core strategies on the roadmap that have not yet been implemented. If you still get OOMs, chunking your data will help. Also consider that few gigs of RAM might not be large enough for your workload. Out of core strategies can only do so much. https://duckdb.org/2023/09/26/announcing-duckdb-090.html https://duckdb.org/2023/09/26/announcing-duckdb-090.html
- riku_iki 3y ago> If you still get OOMs, chunking your data will help. what does chunking data mean here?..
- wenc 3y agoBreaking it up. Instead of running a query for the entire year, run it month by month and stitch the final results back together. Or if you have a unique string ID, calculate an integer hash using hash(ID) % 50 to get 50 chunks which you can process separately without OOMing. A basic assumption is that all the chunks are independent of each other. Chunking is essentially temporary partitioning to fit your processing limitations.
- cmollis 3y agothis part is confusing to me in the doc.. I assume that you're using the httpfs (S3) extensions and perhaps doing scanning of the parquet files (which I think is actually streamed.. e.g. querying for a specific column values in a series of parquet files). We have a huge data set of hive-partitioned parquet files in s3 (e.g. /customerid/year/month/<series of parquet files>). Can i just scan these files using the glob pattern to retrieve data like I can with Athena? The extension doc seems to indicate that I can (from the doc: SELECT * FROM read_parquet('s3://bucket/*/file.parquet', HIVE_PARTITIONING = 1) where year=2013;) Or do I need to know which parquet files I'm looking for in S3 and bring them down to work on locally? If it's the former, then it seems equivalent to Athena..
- wenc 3y agoNo you can definitely use globs in DuckDB. And no you don’t have to know the exact parquet file. You would treat the Hive partitioned data as a single dataset and DuckDB will scan it automatically. (Partition elimination, predicate pushdown etc all done automatically) https://duckdb.org/docs/data/partitioning/hive_partitioning https://duckdb.org/docs/data/partitioning/hive_partitioning
- cmollis 3y agook.. thanks.. I'll try it out. I can think of few use-case that we have where this might be a good alternative to athena.
- datadeft 3y agoExactly the same experience here. I am thinking about getting rid off everything else in our data infra as soon as it hits 1.0.
- treesciencebot 3y agoWould be curious how the performance compares to DataFusion[0] as one of the top contenders to DuckDB on this area (albeit they being different in a lot of parts, I find it one of the closest compared to all others). ClickBench (from ClickHouse) has some benchmarks[1] where it can be compared, but am not super sure how up to date it is. At least a while back, they were majorly out of date and haven't looked too closely on whether they are keeping it fair for everyone else :) [0]: https://github.com/apache/arrow-datafusion https://github.com/apache/arrow-datafusion [1]: https://benchmark.clickhouse.com https://benchmark.clickhouse.com
- jabart 3y agoLooks like a recent PR bumped benchmark.clickhouse.com to DuckDB v0.9 on the 3rd. https://github.com/ClickHouse/ClickBench/pull/141 https://github.com/ClickHouse/ClickBench/pull/141
- leicmi 3y agoA paper on DataFusion is in progress[0]. The draft[1] includes a comparison to DuckDB and preliminary benchmark results. [0]: https://github.com/apache/arrow-datafusion/issues/6782 https://github.com/apache/arrow-datafusion/issues/6782 [1]: https://www.overleaf.com/read/qjhrxqhgksvr https://www.overleaf.com/read/qjhrxqhgksvr
- riku_iki 3y agoWhy do you run benchmarks on such small datasets? It is very hard to judge performance..
- sanderjd 3y agoLooks like DataFusion is included in most of the results in the article?
- treesciencebot 3y agoYou are right! Seems like it is not text-addressable which is why my ctrl+f searches failed.
- esafak 3y agoReminder that you can use Fugue as a unified API and swap out back-ends, including DuckDB: https://fugue-tutorials.readthedocs.io/tutorials/integrations/backends/duckdb.html https://fugue-tutorials.readthedocs.io/tutorials/integration...
- rovr138 3y agoFugue, > Fugue provides an easier interface to using distributed compute effectively and accelerates big data projects. It does this by minimizing the amount of code you need to write, in addition to taking care of tricks and optimizations that lead to more efficient execution on distrubted compute. Fugue ports Python, Pandas, and SQL code to Spark, Dask, and Ray.
- sroerick 3y agoI really love DuckDB for one-offs and analytics and I wonder if anybody here has experience using with medium size data. I still seem to run into the workflow problem where data has to be in proximity to compute in order to function. If I need to run joins on a 5-10 GB parquet / table, unless I have that sitting locally, the performance bottleneck is not the database. I still find myself reaching to Databricks / Spark for most tasks for this reason. I suppose this is what Motherduck is trying to solve? But it just doesn't feel like it's quite there yet for me. Anybody who is better at this stuff than me have thoughts?
- pjot 3y agoI really like Motherducks hybrid execution model. My DS colleagues love abusing our data warehouse - bringing the data down locally to pound on is a win-win
- vgt 3y agoHi, head of Produck at MotherDuck here. Yes, we are indeed a good use case for this. For one, we built a fully-fledged managed storage system on top of DuckDB, with better performance and caching and the like. Two, we're going to be pretty good at reading from S3 because we've optimized that path. Three, our storage has sharing/IAM and is about to have things like zero-copy clone and time travel. Happy to answer any Qs.
- cube2222 3y agoIt would be interesting to see some more benchmarks of e.g. querying multiple files from S3, and how that evolved across versions. When I checked at 0.7.1, when working with ~90 S3 parquet objects (x0000 rows each, so not too many) it was 25-50% faster to first download them in Go and then query them, rather than using the DuckDB S3 extension with those objects directly (the whole execution ran on the order of a couple hundred milliseconds).
- wenc 3y agoYou definitely pay a performance penalty on S3 (S3 is high throughput but high latency storage) so not optimal for (any) database use cases. Local disk will always be faster if you can swing that. It’s not a DuckDB specific issue (although there’s headroom for improvement — I don’t think DuckDB’s S3 connector is highly optimized). It’s S3.
- cube2222 3y agoI might’ve been unclear, so to clarify: The overhead of fetching from S3 via a naive Go implementation (goroutine per object) to disk and then running duckdb on that was lower than using duckdb end-to-end. I was measuring the S3 overhead in both cases.
- wenc 3y agoNo I got you. Like I said S3 is a high throughput high latency storage. When you fetch the S3 object to disk, that’s a high throughput operation and S3 excels at that. Once on disk DuckDB can operate at low latency. If you run DuckDB end to end as a database engine on S3, it has to do partial reads on parquet on S3 etc. and has to deal with S3 latencies and it can end up being slower than what you described above. For long running operations where I can chunk the data, I often copy chunks to local disk before running DuckDB. It’s a lot faster than running DuckDB directly on S3. The downside is I need enough disk space.
- theLiminator 3y ago
- timost 3y agoInterestingly, Dask runs out of memory on many of the tasks of the benchmark.
- pgorczak 3y agoI had similar issues with the Dask scheduler a few months ago. The docs say it encourages depth first behavior in the computation graph but in my case it kept running out of memory on a large ETL task by first trying to load all the input files into memory before moving on to the next stage.
- indeedmug 3y agoI am impressed that Polar is close to DuckDB near the top. It's surprising that a Python library would often out perform everything but DuckDB. DuckDB is very impressive but DataFrames and Python is too useful to give up on.
- deleted 3y ago[deleted]
- shpongled 3y agoPolars is written in Rust!
- wenc 3y agoDuckDB interoperates with polars dataframes easily. I see DuckDB as a SQL engine for dataframes. Any DuckDB result is easily converted to Pandas (by appending .df()) or Polars (by appending .pl()). The conversion to polars is instantaneous because it’s zero copy because it all goes through Arrow in-memory format. So I usually write complex queries in DuckDB SQL but if I need to manipulate it in polars I just convert it in my workflow midstream (only takes milliseconds) and then continue working with that in DuckDB. It’s seamless due to Apache Arrow. https://duckdb.org/docs/guides/python/polars.html https://duckdb.org/docs/guides/python/polars.html
- indeedmug 3y agoWow, what a cool workflow. I looks like the interop promise of Apache Arrow is real. It's a great thing when your computer works as fast as you think as opposed to sitting around waiting for queries to finish.
- theLiminator 3y agoI mean polars is great, but there's nothing fundamentally impossible about polars providing similar performance to DuckDB, polars is written in rust, and really a lazy dataframe just provides an alternative frontend (sql being another frontend). There's nothing in the architecture that would make it so that performance in one OLAP engine is fundamentally impossible to achieve in another.
- riku_iki 3y agoKinda odd that they have max 50GB dataset and run it on machine with 160GB ram, so no testing of out of memory capabilities.
- sega_sai 3y agoI was/am a fan of duckdb, but I recently discovered a bug in 0.9.1 where a fairly innocuous query was silently returning wrong results (issue 9399 on github). That made me much less confident about duckdb and how well tested it is. Maybe it was a one off, but with postgresql for example I don't think I personally encountered cases of simply incorrect query results.
- datarecipes 3y agoJust had a look (https://github.com/duckdb/duckdb/issues/9399 https://github.com/duckdb/duckdb/issues/9399). Yeah it's worrying that such a trivial query returned incorrect results - but credit to the Devs for getting it fixed quickly. To my knowledge the only databases that can be described as "military-grade" in terms of testing are SQLite and Postgres.
- nyanpasu64 3y agoApparently DuckDB requires your real-life name to file an online bug report, bucking every norm of online handles for communication, as well as enabling doxxers and stalkers to find and trace people in real life.
- Cthulhu_ 3y agoIt's the same with a lot of open source contributions; those need to do so for legal and copyright reasons. If you're afraid of doxxing and / or stalking though, at least you have the choice to not contribute. You can still post somewhere else and ask someone else to make the report for you if need be.
- monsieurbanana 3y agoI'm not aware of any other open source project that required a real name just to file a bug
- sega_sai 3y agoYes, that was a surprising requirement when submitting a bug report. (I understand patches may need to be like that due to copyright issues)
- jcuenod 3y agoWhy do so many of the 0.5GB clickhouse benchmarks fail?
- spapas82 3y agoIf you have some data in postgresql and want to query it with duckdb (really fast) you can try extracting the data to a parquet file; this file can then be queried from duckdb with incredible speed. I've written a small program in python that reads from postgresql and exports to parquet for anybody that wanna try it https://github.com/spapas/pg-parquet-py#why https://github.com/spapas/pg-parquet-py#why
- chrisjc 3y agoThere's also the option to use the DuckDB PostgreSQL scanner https://duckdb.org/docs/archive/0.9.1/extensions/postgres_scanner https://duckdb.org/docs/archive/0.9.1/extensions/postgres_sc...
- jcuenod 3y agoReally interesting to compare to Clickhouse's benchmark, when you filter out the non-comparable results. The TLDR is that their benchmark shows DuckDB winning a lot of the races: https://benchmark.clickhouse.com/#eyJzeXN0ZW0iOnsiQXRoZW5hIChwYXJ0aXRpb25lZCkiOmZhbHNlLCJBdGhlbmEgKHNpbmdsZSkiOmZhbHNlLCJBdXJvcmEgZm9yIE15U1FMIjpmYWxzZSwiQXVyb3JhIGZvciBQb3N0Z3JlU1FMIjpmYWxzZSwiQnlDb25pdHkiOmZhbHNlLCJCeXRlSG91c2UiOmZhbHNlLCJjaERCIjpmYWxzZSwiQ2l0dXMiOmZhbHNlLCJDbGlja0hvdXNlIENsb3VkIChhd3MpIjp0cnVlLCJDbGlja0hvdXNlIENsb3VkIChnY3ApIjp0cnVlLCJDbGlja0hvdXNlIChkYXRhIGxha2UsIHBhcnRpdGlvbmVkKSI6dHJ1ZSwiQ2xpY2tIb3VzZSAoUGFycXVldCwgcGFydGl0aW9uZWQpIjp0cnVlLCJDbGlja0hvdXNlIChQYXJxdWV0LCBzaW5nbGUpIjp0cnVlLCJDbGlja0hvdXNlICh3ZWIpIjp0cnVlLCJDbGlja0hvdXNlIjp0cnVlLCJDbGlja0hvdXNlICh0dW5lZCkiOnRydWUsIkNsaWNrSG91c2UgKHpzdGQpIjp0cnVlLCJDcmF0ZURCIjpmYWxzZSwiRGF0YWJlbmQiOmZhbHNlLCJEYXRhRnVzaW9uIChQYXJxdWV0LCBzaW5nbGUpIjpmYWxzZSwiQXBhY2hlIERvcmlzIjpmYWxzZSwiRHJ1aWQiOmZhbHNlLCJEdWNrREIgKFBhcnF1ZXQsIHBhcnRpdGlvbmVkKSI6ZmFsc2UsIkR1Y2tEQiI6dHJ1ZSwiRWxhc3RpY3NlYXJjaCI6ZmFsc2UsIkVsYXN0aWNzZWFyY2ggKHR1bmVkKSI6ZmFsc2UsIkdyZWVucGx1bSI6ZmFsc2UsIkhlYXZ5QUkiOmZhbHNlLCJIeWRyYSI6ZmFsc2UsIkluZm9icmlnaHQiOmZhbHNlLCJLaW5ldGljYSI6ZmFsc2UsIk1hcmlhREIgQ29sdW1uU3RvcmUiOmZhbHNlLCJNYXJpYURCIjpmYWxzZSwiTW9uZXREQiI6ZmFsc2UsIk1vbmdvREIiOmZhbHNlLCJNeVNRTCAoTXlJU0FNKSI6ZmFsc2UsIk15U1FMIjpmYWxzZSwiUGlub3QiOmZhbHNlLCJQb3N0Z3JlU1FMICh0dW5lZCkiOmZhbHNlLCJQb3N0Z3JlU1FMIjpmYWxzZSwiUXVlc3REQiAocGFydGl0aW9uZWQpIjpmYWxzZSwiUXVlc3REQiI6ZmFsc2UsIlJlZHNoaWZ0IjpmYWxzZSwiU2VsZWN0REIiOmZhbHNlLCJTaW5nbGVTdG9yZSI6ZmFsc2UsIlNub3dmbGFrZSI6ZmFsc2UsIlNRTGl0ZSI6ZmFsc2UsIlN0YXJSb2NrcyI6ZmFsc2UsIlRpbWVzY2FsZURCIChjb21wcmVzc2lvbikiOmZhbHNlLCJUaW1lc2NhbGVEQiI6ZmFsc2V9LCJ0eXBlIjp7InN0YXRlbGVzcyI6dHJ1ZSwibWFuYWdlZCI6dHJ1ZSwiSmF2YSI6dHJ1ZSwiY29sdW1uLW9yaWVudGVkIjp0cnVlLCJDKysiOnRydWUsIk15U1FMIGNvbXBhdGlibGUiOnRydWUsInJvdy1vcmllbnRlZCI6dHJ1ZSwiQyI6dHJ1ZSwiUG9zdGdyZVNRTCBjb21wYXRpYmxlIjp0cnVlLCJDbGlja0hvdXNlIGRlcml2YXRpdmUiOnRydWUsImVtYmVkZGVkIjp0cnVlLCJzZXJ2ZXJsZXNzIjp0cnVlLCJhd3MiOnRydWUsImdjcCI6dHJ1ZSwiUnVzdCI6dHJ1ZSwic2VhcmNoIjp0cnVlLCJkb2N1bWVudCI6dHJ1ZSwidGltZS1zZXJpZXMiOnRydWV9LCJtYWNoaW5lIjp7InNlcnZlcmxlc3MiOmZhbHNlLCIxNmFjdSI6ZmFsc2UsImM2YS40eGxhcmdlLCA1MDBnYiBncDIiOnRydWUsIkwiOmZhbHNlLCJNIjpmYWxzZSwiUyI6ZmFsc2UsIlhTIjpmYWxzZSwiYzZhLm1ldGFsLCA1MDBnYiBncDIiOmZhbHNlLCIxOTJHQiI6ZmFsc2UsIjI0R0IiOmZhbHNlLCIzNjBHQiI6ZmFsc2UsIjQ4R0IiOmZhbHNlLCI3MjBHQiI6ZmFsc2UsIjk2R0IiOmZhbHNlLCIxNDMwR0IiOmZhbHNlLCJkZXYiOmZhbHNlLCI3MDhHQiI6ZmFsc2UsImM1bi40eGxhcmdlLCA1MDBnYiBncDIiOmZhbHNlLCJjNS40eGxhcmdlLCA1MDBnYiBncDIiOmZhbHNlLCJtNWQuMjR4bGFyZ2UiOmZhbHNlLCJtNmkuMzJ4bGFyZ2UiOmZhbHNlLCJjNmEuNHhsYXJnZSwgMTUwMGdiIGdwMiI6ZmFsc2UsImRjMi44eGxhcmdlIjpmYWxzZSwicmEzLjE2eGxhcmdlIjpmYWxzZSwicmEzLjR4bGFyZ2UiOmZhbHNlLCJyYTMueGxwbHVzIjpmYWxzZSwiUzIiOmZhbHNlLCJTMjQiOmZhbHNlLCIyWEwiOmZhbHNlLCIzWEwiOmZhbHNlLCI0WEwiOmZhbHNlLCJYTCI6ZmFsc2V9LCJjbHVzdGVyX3NpemUiOnsiMSI6dHJ1ZSwiMiI6dHJ1ZSwiNCI6dHJ1ZSwiOCI6dHJ1ZSwiMTYiOnRydWUsIjMyIjp0cnVlLCI2NCI6dHJ1ZSwiMTI4Ijp0cnVlLCJzZXJ2ZXJsZXNzIjp0cnVlLCJkZWRpY2F0ZWQiOnRydWUsInVuZGVmaW5lZCI6dHJ1ZX0sIm1ldHJpYyI6ImhvdCIsInF1ZXJpZXMiOlt0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlLHRydWUsdHJ1ZSx0cnVlXX0= https://benchmark.clickhouse.com/#eyJzeXN0ZW0iOnsiQXRoZW5hIC...
- alamb 3y agoI do think it was important for duckdb to put out a new version of the results as the earlier version of that benchmark [1] went dormant with a very old version of duckdb with very bad performance, especially against polars. [1] https://h2oai.github.io/db-benchmark/ https://h2oai.github.io/db-benchmark/
- MrPowers 3y agoI think these benchmarks are great, but also quite misleading and should be updated: * the 1 billion row benchmarks are run on a single, uncompressed 50 GB CSV file. 50 GB should be stored in multiple files. * the benchmarks only show the query runtime once the data has been persisted in memory. They should also show how long it takes to persist the data in memory. If query_engine_A takes 5 mins to persist in memory & 10 seconds to run the query and query_engine_B takes 2 mins to persist in memory & 20 seconds to run the query, then the amount of time to persist the data is highly relevant. * benchmarks should also show results when the data isn't persisted in memory. * Using a Parquet file with column pruning would make a lot more sense than a huge CSV file. The groupby dataset has 9 columns and some of the queries only require 3 columns. Needlessly persisting 6 columns in memory is really misleading for some engines. * Seems like some of the engines have queries that are more optimized than others. Some have explicitly casted columns as int32 and presumably others are int64. The queries should be apples:apples across engines. * Some engines are parallel and lazy. "Running" some of these queries is hard because lazy engines don't want to do work unless they have to. The authors have forced some of these queries to run by persisting in memory, which is another step, so that should be investigated. * There are obvious missing query types like filtering and "compound queries" like filter, join, then aggregate. I like these benchmarks a lot and use the h2o datasets locally all the time, but the methodology really needs to be modernized. At the bottom you can see "Benchmark run took around 105.3 hours." This is way to slow and there are some obvious fixes that'll make the results more useful for the data community.
- jandrewrogers 3y agoWhy should 50GB be stored in multiple files?
- MrPowers 3y agoModern query engines are designed to read data in parallel because it's so much faster. The data could be stored in 50 different one-gig files that were read in parallel.
- 3y ago
- benrutter 3y agoI love the increased focus on benchmark testing recently, but I always find it a little weird to read for stuff like spark or dask. Those are written to offer scale over large data so have very different overheads and limits compared to something like duckdb. Seems odd to have them on the same chart. Also, side note but I'd love to see the performance impact of pandas/dask with pyarrow schemas.
- vgt 3y agoWe at MotherDuck are working closely with the DuckDB folks to deliver a serverless analytics service powered by DuckDB [0]. Currently in open Beta and driving towards GA. Our users are certainly recognizing how fast DuckDB is in the cloud. One of the reasons we exist is because DuckDB is meant to be a single-player database. MotherDuck is doing tons of heavy-lifting to turn it into a true multi-player data warehouse, so things like IAM/sharing, persistence, time travel, administration, the ecosystem and so forth. What's magical about MotherDuck is that virtually any DuckDB instance in the wild can connect to MotherDuck by simply running '.open motherduck:' [1], and suddenly you get all these aforementioned benefits. (head of produck at MotherDuck) [0] https://motherduck.com/ https://motherduck.com/ [1] https://motherduck.com/docs/getting-started/connect-query-from-python/installation-authentication#authenticating-to-motherduck https://motherduck.com/docs/getting-started/connect-query-fr...
- tristenharr 3y agoNice job DuckDB team, those are great performance improvements compared to a couple years ago. It’s neat that solutions like this are being offered and developed. At Hasura we’ve been working on a data-connector that’ll wrap around DuckDB. Curious what others would think about having a GraphQL layer on JSON/Parquet/CSV via DuckDB?
- dang 3y agoRecent and related: DuckDB 0.9 - https://news.ycombinator.com/item?id=37657736 https://news.ycombinator.com/item?id=37657736 - Sept 2023 (59 comments)
- kristianp 3y agoModeration note: the actual title of this article is "Updates to the H2O.ai db-benchmark!". It's not about duckdb's performance improvements per se, it's about the change to the aws instance type they're using to get fairer benchmark results by avoiding block storage and preventing noisy neighbors.
- vgt 3y agoFrom the article itself: "The team at DuckDB Labs has been hard at work improving the performance of the out-of-core hash aggregates and joins." Aside from that, since April DuckDB has vastly improved performance of a number of queries, including aggregates, joins, window functions. [0] [0] https://duckdb.org/2023/09/26/announcing-duckdb-090.html#core-system-improvements https://duckdb.org/2023/09/26/announcing-duckdb-090.html#cor...
- kristianp 3y agoFrom: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html > please use the original title, unless it is misleading or linkbait; don't editorialize. Just because they mention X in the article, doesn't mean you should change the title to X, even if X is nice, welcome news.
- vgt 3y agoThat sounds great and i will do that from now on. That said, the substance of my reply holds...
- dang 3y agoChanged now. Thanks! (Submitted title was "DuckDB performance improvements with the latest release")
- vorticalbox 3y agoIt's nice that they put the Tldr at the top, im not sure why but a lot of places, for some weird reason, put it at the end.