Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
szarnyasg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
szarnyasg
4mo ago
Hello, DuckDB devrel here. First, thanks for the kind words :) Second, it's funny you should mention the 20-line awk script. I was making a very similar argument yesterday at the Ubuntu Summit: at some point, using shell scripts with G
2.
▲
by
szarnyasg
5mo ago
Yes, Quack resolves this problem. In particular, your client (likely a DuckDB instance) will talk to a remote DuckDB that both has access to the underlying storage and can also serve as the catalog itself.
3.
▲
by
szarnyasg
5mo ago
Well, we are really working on it: https://github.com/duckdb/ducklake/pull/1151 So you'll be able to test it in a few days.
4.
▲
by
szarnyasg
5mo ago
Hello, DuckDB DevRel here. Quack is independent from MotherDuck. MotherDuck has its own proprietary protocol, which has been around for years and it supports things like dual execution – see more here: https://duckdb.org/qua
5.
▲
DuckDB-Delta Grows Up: Writes, Unity Catalog and Time Travel
(duckdb.org)
3 points
by
szarnyasg
5mo ago
|
0 comments
6.
▲
by
szarnyasg
6mo ago
DuckDB devrel here. You are right. This was in the FAQ but I also added it to the DuckLake documentation's main page at https://ducklake.select/docs/stable/
7.
▲
by
szarnyasg
6mo ago
DuckLake does not use B+ trees and it handles fragmentation with techniques like partial files and compaction upon checkpointing. The developers of DuckLake talks about this here: https://youtu.be/7Su0aVzbb-U?t=689 (Disclai
8.
▲
by
szarnyasg
6mo ago
Hi, DuckDB DevRel here. To have concurrent read-write access to a database, you can use our DuckLake lakehouse format and coordinate concurrent access through a shared Postgres catalog. We released v1.0 yesterday: https://ducklak
9.
▲
by
szarnyasg
7mo ago
That's a difficult question and I would like to avoid giving a direct answer (because I co-lead a nonprofit benchmarking graph databases) but even knowing what you need for a graph database can be a tricky decision. See my FOSDEM 2025
10.
▲
by
szarnyasg
7mo ago
That's a good point. I re-ran the benchmark on two instances: - c8gd.4xlarge - this has a single 950 GB NVMe SSD. - c5ad.4xlarge - this has 2 x 300 GB disks, which I put in a RAID 0 array. There are no c6ad.4xlarge instances, so this i
11.
▲
by
szarnyasg
7mo ago
Indeed, it would have been interesting but I really wanted to get the blog post out on the launch day of the MacBook Neo and did not have the bandwidth to run additional cloud experiments. I ran TPC-DS SF300 now on the c6a.4xlarge. It turns
12.
▲
by
szarnyasg
7mo ago
You're right! I pushed an updated TL;DR block.
13.
▲
DuckDB 1.4.3 LTS with Native Windows ARM64 Support
(duckdb.org)
2 points
by
szarnyasg
10mo ago
|
0 comments
14.
▲
DuckLake 0.3 with Iceberg Interoperability and Geometry Support
(ducklake.select)
2 points
by
szarnyasg
1y ago
|
0 comments
15.
▲
by
szarnyasg
1y ago
Yes - updates on existing rows are supported. (I work at DuckDB Labs.)
16.
▲
by
szarnyasg
1y ago
Great! > About the COPY statement, it means we can drop Parquet files ourselves in the blob storage ? Dropping the Parquet files on the blob storage will not work – you have to COPY them through DuckLake so that the catalog databases is
17.
▲
by
szarnyasg
1y ago
The YouTube video “Apache Iceberg: What It Is and Why Everyone’s Talking About It” by Tim Berglund explains data lakes really well in the opening minutes: https://www.youtube.com/watch?v=TsmhRZElPvM
18.
▲
by
szarnyasg
1y ago
Yes, you can use standard SQL constructs such as INSERT statements and COPY to load data into DuckLake. (diclaimer: I work at DuckDB Labs)
19.
▲
by
szarnyasg
1y ago
AWS started offering local SSD storage up to 2 TB in 2012 (HI1 instance type) and in late 2013 this went up to 6.4 TB (I2 instance type). While these amounts don't cover all customers, plenty of data fits on these machines. But the sof
20.
▲
by
szarnyasg
2y ago
Hi, DuckDB devrel here. DuckDB is an analytical SQL database in the form factor of SQLite (i.e., in-process). This quadrant summarizes its space in the landscape: https://blobs.duckdb.org/slides/goto-amsterdam-2024-duck
21.
▲
by
szarnyasg
2y ago
I'm a co-author of the blog post. I agree that the wording was confusing – apologies for the confusion. I added a note at the end: > The repository does not contain the source code for the frontend, which is currently not available
22.
▲
by
szarnyasg
2y ago
I have observed the Makefile effect many times for LaTeX documents. Most researchers I worked with had a LaTeX file full of macros that they have been carrying from project to project for years. These were often inherited from more senior r
23.
▲
by
szarnyasg
2y ago
I am the author of the original post and I also wrote a followup blog post on it yesterday: https://szarnyasg.org/posts/duckdb-vs-coreutils/ Yes, if you break the file into parts with GNU Parallel, you can easily
24.
▲
by
szarnyasg
2y ago
Hi – DuckDB Labs devrel here. It's great that you find DuckDB useful! On the setup side, I agree that local (instance-attached) disks should be preferred but does EBS incur an IO fee? It incurs a significant latency for sure but it doe
25.
▲
DuckDB in Python in the Browser with Pyodide, PyScript, and JupyterLite
(duckdb.org)
2 points
by
szarnyasg
2y ago
|
0 comments
26.
▲
Spanner Graph: Graph databases reimagined
(cloud.google.com)
3 points
by
szarnyasg
2y ago
|
0 comments
27.
▲
by
szarnyasg
2y ago
DuckDB supports partial reading of Parquet files (also via HTTPS and S3) [1], so it can limit the scans to the required columns in the Parquet file. It can also perform filter pushdown, so querying data in S3 can be quite efficient. Disclai
28.
▲
by
szarnyasg
2y ago
I used Rete in my PhD work, where I designed incremental view maintenance techniques for the openCypher graph query language ( https://szarnyasg.github.io/phd/szarnyasg-phd-dissertation.p... ). Rete is an elegant algorit
29.
▲
Launching open-source language tools for ISO/IEC GQL
(ldbcouncil.org)
2 points
by
szarnyasg
2y ago
|
0 comments
30.
▲
by
szarnyasg
3y ago
Primary keys and foreign keys are rarely used in data science workloads. In our performance guide, we recommend users to avoid using primary keys unless they are absolutely necessary [1]. In workloads running in infrastructure components, k
More ›