3 ms·
Agreed on DuckDB, fantastic for working with most major data formats
by csjh 2y ago
Agreed on DuckDB, fantastic for working with most major data formats
- calderwoodra 2y agoTook your advice and tried DuckDB. Here's what I've got so far: ``` def _get_duck_db_arrow_results(s3_key): con = duckdb.connect(config={'threads': 1, 'memory_limit': '1GB'}) con.install_extension("aws") con.install_extension("httpfs") con.load_extension("aws") con.load_extension("httpfs") con.sql("CALL load_aws_credentials('hadrius-dev', set_region=true);") con.sql("CREATE SECRET (TYPE S3,PROVIDER CREDENTIAL_CHAIN);") results = con \ .execute(f"SELECT * FROM read_parquet('{s3_key}');") \ .fetch_record_batch(1024) for index, result in enumerate(results): print(index) return results ``` I ran the above on a 1.4gb parquet file and 15 min later, all of the results were printed at once. This suggests to me that the whole file was loaded loaded into memory at once.
- wild_egg 2y agoIt's been a long time since I've used python but that sounds like buffering in the library maybe? I use it from Go and it seems to behave differently. When I'm writing to postgres though I'm doing into entirely inside DuckDB with a `INSERT INTO ... SELECT ...` and that seems to stream it over.
- akdor1154 2y agoYou asked it to fetch a batch (15min) then iterated over the batch (all at once). To stream, fetch more batches. What ddb does to get the batches depends on hand wavey magic around available ram, and also the structure of the parquet.