4 ms·
TIL. Usually I just went with SELECT * FROM 'folder/prefix*.parquet'
by ies7 3y ago
TIL. Usually I just went with SELECT * FROM 'folder/prefix*.parquet'
- BohuTANG 3y agoFor http(s) remote file, we can support glob pattern, for example: https://<thehfurl>/resolve/main/data/0000[00-55].parquet https://<thehfurl>/resolve/main/data/0000[00-55].parquet' Databend supports this pattern: https://databend.rs/doc/load-data/load/http#loading-with-glob-patterns https://databend.rs/doc/load-data/load/http#loading-with-glo...
- simonw 3y agoThat works for files on disk but not for files fetched via HTTP - though apparently DuckDB can do that for some situations, eg if they are in an S3 bucket that it can list files in.
- pests 3y agoYep - HTTP has no native support for file listings. In the old days it would be served at the index of a url path if no actual file was available - but that was always a feature of the http server and not anything unique to HTTP. Protocols like S3 and FTP have listings built in.
- booi 3y agoHTTP doesn't but I think you can in WebDAV, HTTP's long lost "file management" extension which actually has great support from most HTTP servers
- chrisjc 3y agoActually... (sorry, not picking on you Simon! Awesome post and I just love reading and talking about this stuff) With duckdb running on python you can register your own file-system adapters. This means that you can do things like intercept globbing, transform urls/path or physically getting files. This means that you could inject whatever listing/lookup that might be needed for your read_{format}() table function use case. http://duckdb.org/docs/guides/python/filesystems http://duckdb.org/docs/guides/python/filesystems However, it is true that there is no standard way to glob over HTTP unless you're using something like S3.
- ies7 3y agoYeah my bad. I only used it with S3 and somehow my brain think http will be the same.