3 ms·
Is the dataset somewhere accessible? Does anyone know more about the "1T challenge", or is it just the 1B challenge moved up a notch? Would be interesting to s
by shinypenguin 11mo ago
Is the dataset somewhere accessible? Does anyone know more about the "1T challenge", or is it just the 1B challenge moved up a notch?
Would be interesting to see if it would be possible to handle such data on one node, since the servers they are using are quite beefy.
- philbe77 11mo agoHi shinypenguin - the dataset and challenge are detailed here: https://github.com/coiled/1trc https://github.com/coiled/1trc The data is in a publicly accessible bucket, but the requester is responsible for any egress fees...
- shinypenguin 11mo agoHi, thank you for the link and quick response! :) Do you know if anyone attempted to run this on the least amount of hardware possible with reasonable processing times?
- philbe77 11mo agoYes - I also had GizmoSQL (a single-node DuckDB database engine) take the challenge - with very good performance (2 minutes for $0.10 in cloud compute cost): https://gizmodata.com/blog/gizmosql-one-trillion-row-challenge https://gizmodata.com/blog/gizmosql-one-trillion-row-challen...
- simonw 11mo agoI suggest linking to that from the article, it is a useful clarification.
- philbe77 11mo agoGood point - I'll update it...
- achabotl 11mo agoThe One Trillion Row Challenge was proposed by Coiled in 2024. https://docs.coiled.io/blog/1trc.html https://docs.coiled.io/blog/1trc.html