5 ms·
I know there are CT search services like crt.sh, but is it practical to download the raw data and search it locally? If the logs are append‐only, it feels like
by cubesnooper 5y ago
I know there are CT search services like crt.sh, but is it practical to download the raw data and search it locally? If the logs are append‐only, it feels like a perfect usecase for rsync.
- middleclick 5y agohttps://github.com/SSLMate/certspotter/ https://github.com/SSLMate/certspotter/ possibly?
- RL_Quine 5y agoCertspotter is basically just crawling for matches, it doesn't retain them.
- RL_Quine 5y agoIt's impractical due to the size of the data. You're talking hundreds of millions, or billions of certificates.
- hannob 5y agoBit more than 6 billion. And yes, it's quite challenging. It would be a nice service to the community if someone would host downloadable dumps of CT logs.
- bushbaba 5y ago100 billion * 2048 bytes per cert is just 205 terabytes. That’s not that impractical. Could easily be stored on gcs and s3 with requestor pays for the api and any associated egress charges. 205 terabytes at 2 cent per GB/month is just 50k/year.
- exikyut 5y agoAnd FTP is a perfectly viable and straightforward alternative to Dropbox. You just need to...
- metadat 5y agoOr purchase 50x4TB drives for $10k USD, or 100 if you want RAID-1. These can easily fit into 4-8U of rack space, at around $1,500 or less per rack U (don't buy new / expensive servers if all you need is high-IO NAS storage hosts for cephFS). This will have minimal recurring fees. But will cost you time to manage instead of paying $CLOUDVENDER $50k/year. Please evaluate the tradeoffs for your scenario but good God don't blindly default to the cloud without first thinking carefully about it. See also related: Today's HN top article is "Just say Yes to self-hosting" https://news.ycombinator.com/item?id=30781536 https://news.ycombinator.com/item?id=30781536
- toomuchtodo 5y agoIt's $12,300/year at Backblaze. https://www.backblaze.com/b2/cloud-storage-pricing.html https://www.backblaze.com/b2/cloud-storage-pricing.html
- sp332 5y agoAccording to https://groups.google.com/g/certificate-transparency/c/iU6SHclGcYI https://groups.google.com/g/certificate-transparency/c/iU6SH..., there is no easy direct download. But you can just use the get-entries endpoint repeatedly and/or in parallel to download however many you want.
- decodebytes 5y agoYes, you can use the certificate-transparency go code to pull down from the trillian API https://github.com/google/certificate-transparency-go/blob/master/client/getentries.go https://github.com/google/certificate-transparency-go/blob/m... You would need to know the index, or you could just iterate over a range
- westurner 5y agoDo you think that CT log data replication would be more secure and efficient if the CT logs were stored by a zero trust distributed application like a blockchain, instead of Merkle signatures in a database owned by one party (that's now discontinuing free indexing, at least)? From "Oak, a Free and Open Certificate Transparency Log" (LetsEncrypt 2019) https://news.ycombinator.com/item?id=19920002 https://news.ycombinator.com/item?id=19920002 : > Trillian is a centralized Merkle tree: it doesn't support native replication [...] According to the trillian README, trillian depends upon MySQL/MariaDB and thus internal/private replication is as good as the SQL replication model (which doesn't have a distributed consensus algorithm like e.g. paxos). And what about indexing and search queries at volume, again without replication? From "A future for SQL on the web" https://news.ycombinator.com/item?id=28158491 https://news.ycombinator.com/item?id=28158491 : > https://thegraph.com/docs/indexing https://thegraph.com/docs/indexing >> Indexers are node operators in The Graph Network that stake Graph Tokens (GRT) in order to provide indexing and query processing services. Indexers earn query fees and indexing rewards for their services. They also earn from a Rebate Pool that is shared with all network contributors proportional to their work, following the Cobbs-Douglas Rebate Function. >> GRT that is staked in the protocol is subject to a thawing period and can be slashed if Indexers are malicious and serve incorrect data to applications or if they index incorrectly. Indexers can also be delegated stake from Delegators, to contribute to the network. >> Indexers select subgraphs to index based on the subgraph’s curation signal, where Curators stake GRT in order to indicate which subgraphs are high-quality and should be prioritized. Consumers (eg. applications) can also set parameters for which Indexers process queries for their subgraphs and set preferences for query fee pricing.
- westurner 5y agoFWIW, here's how OSV affords search queries: https://github.com/google/osv#data-dumps https://github.com/google/osv#data-dumps > For convenience, these sources are aggregated and continuously exported to a GCS bucket maintained by OSV: gs://osv-vulnerabilities > This bucket contains individual entries of the format gs://osv-vulnerabilities/<ECOSYSTEM>/<ID>.json as well as a zip containing all vulnerabilities for each ecosystem at gs://osv-vulnerabilities/<ECOSYSTEM>/all.zip > E.g. for PyPI vulnerabilities: # Or download over HTTP via https://osv-vulnerabilities.storage.googleapis.com/PyPI/all.zip gsutil cp gs://osv-vulnerabilities/PyPI/all.zip Hopefully, with an incentivized Blockchain Indexing service and/or e.g. GCS buckets that you just always `cp` and then load locally and then query locally, we can find a solution for queries of the growing CT Certificate Transparency logs.