4 ms·
Correct, downloading this 40GB file could easily leads to huge s3 bills due to traffic. That's 1$ for a full download.
by mayli 5y ago
Correct, downloading this 40GB file could easily leads to huge s3 bills due to traffic. That's 1$ for a full download.
- Asdrubalini 5y agoI tried downloading the file directly from curl [0] but it seems like it blocks non-partial GET requests, so downloading it is not really straightaway. [0] https://s3.amazonaws.com/static.wiki/db/en.db https://s3.amazonaws.com/static.wiki/db/en.db
- llacb47 5y agocurl -v "https://s3.amazonaws.com/static.wiki/db/en.db" -H "Referer: http://static.wiki/" -H "Range: bytes=0-" Above works but for the sake of OP's wallet I would suggest that you do not download the entire database from his S3 bucket.
- infogulch 5y agoInstead download it from kaggle which is offering to host it for free: https://www.kaggle.com/segfall/markdownlike-wikipedia-dumps-in-sqlite https://www.kaggle.com/segfall/markdownlike-wikipedia-dumps-...
- banana_giraffe 5y ago$2, if my math is right. As pointed out elsewhere, there's some basic work being done to prevent everyone from grabbing the .db file, but it's easy enough to bypass. Also, they're probably seeing non-zero API call charges. S3 only charges $0.0004 per 1000 calls, but when you start to make lots of calls at scale, it can really add up. Still, for this site, it's probably not that bad. If my simple tests are a fair judge, it works out to something like $20 a month for 100k lookups. I've mentioned elsewhere on HN I used a similar technique with Python to query SQLite databases in S3. It works well, but there was an edge case on that project that resulted in millions of requests to S3 per invocation. It didn't break the bank, but it did push the project into the unprofitable range till we fixed that case.