4 ms·
Unless, of course, you read wget documentation and realize most of that stuff is available as flags, you can log errors, you can use GNU parallel and a bunch of
by timonovici 11y ago
Unless, of course, you read wget documentation and realize most of that stuff is available as flags, you can log errors, you can use GNU parallel and a bunch of other specialized tools.
It still complexity in the end, but you just have to factor it all in - I bet you that 90% of those "big data" problems can fit in a small server's RAM. Most of the time is just people making shit up, so they'll have a job.
- halayli 11y agoI've read wget, and I am a cURL contributor as well. I know the capabilities of each very well. But when you have 1M+ urls to download and you care about proper error handling at a large scale, those tools will fall short. Not all errors need to be handled in the same way. Some need retries some don't depending on http response codes for example. Another problem you'll hit is how to make sure that the machines are saturated. How many jobs should be running at the current time depends on what's being downloaded and how much room you have to run additional downloads. Again, it all depends on what you need from the system and how much leeway you have.