15 ms·
Download responsibly
- holowoodman 1y agoJust wait until some AI dudes decide it is time to train on maps...
- nativeit 1y agoI’m looking forward to visiting all of the fictional places it comes up with!
- M95D 1y agoMap slop? That's new!
- jbstack 1y agoAI models are trained relatively rarely, so it's unlikely this would be very noticeable among all the regular traffic. Just the occasional download-everything every few months.
- holowoodman 1y agoOne would think so. If AI bros were sensible, responsible and intelligent. However, the pratical evidence is to the contrary, AI companies are hammering every webserver out there, ignoring any kind of convention like robots.txt, re-downloading everything in pointlessly short intervals. Annoying everyone and killing services. Just a few recent examples from HN: https://news.ycombinator.com/item?id=45260793 https://news.ycombinator.com/item?id=45260793 https://news.ycombinator.com/item?id=45226206 https://news.ycombinator.com/item?id=45226206 https://news.ycombinator.com/item?id=45150919 https://news.ycombinator.com/item?id=45150919 https://news.ycombinator.com/item?id=42549624 https://news.ycombinator.com/item?id=42549624 https://news.ycombinator.com/item?id=43476337 https://news.ycombinator.com/item?id=43476337 https://news.ycombinator.com/item?id=35701565 https://news.ycombinator.com/item?id=35701565
- Waraqa 1y agoIMHO in the long term this will lead to a closed web where you are required to log-in to view any content.
- cadamsdotcom 1y agoDefinitely a use case for bittorrent.
- john_minsk 1y agoIf the data changes, how would a torrent client pick it up and download changes?
- hambro 1y agoLet the client curl latest.torrent from some central service and then download the big file through bittorrent.
- maeln 1y agoA lot of torrent client support various API to automatically collect torrent file. The most common is to simply use RSS.
- Klinky 1y agoPretty sure people used or even still use RSS for this.
- extraduder_ire 1y agoThere's a BEP for updatable torrents.
- Gigachad 1y agoSounds like someone people are downloading it in their CI pipelines. Probably unknowingly. This is why most services stopped allowing automated downloads for unauthenticated users. Make people sign up if they want a url they can `curl` and then either block or charge users who download too much.
- userbinator 1y agoI'd consider CI one of the worst massive wastes of computing resources invented, although I don't see how map data would be subject to the same sort of abusive downloading as libraries or other code.
- Gigachad 1y agoThis stuff tends to happen by accident. Some org has an app that automatically downloads the dataset if it's missing, helpful for local development. Then it gets loaded in to CI, and no one notices that it's downloading that dataset every single CI run.
- account42 1y agoAt some point wilful incompetence becomes malice. You really shouldn't allow network requests from your CI runners unless you have something that cannot be solved in another way (hint: you don't).
- mschuster91 1y agoCI itself doesn't have to be a waste. The problem is most people DGAF about caching.
- account42 1y agoYou don't need caching if your build can run entirely offline in the first place.
- stevage 1y ago
- rossant 1y agoCan't the server detect and prevent repeated downloads from the same IP, forcing users to act accordingly?
- jbstack 1y agoSee: "Also, when we block an IP range for abuse, innocent third parties can be affected." Although they refer to IP ranges, the same principle applies on a smaller scale to a single IP address: (1) dynamic IP addresses get reallocated, and (2) entire buildings (universities, libraries, hotels, etc.) might share a single IP address. Aside from accidentally affecting innocent users, you also open up the possibility of a DOS attack: the attacker just has to abuse the service from an IP address that he wants to deny access to.
- imiric 1y agoMore sophisticated client identification can be used to avoid that edge case, e.g. TLS fingerprints. They can be spoofed as well, but if the client is going through that much trouble, then they should be treated as hostile. In reality it's more likely that someone is doing this without realizing the impact they're having.
- imiric 1y agoIt could be slightly more sophisticated than that. Instead of outright blocking an entire IP range, set quotas for individual clients and throttle downloads exponentially. Add latency, cap the bandwidth, etc. Whoever is downloading 10,000 copies of the same file in 24 hours will notice when their 10th attempt slows down to a crawl.
- tlb 1y agoIt'll still suck for CI users. What you'll find is that occasionally someone else on the same CI server will have recently downloaded the file several times and when your job runs, your download will go slowly and you'll hit the CI server timeout.
- aitchnyu 1y agoDo they email heavy users? We used Nominatim free api for geocoding addresses in 2012 and our email was required parameter. They mailed us and asked us to cache results to reduce request rates.
- crimsoneer 1y agoI continue to be baffled the geofabrik folks remain the primary way to get a clean-ish OSM shapefile. Big XKCD "that one bloke holding up the internet" energy. Also, everyone go contribute/done to OSM.
- omcnoe 1y agoIt's beneficial to the wider community, and also supports their commercial interests (OSM consulting). Win-win.
- marklit 1y ago> primary way to get a clean-ish OSM shapefile Shapefiles shouldn't be what you're after, Parquet can almost always do a better job unless you need to either edit something or use really advanced geometry not yet supported in Parquet. Also, this is your best source for bulk OSM data: https://tech.marksblogg.com/overture-dec-2024-update.html https://tech.marksblogg.com/overture-dec-2024-update.html If you're using ArcGIS Pro, use this plugin: https://tech.marksblogg.com/overture-maps-esri-arcgis-pro.html https://tech.marksblogg.com/overture-maps-esri-arcgis-pro.ht...
- teekert 1y agoWhenever I read about such issues I always wonder why we all don’t make more use of BitTorrent. Why is it not the underlying protocol for much more stuff? Like container registries? Package repos, etc.
- vaylian 1y ago> Like container registries? Package repos, etc. I had the same thoughts for some time now. It would be really nice to distribute software and containers this way. A lot of people have the same data locally and we could just share it.
- maeln 1y agoI can imagine a few things : 1. BitTorrent has a bad rep. Most people still associate it with just illegal download. 2. It requires slightly more complex firewall rules, and asking the network admin to put them in place might raise some eyebrow for reason 1. On very restrictive network, they might not want to allow them at all due to the fact that it opens the door for, well, BitTorrent. 3. A BitTorrent client is more complicated than an HTTP client, and not installed on most company computer / ci pipeline (for lack of need, and again reason 1.). A lot of people just want to `curl` and be done with it. 4. A lot of people think they are required to seed, and for some reason that scare the hell of them. Overall, I think it is mostly 1 and the fact that you can just simply `curl` stuff and have everything working. I do sadden me that people do not understand how good of a file transfer protocol BT is and how it is underused. I do remember some video game client using BT for updates under the hood, and peertube use webtorrent, but BT is sadly not very popular.
- simonmales 1y agoAt least the planet download offers BitTorrent. https://planet.openstreetmap.org/ https://planet.openstreetmap.org/
- lostmsu 1y agoSo does Wikipedia https://meta.m.wikimedia.org/wiki/Data_dump_torrents https://meta.m.wikimedia.org/wiki/Data_dump_torrents Truly the last two open web titans.
- alluro2 1y agoPeople like Geofabrik are why we can (sometimes) have nice things, and I'm very thankful for them. Level of irresponsibility/cluelessness you can see from developers if you're hosting any kind of an API is astonishing, so downloads are not surprising at all...If someone, a couple of years back, told me things that I've now seen, I'd absolutely dismiss them as making stuff up and grossly exaggerating... However, on the same token, it's sometimes really surprising how API developers rarely ever think in terms of multiples of things - it's very often just endpoints to do actions on single entities, even if nature of use-case is almost never on that level - so you have no other way than to send 700 requests to do "one action".
- eleveriven 1y agoHonestly, both sides could use a little more empathy: clients need to respect shared infrastructure, and API devs need to think more like their users
- alias_neo 1y ago> Level of irresponsibility/cluelessness you can see from developers if you're hosting any kind of an API is astonishing This applies to anyone unskilled in a profession. I can assure you, we're not all out here hammering the shit out of any API we find. With the accessibility of programming to just about anybody, and particularly now with "vibe-coding" it's going to happen. Slap a 429 (Too Many Requests) in your response or something similar using a leaky-bucket algo and the junior dev/apprentice/vibe coder will soon learn what they're doing wrong. - A senior backend dev
- alluro2 1y agoThanks for the reply - I did not mean to rant, but, unfortunately, this is in context of a B2B service, and the other side are most commonly IT teams of customers. There are, of course, both very capable and professional people, and also kind people who are keen to react / learn, but we've also had situations where 429s result in complaints to their management how our API "doesn't work", "is unreliable" and then demanding refunds / threatening legal action etc... One example was sending 1.3M update requests a day to manage state of ~60 entities, that have a total of 3 possible relevant state transitions - a humble expectation would be several requests/day to update batches of entities.
- deleted 1y ago[deleted]
- Meneth 1y agoSome years ago I thought, no one would be stupid enough to download 100+ megabytes in their build script (which runs on CI whenever you push a commit). Then I learned about Docker.
- shim__ 1y agoThat's why I build project specific images in CI to be used in CI. Running apt-get every single time takes too damn long.
- Havoc 1y agoAlternatively you can use a local cache like AptCacherNg
- jve 1y agoWait, docker caches layers, you don't have to rebuild everything from scratch all the time... right?
- kevincox 1y agoIt does if you are building on the same host with preserved state and didn't clean it. There are lots of cases where people end up with with an empty docker repo at every CI run or regularly empty the repo because docker doesn't have any sort of intelligence space management (like LRU).
- the8472 1y agoTo get fine-grained caching you need to use cache-mounts, not just cache layers. But the cache export doesn't include cache mounts, therefore the docker github action doesn't export cache mounts to the CI cache. https://github.com/moby/buildkit/issues/1512 https://github.com/moby/buildkit/issues/1512
- eleveriven 1y agoIt's like, once it's in a container, people assume it's magic and free
- trklausss 1y agoI mean, at this point I wouldn't mind if they rate-limit downloads. A _single_ customer downloading the same file 10.000 times? Sorry, we need to provide for everyone, try again at some other point. It is free, yes, but there is no need to either abuse it or give as much resource for free as they can.
- k_bx 1y agoThis. Maybe they could actually make some infra money out of this. Make token-based free tier download, pay if you break it.
- stevage 1y ago>Just the other day, one user has managed to download almost 10,000 copies of the italy-latest.osm.pbf file in 24 hours! Whenever I have done something like that, it's usually because I'm writing a script that goes something like: 1. Download file 2. Unzip file 3. Process file I'm working on step 3, but I keep running the whole script because I haven't yet built a way to just do step 3. I've never done anything quite that egregious though. And these days I tend to be better at avoiding this situation, though I still commit smaller versions of this crime.
- xmprt 1y agoMy solution to this is to only download if the file doesn't exist. An additional bonus is that the script now runs much faster because it doesn't need to do any expensive networking/downloads.
- stanac 1y ago10,000 times a day is on average 8 times a second. No way someone has 8 fixes per second, this is more like someone wanted to download a new copy every day, or every hour but they messed up milliseconds config or something. Or it's simply malicious user. edit: bad math, it's 1 download every 8 seconds
- gblargg 1y agoWhen I do scripts like that I modify it to skip the download step and keep the old file around so I can test the rest without anything time-consuming.
- BenjiWiebe 1y agoAn easy way to avoid this is to have several scripts (bash for example): getfile.sh processdata.sh postresults.sh doall.sh And doall.sh consists of: ./getfile.sh ./processdata.sh ./postresults.sh
- cjs_ac 1y agoI have a funny feeling that the sort of people who do these things don't read these sorts of blog posts.
- globular-toast 1y agoAh, responsibility... The one thing we hate teaching and hate learning even more. Someone is probably downloading files in some automated pipeline. Nobody taught them that with great power (being able to write programs and run them on the internet) comes great responsibility. It's similar to how people drive while intoxicated or on the phone etc. It's all fun until you realise you have a responsibility.
- vgb2k18 1y agoSeems a perfect justification for using api keys. Unless I'm missing the nuance of this software model.
- kevincox 1y agoBut that raises the complexity of hosting this data immensely. From a file + nginx you now need active authentication, issuing keys, monitoring, rate limiting... Yes, this the the "right" solution but it is a huge pain and it would be nice if we could have nice things without needing to do all of this work. This is tragedy of the commons in action.
- bombcar 1y agoThere’s a cheapish middle ground - generate unique URLs for each downloaded, which basically embeds a UUID “API” key. You can paste it into a curl script, but now the endpoint can track it. So not example.com/file.tgz but example.com/FCKGW-RHQQ2-YXRKT-8TG6W-2B7Q8/file.tgz
- hyperdimension 1y agoYeah, but everyone knows that one. ;)
- account42 1y agoEveryone also knows the API keys that are used for requests from clients (apps/websites/etc.). ;)
- woodpeck 1y agoSpeaking as the person running it - introducing API keys would not be a big deal, we do this for a couple paid services already. But speaking as a person frequently wanting to download free stuff from somewhere, I absolutely hate having to "set up an account" just to download something once. I started that server well over a decade ago (long before I started the business that now houses it); the goal has always been first and foremost to make access to OSM data as straightforward as possible. I fear that having to register would deter many a legitimate user.
- unwind 1y agoMeta: the title could be clearer, e.g. "Download OSM Data Responsibly" would have helped me figure out the context faster as someone not familiar with the domain name shown.
- Oleh_h 1y agoGood to know. Thanks!
- sceptic123 1y agoWhy put a load of money into infra and none into simple mitigations like rate limiting to prevent the kind of issues they are complaining they need the new infra for?
- Joel_Mckay 1y agoSounds like someone has an internal deployment script with the origin mirror in the update URI. This kind of hammering is also common for folks prototyping container build recipes, so don't assume the person is necessarily intending anyone harm. Unfortunately, the long sessions needed for file servers do make them an easy target for DoS, and especially if supporting quality of life features like media fast-forward/seek features. Thus, a large file server normally does not share a nimble website or API server (CDNs still exist for a reason.) Free download API keys with a set data/time Quota and IP rate limit are almost always necessary (i.e. the connection slows down to 1kB/s after a set daily limit.) Ask anyone that runs a Tile server or media platform for the connection-count costs of free resources. It is a trade-off, but better than the email-me-a-link solution some firms deploy. Have a wonderful day =3
- Havoc 1y agoThey should rate limit it with a limit high enough that normal use doesn't hit it Appeals to responsibility aren't going to sink in to people that are clearly careless
- kevincox 1y agoBut rate-limiting public data is a huge pain. You can't really just have a static file anymore. Maybe you can configure your HTTP server to do IP based rate limiting but that is always ineffective (example public clouds where the downloader gets a new IP every time) or hits bystanders (a reasonable download the egresses out of the same IP or net block). So if you really want to do this you need to add API keys and authentication (even if it is free) to reliably track users. Even then you will have some users that find it easier to randomly pick from 100 API keys rather than properly cache the data.
- theshrike79 1y agoRate limit anonymous, unlimited if you provide an API key. This way you can identify WHO is doing bad things and disable said API key and/or notify them.
- eleveriven 1y agoThis is a good reminder that just because a server can handle heavy traffic doesn't mean it should be treated like a personal data firehose
- elAhmo 1y agoChances of those few people doing very large amounts of downloads reading this are quite small. Basic rate limits on IP level or some other simple fingerprinting would do a lot of good for those edge cases, as those folks are most likely not aware of this happening. Yes, IP rate limiting is not perfect, but if they have a way of identifying which user is downloading the whole planet every single day, that same user can be throttled as well.
- falconertc 1y agoSeems to me like the better action would be to implement rate-limiting, rather than complain when people use your resource in ways you don't expect. This is a solved problem.
- booleandilemma 1y agoI wonder if we're going to see more irresponsible software from people vibe coding shit together and running it without even knowing what it's doing.
- ranzhh 1y agoOh hey, it's me, the dude downloading italy-latest every 8 seconds! Maybe not, but I can't help but wonder if anybody on my team (I work for an Italian startup that leverages GeoFabrik quite a bit) might have been a bit too trigger happy with some containerisation experiments. I think we got banned from geofabrik a while ago, and to this day I have no clue what caused the ban; I'd love to be able to understand what it was in order to avoid it in the future. I've tried calling and e-mailing the contacts listed on geofabrik.de, to no avail. If anybody knows of another way to talk to them and get the ban sorted out, plus ideally discover what it was from us that triggered it, please let me know.
- woodpeck 1y agoHey there dude downloading italy-latest every 8 seconds, nice to hear from you. I don't think I saw an email from you at info@geofabrik, could you re-try?
- ranzhh 1y agoAbsolutely, a couple of them got bounced but one should've gone through. I'll retry as soon as I get into the office. Edit: I sent you the email, it got bounced again with a 550 Administrative Prohibition. Will try my university's account as well. Edit2: this one seems to have gone through, please let me know if you can't see it.
- deleted 1y ago[deleted]
- 1vuio0pswjnm7 1y ago"There have been individual clients downloading the exact same 20-GB file 100s of times per day, for several days in a row. (Just the other day, one user has managed to download almost 10,000 copies of the italy-latest.osm.pbf file in 24 hours!) Others download every single file we have on the server, every day." This sounds like problem rate-limiting would easily solve. What am I missing. The page claims almost 10,000 copies of same file were downloaded by the same user The server operator is able to count the number of downloads in a 24h period for an individual user but cannot or will not set a rate limit Why not Will the users mentioned above (a) read the operator's message on this web page and then (b) change their behaviour I would be bet against (a) and therefore (b) as well
- woodpeck 1y agoGeofabrik guy here. You are right - rate limiting is the way to go. It is not trivial though. We use an array of Squid proxies to serve stuff and Squid's built-in rate limiting only does IPv4. While most over-use comes from IPv4 clients it somehow feels stupid to do rate limiting on IPv4 and leave IPv6 wide open. What's more, such rate-limiting would always just be per-server which, again, somehow feels wrong when what one would want to have is limiting the sum of traffic for one client across all proxies... then again, maybe we'll go for the stupid IPv4-per-server-limit only since we're not up against some clever form of attack here but just against carelessness.
- 1vuio0pswjnm7 1y agoStick tables work with either IPv4 or IPv6
- Incipient 1y ago>one user has managed to download almost 10,000 copies of the italy-latest.osm.pbf file in 24 hours! Trash code most likely. Rate limit their IP. No one cares about people that do this kind of thing. If they're a VPN provider...then still rate limit them.
- mvdtnz 1y ago> If you want a large region (like Europe or North America) updated daily, use the excellent pyosmium-up-to-date program This is why naming matters. I never would have guessed in a million years that software named "pyosmium-up-to-date" does this function. Give it a better name and more people will use it.