7 ms·
A sysadmin's rant about feed readers and crawlers (2022)
- yapyap 2y agohttps://web.archive.org/web/20241205224611/http://rachelbythebay.com/w/2022/03/07/get/ https://web.archive.org/web/20241205224611/http://rachelbyth...
- croisillon 2y agofor a feed reader i built that year (2022) i was polling each feed every second day, except the ones which had a new item within 50 days -> every day
- flymasterv 2y agoMy strategy is to have an exponential backoff. I start a new feed set to query every 30 minutes, and if there’s no new post, I double the period. If there’s is a new post, I halve it. My reader goes through every feed every half hour, and randomizes which feeds it checks: a 1/4 chance for a 2 hour feed, a 1/48 chance for a 24 hour feed, etc.
- cellularmitosis 2y agoSeveral people have mentioned exponential back off. What upper limit would you suggest? Someone might not post for 6 months, then resume posting daily. You might miss those posts for months.
- croisillon 2y agoin this case polling every few days is more than reasonable
- flymasterv 2y agoI am not worried about missing posts, on my own reader. I specifically only grab the single newest post, so I may be missing posts by design. That said, I think 3 or 4 days seems reasonable.
- tomrod 2y agoI always appreciate Rachel's writings. I don't know much about her, but my takeaway is that she has worked at some of the hardest sysadmin jobs in the past few decades and writes to her experience super well.
- theshrike79 2y agoThis is a good lesson on being a good citizen of the Internet. It's easy to just curl a feed every second, but should you? (Of course not) Take it as a challenge to make your reader as fancy as possible, use every trick in the book to optimise how it fetches content. Analyse the patterns of releasing new content per feed and adjust the fetch frequency based on that. And if you're building a reader for distribution, don't let the user set a refresh interval that doesn't make sense.
- flir 2y agoLike writing lift control software. Minutes to learn, a life time to master. But I bet you can get 95% of the benefit with a simple exponential backoff scheme.
- preinheimer 2y agoI feel like polling twice a day (with the right etags) will beat exponential back offs. Exponential back off is great for a lot of problems, but irregularly updated blogs doesn’t seem like one of them.
- flir 2y agoThinking about it, it's not an easy problem to define because it's got a tradeoff. Getting the content quickly (easy: poll in a loop) vs not using server resources (easy: never poll). We have to define "better" before we can decide which solution is better.
- account42 2y agoIf-Modified-Since and ETag are nice and everyone should implement them but IME the implementation status is much better on the reader side than on the feed side. Trim your (main) feed to only recent posts and use Atom's paginatio to link to the rest for new subscribers and the difference in data transferred becomes much smaller. > Besides that, a well-behaved feed will have the same content as what you will get on the actual web site. The HTML might be slightly different to account for any number of failings in stupid feed readers in order to save the people using those programs from themselves, but the actual content should be the same. Given that, there's an important thing to take away from this: there is no reason to request every single $(&^$(&^@#* post that's mentioned in the feed. > If you pull the feed, don't pull the posts. If you pull the posts, don't pull the feed. If you pull both, you're missing the whole point of having an aggregated feed! Unfortunately there are too many feeds that don't include the full content for this to work. And a reader won't know if the feed has the full content before fetching the HTML page. This can also change from post to post so it can't just determine this when subscribing. > Then there are the user-agents who lie about who they are or where they are coming from because they think it's going to get them special treatment somehow. These exist because of misbehaved web servers that block based on user agen't or send different content. And since you are complaining about faked user agents that probably includes you. > Sending referrers which make no sense is just bad manners. HTTP Referer should not exist. And has been abused by spammers for ages.
- spiderfarmer 2y ago> These exist because of misbehaved web servers that block based on user agen't or send different content. And since you are complaining aber faked user agents that probably includes you. That's a niche. It's about 1 million percent more likely a fake request is coming from an overzealous AI scraper nowadays. I have blocked hundreds of them and I'm on the verge of giving up and handing over money to Cloudflare just for their AI scraping protection.
- dijit 2y agoUsing HTTP meta-headers is actually something we seem to have forgotten how to do. The one that annoys me most is the accept-language header which is almost entirely ignored in favour of GeoIP lookups to figure out regionality... which I find super odd; as if people are walking around using a browser in a language they don't speak. (or, an operating system configured for a language they don't speak). ETAG's though, are a bit fraught- if you're a company, a security scan will fire if an etag is detected because you might be able to figure out the inode on the filesystem based on it... which, idk why that's a security problem eitherway[0], but it's common for there to be false-positives[1]... which makes people not respect the header. Last-Modified should work though, I love the idea of checking headers and not content. I think people don't care to imagine the computer doing as little as possible to get the job done, and instead use the near unlimited computing power to just avoid thinking about consequences. [0]: https://www.pentestpartners.com/security-blog/vulnerabilities-that-arent-etag-headers/ https://www.pentestpartners.com/security-blog/vulnerabilitie... [1]: https://github.com/sullo/nikto/issues/469 https://github.com/sullo/nikto/issues/469
- balamatom 2y ago>as if people are walking around using a browser in a language they don't speak. (or, an operating system configured for a language they don't speak Well, yes, they are! Computers translated in my native language sound dumb. That's how a whole generation of my world learned better English than native speakers, ffs! Half of the time it's just translated wrong. You think anyone has any incentive to translate any technology to a language with a couple million speakers, all of whom are obligate pirates? And it seems like you might be surprised to hear that people speak more than one language. Then where's my global setting to tell the browser what languages I speak, so it'd know what header to send? Same place that lets me configure what ads I'm actually interested in. Nowhere. >I think people don't care to imagine the computer doing as little as possible to get the job done, and instead use the near unlimited computing power to just avoid thinking about consequences. This, friend, is what computers are for in the XXI century. "Bicycle for the mind", ha...
- dijit 2y ago
- theandrewbailey 2y agoRSS feeds have a TTL inside the feed. Do feed readers respect it? https://www.rssboard.org/rss-draft-1#element-channel-ttl https://www.rssboard.org/rss-draft-1#element-channel-ttl
- ozarker 2y agoI was just thinking a header with a suggested poll rate might be nice
- moebrowne 2y agoThe GitHub Event API uses a `X-Poll-Interval` header for this purpose. There is also Retry-After but that seems more targeted towards error states. EDIT: There is the `Cache-Control` header, it seems ideal for this use-case - https://docs.github.com/en/enterprise-cloud@latest/rest/activity/events?apiVersion=2022-11-28 https://docs.github.com/en/enterprise-cloud@latest/rest/acti... - https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Retry-After https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Re... - https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Cache-Control https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Ca...
- hahn-kev 2y agoI would ask "Has a feed ever tried to abuse that to get readers to go away" the answer is yes, at which point the reader would ignore it.
- balamatom 2y agoSo nice to see RSS making a comeback!
- someothherguyy 2y agoIMO, it is also unreasonable to have ultra-restrictive rate limits, like blocking a client after one request. https://rachelbythebay.com/w/atom.xml https://rachelbythebay.com/w/atom.xml
- ing33k 2y agohaha, true. Just got blocked because I opened the link once and clicked on refresh.
- rmholt 2y agoMaybe a warning after 3 requests and a ban on 4 per 24hr, but I understand the sentiment
- horsawlarway 2y agoI'm with you. Especially as the cost to serve this content approaches zero. I find the take in the blog to be relatively hostile. It's a "technically correct" rant. Not wrong, but mostly missing the point, and being a bit of a dick in the process. Sure - block the readers that make a request every 10 seconds. It's perfectly reasonable to block clients if they hit a limit like 20 to 50 requests in a day. It's damn hostile to block for 24 hours after a single request. If the 10MB of traffic for 20 requests is going to break the bank... maybe don't host an atom or RSS feed at all? --- That said - weirdos can weird on their own sites as they like. It's not a public service. But I bucket this into the same category of weird as posting a whole bunch of threatening "no trespassing", "beware of dog", "homeowner is armed", "Solicitors not welcome", etc style signs all over their property. Like - point out on the doll where the rss client hurt you. Because something's up.
- jstanley 2y ago> If you pull the feed, don't pull the posts. If you pull the posts, don't pull the feed. If you pull both, you're missing the whole point of having an aggregated feed! People probably do this because some sites only give you a preview in the feed, to force you to go to the site and view the ads. So if you want the full post in the feed reader, you need to pull the post as well.
- hylaride 2y agoThis. My feed reader pulls a "reader" view so I don't have to leave the app. I normally wouldn't mind going to the website, except that to do so would mean waiting for it to fully load, dealing with javascript popups, and often bad scrolljacking. This person isn't thinking as a user.
- Aeolun 2y ago160 gigabytes of feed over the course of a month (when polling a 640kb feed every 10 seconds), in case anyone else was wondering.
- benwerd 2y agoWebSub is your friend here: https://www.w3.org/TR/websub/ https://www.w3.org/TR/websub/ This adds a nice publish-subscribe model to RSS. Ping the WebSub server when there are changes; subscribing services are easily notified; nobody has to worry about excessive polling. Hooray.
- internetter 2y ago> If you pull the feed, don't pull the posts. If you pull the posts, don't pull the feed. If you pull both, you're missing the whole point of having an aggregated feed! In some cases the reader should fetch both the feed and the pages. Unfortunately, none do https://github.com/miniflux/v2/issues/3084 https://github.com/miniflux/v2/issues/3084
- gildas 2y agoBonus point for clients that don't support the HTTP “Accept-Encoding” header [1] and consume all your bandwidth. [1] https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Accept-Encoding https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Ac...
- fc417fc802 2y agoSeems like a reasonable case for disregarding the client preference. If you're able to speak TLS then you're able to load up a public domain (de)compression library.
- shadowgovt 2y agoRachel makes an excellent point here about feed change frequency. Seems like it'd be straightforward to implement a backoff strategy based on how frequently the feed content changed into most readers. For a regular, periodic fetch, if the content has proven it doesn't update frequently, just back off the period for that endpoint.
- ffjffsfr 2y agoI’m 100% sure there are many badly written inefficient crawlers that are wasting server resources and resources where they run but I use feed readers a lot and it is very hard to find well maintained feeds. Many servers also use cache related headers incorrectly or don’t use them at all.