5 ms·
>This dreary state of affairs is due in part to the fact that "checking" an RSS just means downloading the entire file containing all posts the author would lik
by bArray 8y ago
>This dreary state of affairs is due in part to the fact that "checking" an RSS just means downloading the entire file containing all posts the author would like to make public, regardless of how many are new to the reader. This insanely unnecessary bandwidth usage penalizes sites that have long (>10) feeds with complete posts. (The single RSS file on my blog takes up the majority of my bandwidth costs.)
Still, exceptionally low cost compared to running pretty much any website's massive CSS and JS files. A single image in most cases will take more than the entire RSS feed before compression.
That said, I would like to see a new standard (a new one would be needed [1]) that only gets the difference from what you last read - I think that would really take RSS feeds to new places of usefulness. There's no reason why you couldn't send the server an ID (not timestamp to avoid issues with timezones, clock stretching, forward/backward time setting, etc) of the last request and have it send back everything since (within reason).
[1] https://www.w3schools.com/XML/xml_rss.asp https://www.w3schools.com/XML/xml_rss.asp
- abecedarius 8y agoThat’d impose a cost on the server, remembering an ID map. Instead make the client send back a value given by the server last time.
- Izkata 8y agoI think that's what they meant by "an ID [...] of the last request"
- jerf 8y agoThat's something both RSS consumers and produces should already support, but don't always, called an E-Tag, a standard part of the HTTP standard. However, the E-Tag is all-or-nothing; either it matches, and the entire request is essentially aborted, or it doesn't match, and the entire file is served up. Passing "the ID of the last piece of content I saw" would allow the server to return just the updated stuff, or abort early like an E-Tag. However, counterintuitively, as is often the way in computer science, I'm not sure it would be that big a win to be able to return partial content. The vast bulk of the win on most blogs will just be the ability to abort at all, provided just fine by E-Tags. I would say that if your site is getting hammered by HTTP requests for your RSS, do double-check that you've got E-Tags set up and working correctly. It is in the best interests of the big scrapers to support that properly, as they are paying for that bandwidth too. RSS aggregators don't have to get too large before this becomes a top-priority feature request. Unless the feed is literally changing on roughly the same frequency as it is scanned, it shouldn't be the dominant factor in your bandwidth bill.
- jessriedel 8y ago> Still, exceptionally low cost compared to running pretty much any website's massive CSS and JS files. My blog has both RSS and browser readers, but the bandwidth is dominated by the RSS feed. Since the RSS file has no images (just HTML pointing to the images), I think it's more likely that for my Wordpress blog the CSS/JS overhead is just not that much (as opposed to the alternative hypothesis that I have many many times more RSS readers who never end up downloading the images).
- prepend 8y agoOr you have way more users via rss than through their browser. This could be a good thing as it results in less bandwidth than if they all hit you directly through browser.
- jessriedel 8y agoMy parenthetical comment raised this hypothesis and rejected it with a reason. So you should address that reason if you want to disagree.
- bArray 8y ago>the bandwidth is dominated by the RSS feed Exceptionally low cost per hit, as opposed to overall bandwidth. Overall bandwidth will probably fair off worse as you say due to the polling nature of RSS. I think in general RSS readers could do a better job of fetching heads and checking whether or not there is a change worth fetching, that would save a tonne of bandwidth. Also in general, I would be tempted to make an RSS feed more minimalist in terms of content and markup. It should just be a short `<description>` and a link to the main article (which would still allow you to potentially monetize your content or gauge interest more accurately). >I think it's more likely that for my Wordpress blog the CSS/JS overhead is just not that much Also, bandwidth is just one resource - potentially each call to a page is a database read, whereas your RSS feed should be static (not sure about the WordPress implementation, but I would hope for static caching with something that doesn't change for long periods of time). I've seen WordPress database lockups with modest amounts of traffic (again, most of the time this could have been easily statically cached - but doesn't appear to be by default).
- elFarto 8y agoYou should already be able to implement this using the If-Modified-Since HTTP header. The server only needs to send you back articles in the RSS feed that have been added after that date. The If-Modified-Since value is meant to come from the previous responses Last-Modified (not a timestamp you make up), so that side-steps timezone issues.
- bArray 8y agoI think that still requires a change to both the server and the reader - although at least not a new standard.