5 ms·
This is something I'll be working on soon enough. There are many, many dependencies (not only for the web app, but there are three DBs that you have to have ins
by conesus 14y ago
This is something I'll be working on soon enough. There are many, many dependencies (not only for the web app, but there are three DBs that you have to have installed as prerequisites - mongo, postgres/mysql, and redis).
The real problem with setting up your own instance of NewsBlur is that you'll have to do your own feed fetching. This is effectively what you're paying for when you pay for NewsBlur. Let me break it down.
500 feeds updated once every five minutes is 144,000 feed fetches a day. Couple that with the original page and icon fetches, you're looking at almost two feeds every second just to keep your feeds up to date. And the average feed takes 5 seconds to completely update (feed, page, differences in stories, updating unread counts). So you have to run 10 processes in parallel, and that's already beyond the capabilities of most single machines. Then good luck keeping your DBs running clean and backed up.
Or you could pay $2 / month and never have to worry about setting up a very big stack. And you get the social community of shared stories on newsblur.com.
- georgeoliver 14y agoDoes NB fetch feeds that have no new articles?
- UberMouse 14y agoDoesn't it have to fetch feeds to determine there are no new articles?
- tcwc 14y agoFeed readers can send If-Modified-Since or If-None-Match as part of the request so the server only sends back the full feed if there's something new (Otherwise a 304 Not Modified)
- mikro2nd 14y agoAssuming that the feed implementer bothered... many news CMSs don't bother. The world of RSS is a horrorshow of standards noncompliance; Atom only somewhat better. (My experience is, admittedly,a couple of years out of date, but feed generation tends to be a minimum spend item for news organisations.)
- boyter 14y agoJust to add to this, the size of the data grows far more quickly then you would expect. I crawl 30 or so feeds for a personal project and database after 2 months of operation is over 200 meg in size. For the cost of hosting/development/work you are MUCH better paying for newsblur and forgoing 6 outside coffee's a year.
- duncan_bayne 14y agoThree DBs? I don't suppose you have a blog post or something that explains your architecture decisions? Not criticising - I'm sure you had good reasons - but I'm rather curious as to what they were :)
- bsg75 14y agoI too am interested, esp what goes into Mongo and what into Postgres (assuming Redis is a form of cache?) I suppose I could also look at the source ;)
- BoyWizard 14y agoBased on the Github page, I would imagine Redis as a cache, MongoDB to store the 'non-relational data' (quoted from the GH page - I'm guessing the actual articles and such) and Postgres for the other application stuff - accounts, preferences, subs, etc.
- ceol 14y agohttps://github.com/samuelclay/NewsBlur/blob/master/apps/rss_feeds/models.py https://github.com/samuelclay/NewsBlur/blob/master/apps/rss_... It looks like MongoDB is being used for fetch/push history, feed icon data, and article data. It's also being used for some other stuff: https://github.com/search?l=Python&q=%22mongo.Document%22+repo%3Asamuelclay%2FNewsBlur&ref=advsearch&type=Code https://github.com/search?l=Python&q=%22mongo.Document%2...
- mnutt 14y agoHave you considered releasing VM images? I've thought about it for open source projects that have a bunch of dependencies, but haven't actually tried it. On the other hand, you may end up with people running the app that have no way to maintain/support it.
- rssident 14y agoYou should make an intelligent feed updating algorithm. If you have at least fifty previous entries from a feed it's easy to predict fairly accurately when that feed will need to be updated again. That's how the indexer for http://rssident.com http://rssident.com works. Saves you tons of cpu cycles.
- sergiosgc 14y agoMost blog providers support pubsubhubbub(the worst protocol name ever). It allows you to avoid polling, by having the content producer notify you of feed updates.
- dpcx 14y agoExcept that, even with PSHB, you get polling. I wrote about it here http://www.dp.cx/blog/pubsubhubbub-and-polling.html#.UUcPWLp4Yag http://www.dp.cx/blog/pubsubhubbub-and-polling.html#.UUcPWLp...
- sergiosgc 14y agoThe fact that two feed clients, Superfeedr and Guzzle, implement PuSH wrong in no way leads to the conclusion that, when creating a feed reader, you can avoid polling for every content source that does support PuSH.
- laureny 14y ago> So you have to run 10 processes in parallel, and that's already beyond the capabilities of most single machines. Er... really? My laptop is three years old and I run 150 threads on it every time I launch the server I'm working on (dozens of times a day). No problems at all. It's in Java, in case it matters.
- MikeAmelung 14y agoI believe he means that each process requires ~5 seconds of CPU/system time per update, meaning that even a single 8-core machine wouldn't keep up. It's not really related to how many threads or processes can be created, rather how many can actually be doing work in parallel.
- VLM 14y agoThere's an old saying that the true definition of supercomputing merely means moving the bottleneck from the CPU to the IO subsystem. The modern "big data" world seems about the same way. On a regular basis I find strange new ways at work to saturate any I/O system I get access to... Give me a faster I/O system, I'll find an exciting new way to bring it to its knees.
- lgp171188 14y agoI tried very hard to setup my own instance as well going through the fabfile.py to figure out things, but I hit a point where I simply couldn't go ahead on my Debian Squeeze installation. I'd love to have a very good documentation on the installation process. Also there seem to be a lot of stuff that need to be installed and configured that might not be necessary for single-user/<10 user instances. Yeah the 'fetching the feed part' still needs to be done, but I guess right now it is close to impossible for anyone apart from yourself to be able to understand how to do that. Due to the mass exodus from Google Reader, I understand that NewsBlur has faced challenges in scaling and hence is slow. But things can only get better from here. Since it in open source software, you should at least setup a getting started developer's guide so that you will get a lot more contributions from the community.
- Flenser 14y agoCould you federate some of the feed fetches to your users? Get the UI to check for updates to some subset of that user's feeds and then ping the newsblur server to tell it when there's been an update.