14 ms·
Show HN: Python package to collect normalized news from almost any website
- spiritplumber 6y agoThis looks amazing.
- artembugara 6y agoThanks!
- anonytrary 6y agoThe demo gif is really long (a few minutes), but it's well worth watching and summarizes the capabilities of the library well. It'd be cool if you could register additional sources via a registration API or some passed in configuration.
- pydry 6y agoThis would be very cool - a bit like how youtube-dl handles non YouTube websites.
- carbolymer 6y agoSo, the centralized, paid version of RSS?
- artembugara 6y agoIt's a Python package to collect news data. Nothing paid
- oceanbreeze83 6y agoFor any news junkies heres, I've built https://maagnit.com https://maagnit.com which gathers both Left and Right leaning sources for any story and displays them altogether on one page. My approach has been, if we can't get neutral/objective coverage, getting comprehensive, 360 desgree coverage is a good alternative. These days, news bias is not only in the way a story is covered, but also which stories are covered. So, maagnit automatically collects stories from the left and right.
- jsilence 6y agoThat is a brilliant idea!
- heisenzombie 6y agoI’m sure this is not an original comment, but it’s interesting to see what you’ve classified as left/right. It must be difficult given that there is not really one axis of left/right and that “the centre” is highly relative. To me, seeing the BBC and Euronews in the “left” section is pretty funny, but I guess it’s true relative to US politics. Is there anything you’ve learned about “the left media” and “the right media” from doing this? Do you think your sources are equidistant from “the centre”?
- scrollaway 6y agoThe fact US politics call them "left" and "right" is meaningless even, given how right-shifted US politics are (By EU standards, only US extreme-progressives are actually in the european "Left"). Not to mention americans have demonized "the center" as some "if you're somehow trying to consider all the facts you're a coward who can't decide" type of thing. The two sides being "at war" drives TV/website engagement and that's all that matters to the people writing the headlines. It just so happens that currently, one of the two "sides" relies heavily on disinformation; so anything that tries to fight disinformation (including remaining impartial) is that side's enemy. So those things end up being considered "left-wing". Truly, the united states has four political parties: The Media Left, the Media Right, the Political Left, and the Political Right. Nearly every american you know is part of the first two; the last two don't make for good TV.
- cbnotfromthere 6y ago"how right-shifted US politics are As an EU citizen myself, I actually believe it is the Western EU which is left-shifted. "By EU standards" Perhaps you mean Western EU standards; I'm sure the majority of Polish / Hungarians / Romanians / Baltics will be shocked of your ideas. "only US extreme-progressives are actually in the european "Left"" The US extreme-progressives would be spat on as disgusting Commies in Eastern EU, where the left-wing Democrats would be the equivalent of Social-Dmeocrats.
- permanent 6y agoHi looks interesting and useful! What's the differences between newscatcher (python) vs. newscatcherapi.com? Is there any limit on using newscatcher (python) as shown in the pricing pages? Also, I was looking at https://newsapi.org https://newsapi.org. How does your python API compare? I see that in newsapi, it also get old articles; do you implement similar features?
- kizy25 6y agoHey. Newscatcher (python) is an open-sourced package that we developed for users' side projects. There are no limits. You can even modify it for your own needs. Newscatcher API is a product that allows you not only get the latest articles but also search by keywords, topic, country etc. Basically, the main feature is that you search for articles that contain a specific word or phrase. This can have an added value for professionals and companies. There will be a free plan for developers with limits and chargeable for more heavy usage. Compare to newsapi, we are less expensive for the content. We will not be able to get old articles from today. We began to stock data couple at the release. Hope I answered all of your questions.
- darwinwhy 6y agoDo you guys have any plans to create webhooks that let you know when a feed/search is updated with a new article in real time?
- artembugara 6y agoHey. Not yet. Let us think about that.
- darwinwhy 6y agoIt's probably not worth it. Anyone needing near-real-time feed webhook updates will probably build their own scraper.
- artembugara 6y ago
- marban 6y agoSelf-Plug: https://www.hvper.com https://www.hvper.com (Official Successor of popurls which more or less started the single page aggregator craze.)
- rawoke083600 6y agoLooks good ! How do I get the BIG pink bar away ? I would rather see ads (to support the free version) than that BIG PINK BAR ?
- marban 6y agoYou can either cover my monthly 5 digit server bill or upgrade to a paid version. Sorry to be the bearer of bad news ;)
- zwaps 6y agoThat pink bar takes up almost half of my screen on my iphone 11 and makes the website defacto unusable. Not to be unkind, but you should fix that. Instead of being incentivized to pay you, my only instinct was to close that page asap.
- zo1 6y agoHow about you switch to desktop mode on your browser instead of forcing sites to be mobile friendly so you can consume it on a tiny screen. Half the web is broken because of "responsiveness".
- napolux 6y agoHey, looks cool. Which news sources are available for the italian market/language?
- artembugara 6y agohey, you can get the news sources by `it` language and check it
- napolux 6y agothx!
- rootuid123 6y agoThis article is just a promotion for a commercial service. Shame on you.
- artembugara 6y agoAlright
- nmstoker 6y agoHow complete are the articles that are returned? Last time I looked into RSS news feeds a lot of sources would just put abbreviated / teaser content in and then try to get you to click through to the story on their site. That obviously didn't make for the best experience with an RSS downloader though.
- k1m 6y agoI work on Full-Text RSS which can help convert abbreviated feeds into full-text versions. The idea is you'll get a new feed URL from Full-Text RSS to use instead of the original partial feed in your news reader or application. Free to try here: http://ftr.fivefilters.org/ http://ftr.fivefilters.org/ and code for a slightly older version available here: https://bitbucket.org/fivefilters/full-text-rss/src/master/ https://bitbucket.org/fivefilters/full-text-rss/src/master/
- raziel2p 6y agoThis doesn't "collect" anything, it's just a python package wrapped on top of a sqlite database. If you choose an arbitrary website that doesn't have information in said database you just get "website not supported". There's not even any logic to guess a website's RSS feed URL... And in the end it's just an RSS feed reader?
- artembugara 6y agoIt's just a python package wrapped on top of a sqlite database. Yes, it is! Though thousands of ppl found it useful
- justaguyhere 6y agoWhat would the logic to guess the RSS feed URL look like? I suppose it is easy for wordpress sites, might not be so simple for others?
- slightwinder 6y agoCommon way is to set a link-tag in html with type rss, so others can discover the feed to the active url. If this is not set, chances are slow that there is a proper feed available. Not that you can't still try guessing and googling for it...
- pmyteh 6y agoDepending on what you're trying to do, there's also newspaper3k: https://github.com/codelucas/newspaper https://github.com/codelucas/newspaper It's quite easy to get "good" extraction for large numbers of outlets/articles without a massive amount of special-casing, as news articles are nearly universally marked up with RDF metadata (partly for Google News's benefit). Article discovery, and perfect parsing, is quite a bit harder. I ended up rolling a new Scrapy project with site-specific parsing code for an academic project as I had quite specific requirements.
- jerzyt 6y agoIs your code on github? I'm actually working on something very similar, and would love to get some ideas from what you've done.
- pmyteh 6y agoYes! https://github.com/pmyteh/RISJbot https://github.com/pmyteh/RISJbot
- totemandtoken 6y agoHey, I used this for my newsbetting site - https://www.rashomonnews.com/ https://www.rashomonnews.com/ I haven't pulled any articles in a while, so it's a little outdated but I love newspaper3k.
- forgingahead 6y agoNeat project! Definitely useful to have this, folks can build "headline-edit trackers" or services that collect news using other filters more easily using this.
- seemslegit 6y agoI'm not sure how to parse the following sentence: By newscatcherapi.com (this package is fully self-sufficient, you can just use it. No dependency on external services/API)
- stevenjohns 6y agoI interpret this as Created by <api provider> but you do not need anything from us or from anyone else to get the software going, it just works out of the box.
- artembugara 6y agoI am going to put it on a README. Thanks a lot!
- seemslegit 6y agoWell the 'anyone else' part is wrong, someone has to provide the news.
- artembugara 6y agoWe tried to make sure that everyone understands that it does not depend on newscatcherapi.com products/services)
- kuu 6y agoThis seems really cool! How does it work internally? Is it downloading the news from a RSS or is it crawling the content of the website? Or is the content coming from an external service? How are the feeds selected? Can we add more? who is maintaining them (in case the data is crawled)? Thanks!
- kuu 6y agoI saw this at the bottom of the README: The package itself is nothing more than a SQLite database with RSS feed endpoints for each website and some basic wrapper of feedparser.
- artembugara 6y agohey, there is a sqlite DB that stores the RSS endpoints. Then we use feedparser python package to parse it.
- slightwinder 6y agoThen why should I use this instead of a real Feedreader? What advantage has this?
- artembugara 6y agoI think you mean feedparser. It knows how to parse the feed. So you have to give it the feed’s url
- slightwinder 6y agoNo, I mean feedreader. A Feedreader is a service which collects and parses rss-files, then makes them accessable in an interface for the user. Similar to a mailclient, but for rss. This package seems to do the first part, collecting and parsing, but lacks the interface. So what is the point of it? And if this is maintaining it's own config for sources, can I even add my own sources? Or is this just an elaborated OPML-file with attatched business-logic?
- nreece 6y agoThis looks great! I read your blog post about how open-sourcing helped you find testers, which is a great step for such tools. * Shameless plug *: Our web service, Feedity - https://feedity.com https://feedity.com, helps create a custom feed for any news webpage, via a point-and-select feed builder and REST API.
- nickthemagicman 6y agoHow do subscriptions work with this? Or is it treated like anonymous browsing?
- slightwinder 6y agoI would be more interessted in a ready-to-use server which I can selfhost and to which I can throw any script for collecting data, filter and process them, which also handles errors and storage. At the moment the only real solution for this seems to be using your own RSS-Reader (Tiny Tiny RSS for me at the moment), which has the disadvantage of being limited to RSS-sources, as also not allowing much filtering and processing. But with more and more sources moving avway from RSS, I want something which can be fetch alternative sources and integrate them into a unified interface. In best case it would be even work with any language, be it a shellscript or python, ruby or even java.
- artembugara 6y agoWell, we hat is quite similar to what we are doing at newscatcherapi.com. We collect tons of news data and let you query it with an API
- slightwinder 6y agoNot selfhosted, isn't it? By which I mean it costs money.
- artembugara 6y agoYeah. Though. It will have a free plan. And our hosting cost quite a lot so (imho) it’s almost always must cheaper to buy a ready solution.
- kingofpandora 6y ago> Programmatically collect normalized news from (almost) any website. By this you mean ... Fetch RSS feeds from the websites that we've hardcoded into a Python library? It seems like the work here was in collecting and cataloging a lot of RSS feeds. You don't seem to be able to arbitrarily "collect news from any website".
- gitgud 6y ago> "Almost any website" Is there a list of supported websites?
- 0x006A 6y agoWhy did you choose to ship the list of RSS feeds as an SQL database? This makes it hard to keep it up to date and submit pull requests with additional sources. Would it not be better to keep that info in a json file or a dict / list in a python module?
- artembugara 6y agoHey. Yes. You are right. We will change that!
- remram 6y agoWill you? https://github.com/kotartemiy/newscatcher/issues/3 https://github.com/kotartemiy/newscatcher/issues/3
- pictuga 6y agoyeah, git isn't really meant to handle binary files
- 0x006A 6y agoYour git repo contains the dist folder even though its in .gitignore, might be good to remove that. No need to checking generated artifacts.
- artembugara 6y agoThx. Will check.
- tyingq 6y agoThe metadata of language, topic, rss url, etc, is nice work. For use outside of python, here's a gist with a sorted list of sites and a sqlite dump of the site data: https://gist.github.com/tyingq/8e921eed10bf2ecf9c40ebdd70ff1871 https://gist.github.com/tyingq/8e921eed10bf2ecf9c40ebdd70ff1...
- jerzyt 6y agoFirst, kudos to the OP. If you don't mind, would you compare and contrast it to newspaper3k?
- uptown 6y agoNeat! Pair this with some sketchy ad tech, and you can start raking in the bucks. https://www.cnbc.com/2020/05/17/broken-internet-ad-system-makes-it-easy-to-earn-money-with-plagiarism.html https://www.cnbc.com/2020/05/17/broken-internet-ad-system-ma...
- A4ET8a8uTh0 6y agoMy initial reaction is very positive. I believe there is a market for a tool like this. Lets see how it handles marketwatch.
- nojito 6y agoThis is literally just a sqlite db with rss feeds and a python script with a very misleading description. The real purpose of this post is to get traffic for their paid news api. Also this closed issue is hilarious! https://github.com/kotartemiy/newscatcher/issues/3 https://github.com/kotartemiy/newscatcher/issues/3
- fuball63 6y agoThis looks like a cool library, but I have a question about the newscatcher API. How does the licensing work for the content? Seems odd (but great) that I can just read the news in my terminal from NYT but not pay a subscription or see ads. I read in some of the comments it's an RSS feed, is that freely available all the time? Surely even the RSS feed is protected with copyright and has restrictions on republishing? If that's not the case, the larger implication here that the news is free if it's in a format that is not as widely used (RSS) compared to what the mass populous uses (mobile browsers/app). Cool library, thanks for sharing!
- dunefox 6y agoWhy is this being downvoted? Seems like a viable question.
- artembugara 6y agoHi. Co-founder is here. Short answer. We do not know if it is legal!
- hundchenkatze 6y agoIt seems a little risky to build a paid service that you're not sure is legal. Also, the Terms of Service, Privacy Policy, and GDPR Policy links in the footer of your site don't work. They all have empty hrefs.
- amelius 6y agoPerhaps nice to show news headlines as an alternative to /etc/motd
- danmorpius 6y agoprevious discussion on the same repo https://news.ycombinator.com/item?id=22407835 https://news.ycombinator.com/item?id=22407835
- pheme1 6y agoWhat does the normalized news means? Able to query from a predefined RSS feed? Because I was expecting news coverage from different partisan source pointing to the same news event.
- polymorph1sm 6y agoSeems like a repost of previous Show HN [1]. [1] https://news.ycombinator.com/item?id=22407835 https://news.ycombinator.com/item?id=22407835
- unixhero 6y agoWill this work for fetching financial data from the public bloomberg.com website?
- unixhero 6y agoI see your downvotes. It means I am on to something. I will try.
- hitpointdrew 6y agoWhat is "normalized news"? What does that even mean?
- zzuko 6y agoThere is a similar library for this called newspaper which I had used in my undergrad thesis. Not dismissing this work, but I am curious about what it is offering on top of it? It doesn't offer a comparison to the newspaper, at least not in the Github page.
- ryanalam 6y agoI've been experimenting with a side project: www.glancereport.com, which pulls the top headline from a variety of news sources. You can also sort these headlines by neutral, progressive, or conservative sources. I'm still learning how to play with Python and Redis, and it's not perfect, but would love feedback.
- modlinska 6y agoI see that the Neutral/Progressive/Conservative filter is most applicable to Politics news, but I'd be interested in seeing if you can filter by topic like Business, Sports, Lifestyle etc.
- charlesdaniels 6y agoAs others have noted, this doesn't seem to collect the full article text, just stuff that you would get from an RSS fee. From the title, I expected something more like newspaper3k[0]. I used that for an NLP class during undergrad to collect full-text news articles, in conjunction with Selenium (many mainstream sites don't work with just plain wget or requests). Lately I've starting using EpubPress[1] to grab full-text articles and generate an ePub, which happens every night via cron. Then I can get a full digest on my iPad over sftp at my leisure. Sadly EpubPress is not very sophisticated, sites like Bloomberg or ArsTechnica return "are you a robot" challenges which it can't bypass. I wish there was some kind of community driven library for retrieving full-text articles from common sites. In my vision of how that would work, users would contribute hand-crafted Selenium scripts to download and extract the article text, bypassing the bot-detection for each site. Then something like EpubPress would work a lot better. The "modern web" just has too much junk to be interesting any more. Sometimes news sites publish articles I would like to read, but I'm not interested in dealing with 1000 different implementations of crappy mobile UIs, advertisements, animations, etc. I know reader view exists, but you still have to wait for the page to load, and it doesn't work very well for some sites. For the sites it doesn't break with, the experience with EpubPress is much better. 0 - https://github.com/codelucas/newspaper https://github.com/codelucas/newspaper 1 - https://epub.press/ https://epub.press/