5 ms·
Unfortunately, the Facebook bridge is very broken and no longer maintained. From the public pages I've tried, it only intermittently works on at most half of th
by sitta 5y ago
Unfortunately, the Facebook bridge is very broken and no longer maintained. From the public pages I've tried, it only intermittently works on at most half of them.
I don't blame them for not maintaining it though. I grabbed the code to see if I could get it working and realized Facebook has made their newer pages fiendishly difficult to scrape. Assuming one does put in the effort to make it work, how long will it last till it's broken again?
I wish there was a viable alternative platform for organizations where "public" content is actually accessible publicly.
- trog 5y agoI have been wondering if, instead of bothering to try to scrape Facebook via DOM, it might be worth pursuing a basic image matching approach to divide a Facebook page up into regular "chunks" that are just saved as images. I'd personally be happy with an RSS feed of Facebook as just a series of images in a feed, but I guess once you've gone that far you could relatively simply run it through OCR to get most of the text.
- Karrot_Kream 5y agoI've toyed around with scraping Facebook through a headless chrome myself but haven't done it yet. Have you taken a look at https://mbasic.facebook.com https://mbasic.facebook.com ? It might be simpler to parse.
- tfsh 5y agoUnfortunately the RSS bridge doesn't work and also seems to be unmaintained, this is surprising given it seems to be set as one of the "core" bridges. Data archival, retention and accessibility is absolutely fundamental and it's unfortunate to know that so many companies are hell bent on stopping individuals from using open-means to access such data (although I am not surprised, and from a rational perspective I can understand their reasons for making it difficult).
- doom277 5y ago> Unfortunately, the Facebook bridge is very broken and no longer maintained. True. There is no maintainer for FacebookBridge. As for now, the only way to fix FacebookBridge that I see is: 1. using existing Facebook account 2. fetch pages using Selenium (simply speaking - web browser) This is mentioned in https://github.com/RSS-Bridge/rss-bridge/issues/2400 https://github.com/RSS-Bridge/rss-bridge/issues/2400
- deleted 5y ago[deleted]
- 1vuio0pswjnm7 5y ago"... Facebook has made their newer pages fiendishly difficult to scrape." This has not been my experience. It is still as easy as ever using the mobile sites. However I suspect one person's notion of "scraping" is not always the same as another's. I prefer to work on the command line, in textmode. Thus when I check Facebook I want text-only, no graphics. First I extract the all the story.php URLs and sort them by date. Then I retrieve the contents stored at these URLS and dump into a text-only format so I can read through comments. If there is something that looks interesting I can save the story URL and check out the photos later on a computer that has graphics layer loaded. TBH, with Facebook I mainly just check messages and notifications. Saving story URLs in chronological order is useful for me because it makes it easy to go back and find items from the past. But if I were really serious about monitoring a "feed" I can create one myself, better than Facebook's, by retrieving the profile page of each friend and extracting story URLs from their source instead of relying on Facebook's manipulative algorithms that deliberately hide stories and re-order what they do show, non-chronologically, in a way that suits Facebook's interests over the user's. I do not use any fancy software, just a local TLS proxy, netcat (or equivalent) and ubiquitous base userland UNIX text processing utilities. I use a text-only browser (not lynx) for reading HTML and a pager (less) for reading formatted text. I use the m.facebook.com or mbasic.facebook.com sites, not www.facebook.com because the site on the www subdomain uses GraphQL instead of hyperlinks. Not to mention the mbasic subdomain has no ads. The way websites use GraphQL instead of simple hyperlinks is unnecessarily complex, reminiscient of a Rube Goldberg cartoon.
- RobSm 5y agoWhat's the point of TLS proxy?
- 1vuio0pswjnm7 5y agoIn this instance it is to enable use of any TCP client, namely ones that do not support TLS, e.g., original netcat, tcpclient or an early version of a text-only browser. I use a variety of clients and with only one exception I only trust the proxy to make remote TLS connections. If one uses a client with acceptable TLS support, then there is no need for a TLS proxy.
- deleted 5y ago[deleted]
- anthropodie 5y agoI say instead of maintaining that bridge we ask whoever is still using Facebook to move somewhere else. That would be easier instead of maintaining that bridge because Facebook will actively try to patch any holes in their walled garden.