3 ms·
How is it that you are able to scrape all of the content of Reddit, while 3rd party apps have almost stopped working.
by roshin 3y ago
How is it that you are able to scrape all of the content of Reddit, while 3rd party apps have almost stopped working.
- rglullis 3y ago- reddit still allows 600 requests per 10 minutes free of charge. - third-party apps need to make requests on behalf of all their users in real-time, this is making requests for a single user (the backend) and can be batched. Right now it it fetches data from 10 subreddits per request and it can be made to get even more at once. - I am not scraping all of reddit, only about 80 for now. It can be done every 3-4 minutes well within the rate limits. - if necessary, I could take the same approach used by the lemmit.online dev, and just request the JSON feeds (which do not require an API key) and parse the results on my own. - If more people run this service, we can make them cooperate so that we do a "MapReduce" on the scraping: the requests to different subreddits are spread between the fediverser instances, and they can gossip back the results among themselves.