5 ms·
OP again, we're way overloaded. We got 200 profile requests already in an hour and we didn't see it coming. We're extremely thankful and sorry to disappoint you
by lourot 8y ago
OP again, we're way overloaded. We got 200 profile requests already in an hour and we didn't see it coming. We're extremely thankful and sorry to disappoint you at the same time. We're closing the "Get your profile" feature for now. Spinning up many EC2 instances to process the 200 profile requests faster in parallel as we speak. And we'll come back soon with a system that can handle this load.
Many many thanks!
- KenanSulayman 8y agoIsn't 200 requests in an hour ... just three requests per minute? What is this service doing?
- iokanuon 8y agoMaybe it's cloning all GitHub repos of a user to compute activity?
- brillout 8y agoWe don't clone repos. Instead we use GitHub's API to get activity info. E.g. https://api.github.com/repos/aurelienlourot/ghuser.io/contributors https://api.github.com/repos/aurelienlourot/ghuser.io/contri...
- brillout 8y agoGitHub's API does't provide the exhaustive list of all your contributions. Instead we crawl GitHub's website which is slow. Details at https://github.com/AurelienLourot/github-contribs#how-does-it-work https://github.com/AurelienLourot/github-contribs#how-does-i... We are in talks with GitHub and they know that we are crawling GitHub.
- ChristianBundy 8y ago> Instead we crawl GitHub's website which is computationally expensive I'm surprised to hear that you're bottlenecking on CPU time. Could you verify that my understanding is correct? I would've thought your bottleneck would be networking and connectivity as you have to wait for GitHub to process all of the requests.
- rjacksonm1 8y agoLooks like they're hitting a GitHub rate limit, rather than bottlenecking on CPU.
- brillout 8y agoYou're right. I edited my answer. The GitHub's unofficial API we are using is slow and per IP rate limited. We spin up several servers to have several IPs to circumvent the rate limit. (GitHub knows that we do that and we are in contact with them.)
- exikyut 8y agoIf GitHub are okay with you using multiple IPs to get that data then it's not inherently expensive on their side for you to be using this. Surely a rate-limit exception could be in order, then? And perhaps you could help them alpha-test a new API endpoint that just so happens to include all the info rolled up as one URL :D (Hmmmmm.... GraphQL....)
- lourot 8y agothis is the bottleneck: https://github.com/AurelienLourot/github-contribs https://github.com/AurelienLourot/github-contribs It uses a very slow non-official GitHub API and so it takes several hours to do the initial crawling of one single profile, and is limited by IP so you need several machines/instances in order to parallelize. We plan to use AWS Fargate for the future. (We thought this future would be much farther away)
- sethish 8y agoYou may want to delete my profile request, I have somewhere between 50k and 60k repositories and it tends to crash services. (user: sethwoodworth)
- lytedev 8y agoHaha! That is amazing! How, exactly, do you have so many?!
- ishanjain28 8y agoCan we clone the repo and run it locally on my own profile? I arrived late..
- brillout 8y agoNot currently but we are thinking about it, see https://github.com/AurelienLourot/ghuser.io/issues/103 https://github.com/AurelienLourot/ghuser.io/issues/103
- sam0x17 8y agoWould it be practical to use puppeteer / headless chrome from within a Google Cloud Function to do this? Then you could do literally millions of profiles per hour. The only requirement is that it never takes more than 540 seconds for any one profile, and that you could work around if you write in the ability to export the execution state and send it off to a new function invocation if you are close to the deadline. This seems like a much easier problem if you can get out of VPS land and into lambda/cloud functions since its something like $0.00001 per invocation. https://cloud.google.com/blog/products/gcp/introducing-headless-chrome-support-in-cloud-functions-and-app-engine https://cloud.google.com/blog/products/gcp/introducing-headl...