3 ms·
I don't understand the following: "Yelp passed 100 million reviews in March 2016. Imagine asking two questions. First, “Can I pull the review information from
by blergh999 10y ago
I don't understand the following:
"Yelp passed 100 million reviews in March 2016. Imagine asking two questions. First, “Can I pull the review information from your service every day?” Now rephrase it, “I want to make more than 1,000 requests per second to your service, every second, forever. Can I do that?” At scale, with more than 86 million objects, these are the same thing."
Who is making 1000 requests per second to retrieve all of the 100 million reviews? Other services within Yelp? Why would they pull all reviews every time, and not just the reviews they haven't already processed?
How is 1000 requests per second the same thing as pulling all review information every day?
This section is really confusing and needs some clearer explanation.
- j-m-o 10y agoI think the unstated assumption is that there's some sort of processing that occurs on each review every day. So with 100 million reviews every 24 hours, that's just over 1157 requests per second.
- artpepper 10y agoI think it's saying that a single iteration over the entire set, would translate to 1000 requests per second for a day (if done naively as one request per object). It's really talking about the N+1 problem.
- vhost- 10y agoThis is how I read it too. I think op was taking that too literally when they were just stating that the two are essentially equivalent at that scale.
- AznHisoka 10y agoIt has to be internal because they are incredibly protective of their API (hell they sold their firehouse to just 1 darn company). My guess it's NLP type processing. Things like review highlights, recommendations, etc. that's probably 99% of their load. The user-centric stuff like submitting reviews, comments is trivial - as another user said , a large single instance is enough.