7 ms·
This is such a clever way of sampling, kudos to the authors. Back when I was at Pew we tried to map YouTube using random walks through the API's "related videos
by pvankessel 3y ago
This is such a clever way of sampling, kudos to the authors. Back when I was at Pew we tried to map YouTube using random walks through the API's "related videos" endpoint and it seemed like we hit a saturation point after a year, but the magnitude described here suggests there's a quite a long tail that flies under the radar. Google started locking down the API almost immediately after we published our study, I'm glad to see folks still pursuing research with good old-fashioned scraping. Our analysis was at the channel level and focused only on popular ones but it's interesting how some of the figures on TubeStats are pretty close to what we found (e.g. language distribution): https://www.pewresearch.org/internet/2019/07/25/a-week-in-the-life-of-popular-youtube-channels/ https://www.pewresearch.org/internet/2019/07/25/a-week-in-th...
- 0x1ceb00da 3y agoThis technique isn't new. Biologists use it to count the number of fish in a lake. (Catch 100 fish, tag them, wait a week, catch 100 fish again, count the number of tagged fishes in this batch)
- dclowd9901 3y agoI made the same connection but it’s still the first time I’ve seen it used for reverse looking up IDs.
- fergbrain 3y agoIsn’t this just a variation of the Monte Carlo method?
- deleted 3y ago[deleted]
- zellyn 3y agoDo you get the same 100 dumb fish?
- egeozcan 3y agoCatching fish is theoretically not perfectly random (risk-averse fish are less likely to get selected/caught) but that's the best method in those circumstances and it's reasonable to argue that the effect is insignificant.
- nkurz 3y agoYou make a very weak argument, and are simply assuming the conclusion. What makes it the "best method"? Would it be better to use a seine, or a trap, or hook-and-line? How would we know if there are subpopulations that have different likelihood of capture by different methods? To say it's "reasonable to argue that the effect is insignificant" is purely assertion. Why is it unreasonable to argue that a fish could learn from the first experience and be less likely to be captured a second time? If what you mean is that it's better than a completely blind guess, then I'd agree. But it's not clearly the best method nor is it clearly unbiased.
- egeozcan 3y agoFair points. But, mark-recapture is about practicality. It's not perfect, but it's a solid compromise between accuracy and feasibility (so I mean best in these regards, to be 100% clear). Sure, different methods might skew results, but this technique is about getting a reliable estimate, not pinpoint accuracy. As for learning behavior in fish, that's considered in many studies (and many other things, like listed here: https://fishbio.com/fate-chance-encounters-mark-recapture-studies-open-fish-population/ https://fishbio.com/fate-chance-encounters-mark-recapture-st... ), but overall, it doesn't hugely skew the population estimates. So, again, it's about what works best in the field, not in theory.
- lanstin 3y agoIn my experience conservation biologists are really good at finding animals in the wild. Much better than a typical SWE or typical business person.
- 3y ago
- deleted 3y ago[deleted]
- justinpombrio 3y agoThat's not actually the technique the authors are using. Catching 100 fish would be analogous to "sample 100 YouTube videos at random", but they don't have a direct method of doing so. Instead, they're guessing possible YouTube video links at random and seeing how many resolve to videos. In the "100 fish" example, the formula for approximating the total number of fish is: total ~= caught / tagged (where caught=100 in the example) In their YouTube sampling method, the formula for approximating the total number of videos is: total ~= (valid / tried) * 2^64 Notice that this is flipped: in the fish example the main measurement is "tagged" (the number of fish that were tagged the second time you caught them), which is in the denominator. But when counting YouTube videos, the main measurement is "valid" (the number of urls that resolved to videos), which is in the numerator.
- ad404b8a372f2b9 3y agoDid you understand where the 2^64 came from in their explanation btw? I would have thought it would be (64^10)*16 according to their description of the string. Edit: Oh because 64^10 * 16 = (2^6)^10 * (2^4)
- dajonker 3y agoThe YouTube identifiers are actually 64 bit integers encoded using url-safe base64 encoding. Hence the limited number of possible characters for the 11th position.
- deleted 3y ago[deleted]
- pants2 3y agoThat's typically the Lincoln-Petersen Estimator. You can use this type of approach to estimate the number of bugs in your code too! If reviewer A catches 4 bugs, and reviewer B catches 5 bugs, with 2 being the same, then you can estimate there are 10 total bugs in the code (7 caught, 3 uncaught) based on the Lincoln-Petersen Estimator.
- mewpmewp2 3y agoBut this implies that all bugs are of equal likelihood of being found which I would highly doubt, no?
- pants2 3y agoYes, it's obviously not a perfect estimate, but can be directionally helpful. You could bucket bugs into categories by severity or type and that might improve the estimate, as well.
- cpeterso 3y agoA similar approach is “bebugging” or fault seeding: purposely adding bugs to measure the effectiveness of your testing and to estimate how many real bugs remain. (Just don’t forget to remove the seeded bugs!) https://en.m.wikipedia.org/wiki/Bebugging https://en.m.wikipedia.org/wiki/Bebugging
- rightbyte 3y agoOh this is a really interesting concept. I guess it underestimates the number hard to find bugs though since it assumes same likelyhood to be found.
- krackers 3y agoAlso related is the unseen species problem (if you sample N things, and get Y repeats, what's the estimated total population size?). https://en.wikipedia.org/wiki/Unseen_species_problem https://en.wikipedia.org/wiki/Unseen_species_problem http://www.stat.yale.edu/~yw562/reprints/species-si.pdf http://www.stat.yale.edu/~yw562/reprints/species-si.pdf
- midasuni 3y agoIt’s not even new in the YouTube space as they acknowledge from 2011 https://dl.acm.org/doi/10.1145/2068816.2068851 https://dl.acm.org/doi/10.1145/2068816.2068851
- layer8 3y agoThat's only vaguely the same. It would be much closer if they divided the lake into a 3D grid and sampled random cubes from it.
- neurostimulant 3y ago> You generate a five character string where one character is a dash – YouTube will autocomplete those URLs and spit out a matching video if one exists. Won't this mess up stats though? It's like a lake monster randomly swapping an untagged fish with tagged fish as you catch them.
- MBCook 3y agoThis would find things like unlisted videos which don’t have links to them from recommendations.
- trogdor 3y agoThat’s a really good point. I wonder if they have an estimate of the percentage of YouTube videos that are unlisted.
- gaucheries 3y agoI think YouTube locked down their APIs after the Cambridge Analytica scandal.
- herval 3y agoin the end, that scandal was the open web's official death sentence :(
- m1sta_ 3y agoThe issue wasn't the analytics either. The issue was the engagement algorithms and lack of accountability. Those problems still exist today.
- TeMPOraL 3y agoSo as usual, the exploitative agents get to destroy the commons and come out on top. We need to figure out how to target the malicious individuals and groups instead of getting creeped out by them to the point of destroying most of the so praised democratizing of computing. Between this and locking down the local desktop and mobile software and hardware, we've never got to having the promised "bicycle for the mind".
- ganzuul 3y ago[flagged]
- bossyTeacher 3y agono one promised you anything
- fallingknife 3y agoAnd what kind of accountability is that? An engagement algorithm is a simple thing that gives people more of what they want. It just turns out that what we want is a lot more negative than most people are willing to admit to themselves.
- hipadev23 3y ago[flagged]
- blackle 3y agoIt is a little more sophisticated. They say they use an exploit that was found where a URL with five characters with a dash will get autocompleted by YouTube (I wonder why that is.) That improves sampling by 32,000 times apparently
- deleted 3y ago[deleted]
- m463 3y ago> Google started locking down the API almost immediately after we published our study Isn't this ironic, given how google bots scour the web relentlessly and hammer sites almost to death?
- dotandgtfo 3y agoThis is one of the most important parts of the EUs upcoming digital services act in my opinion. Platforms have to share data with (vetted) researchers, public interest groups and journalists.
- vasco 3y agoFor aggregated data and stats like this I think it could be fully publicly available.
- hotstickyballs 3y agoVetted always means people with the time, resources and desire to navigate through the vetting process, which makes them biased.
- lillecarl 3y agoI would argue it's better than nothing, and what are they going to be biased towards?
- fsckboy 3y agoAre you talking about Europe? they're certainly going to be biased against Google and any US tech giant. I'm biased against Google, but I'm honest about it. I don't ask "what could I possibly be biased about?"
- weeblewobble 3y agoYou might say the same thing about doing research in general
- 3y ago