4 ms·
I don't think you truly grasp how small this number is, this is actually good, like really really good. Github has about half a billion repositories.
by UrineSqueegee 3y ago
I don't think you truly grasp how small this number is, this is actually good, like really really good. Github has about half a billion repositories.
- Cupprum 3y agoGetting the actual number is probably very hard. These are the infected repos the OP found during their research.
- simonw 3y agoFor public repos you can get an approximate number by querying various public datasets. SELECT uniqHLL12(repo_name) FROM github_events; Against https://play.clickhouse.com/play?user=play#U0VMRUNUIHVuaXFITEwxMihyZXBvX25hbWUpIEZST00gZ2l0aHViX2V2ZW50czs= https://play.clickhouse.com/play?user=play#U0VMRUNUIHVuaXFIT... returns: 361648383
- jonnat 3y agoThey probably mean that the actual number of malicious repos is probably very hard to get. The article reaches the 100K number by searching for repos with patches with a particular string contained in this specific attack, so it's likely missing many malicious repos that use different methods of infection.
- arp242 3y agoNot only that, millions of these type of repos get created, and the vast majority are caught and deleted. The article mentions this: "Most of the forked repos are quickly removed by GitHub, which identifies the automation. However, the automation detection seems to miss many repos, and the ones that were uploaded manually survive. Because the whole attack chain seems to be mostly automated on a large scale, the 1% that survive still amount to thousands of malicious repos."
- mgii 3y agoNotice that as it seems, the vast majority are caught and deleted due to the intense automation, not the detection of malicious contents. If the actor was to run a smoother automation process, probably nothing would have been deleted. (disclaimer: author this article)
- bastardoperator 3y agoExactly, GitHub claims to have 400M+ repos making this number 0.025% of repos. I'm sure they could get it lower but less than half of 1% is pretty damn good. As a developer I have to do some due diligence about where I'm getting my data from. If I'm slurping in random repos because the name matches that's a people problem, not a github specific problem.