7 ms·
Getting 10TB of GitHub logs and extracting details of all users and repositories
- atum47 3y ago[flagged]
- amusingimpala75 3y agoEven https://archive.ph/atw1q https://archive.ph/atw1q didn’t work correctly on the page, it just ceases to scroll after a point.
- George83728 3y agoThe website works fine unless you enable javascript. That's usually the way it is with these sort of things. The webdev or CMS creates a perfectly functional website using HTML and CSS, then some javascript is added to shit the whole thing up. Disable javascript by default for a better web experience.
- zaric 3y agoThe noise is gone, enjoy your read! :)
- deleted 3y ago[deleted]
- mtmail 3y agoIMHO it's in the guidelines "Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- atum47 3y agodid not know about this part of the guidelines, thanks. a while back I was thinking about creating an extension that deals with this issue but I've heard some browsers are already working on that.
- didntcheck 3y agoI'm a big fan of neat statistical analyses, but when I look at the archive site https://www.gharchive.org/ https://www.gharchive.org/ , the overwhelming feeling I have is "creeped out". Taking periodic snapshots of repositories and their issues and wikis sounds good, but do we really need a log of every time someone watches an issue, and every commit message being irrevocably set on public record? That level of details on individual activity seems to have little value outside of cyberstalking I did have a look at the bottom of the page where prominent uses are listed, but nothing stands out as actually useful tbqh
- zx8080 3y ago> That level of details on individual activity seems to have little value outside of cyberstalking A selling point for business-level is that they can do whatever OKR they want on top of that details. To make any employee dance whatever business can imagine. Any amount of Tb seems OK until it helps selling "Business" tier github.
- ziml77 3y agoI have a GitHub account under my real name, but recently I've started using GitHub under a couple of other names instead. There's so much stuff you do in public on GitHub that I want to avoid people doing exactly this kind of analysis on. I wish using multiple identities was at least some level of foolproof though. I have to be careful to configure my local copies of repos to use the correct username, masked email, and PGP signing key. It would be super handy if git had a global config of multiple identities and if I could have it prompt me to select one either when I clone a repo or the first time I push to it.
- sleepytimetea 3y agoThe background with static noise really bothers me. Will have to skip reading till they provide a disable button.
- giancarlostoro 3y agoOn Firefox and even Microsoft Edge there is a "reader mode" option for most websites. I click on that often enough when I expect an article, to remove noise from ads.
- zaric 3y agoNo more noise!
- sleepychu 3y agoWas this written by GPT? I was quite interested in the topic of the article but I started to get the brain fog I associate with parsing ChatGPTs convoluted sentences.
- thefourthchime 3y agoIt also seems like a thinly veiled piece of product marketing.
- thomasjudge 3y agoThis
- wspeirs 3y agoAgreed... I wanted to understand what it was all about, but really struggled to follow. They talk about the whole thing taking around 24 hours, but some part took over 30. Also that it ran on a 4GB of RAM machine, but they needed larger ones to do all the parsing. Also in the end, unsure of what the actual results are. Maybe I missed clicking on something.
- zX41ZdbW 3y agoThe article leaves a bitter taste of unnecessary complexity. Data engineering should not be hard. For example, you can load the GitHub Archive to ClickHouse, and it will be accessible with interactive real-time queries: https://ghe.clickhouse.tech/ https://ghe.clickhouse.tech/ See also https://til.simonwillison.net/clickhouse/github-explorer https://til.simonwillison.net/clickhouse/github-explorer
- e9 3y agoYou can get this data for free on Snowflake: https://app.snowflake.com/marketplace/listing/GZTSZAS2KJ3 https://app.snowflake.com/marketplace/listing/GZTSZAS2KJ3 Caveat: you have to have Snowflake account