5 ms·
Bluesky Social Dataset (235M posts from 4M users)
- 7d7n 2y agoPollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-generated content from Bluesky Social. The dataset contains the complete post history of over 4M users (81% of all registered accounts), totaling 235M posts. We also make available social data covering follow, comment, repost, and quote interactions. Since Bluesky allows users to create and bookmark feed generators (i.e., content recommendation algorithms), we also release the full output of several popular algorithms available on the platform, along with their timestamped “like” interactions and time of bookmarking. This dataset allows unprecedented analysis of online behavior and human-machine engagement patterns. Notably, it provides ground-truth data for studying the effects of content exposure and self-selection, and performing content virality and diffusion analysis.
- yawnxyz 2y agoI wonder how much time it takes to run this / what the script is / how resource intensive it is? Bsky is public right, so do you get rate limited? Do you scrape or use an official API? So many questions Also, I feel like only recently there's been an influx of people who have actually interesting things to say so I'd love to see nextyear's dataset
- viccis 2y agoNot sure about bulk export but you can set up a full stream of all activity without even registering an account.
- unshavedyak 2y agoBlows my mind that they can send that much for free.
- ks2048 2y agoI was checking out the Python API today (the "firehouse" via "atproto" package) and got 5000 posts in 7.5 seconds.
- verdverm 2y agoI believe they are enabling(ed?) filters so you can control how much and what you actually get from the firehose
- srik 2y ago[deleted]
- paxys 2y ago"Personal data" that was voluntarily published on a public microblogging platform with the explicit intention to share it with the world?
- srik 2y ago[deleted]
- ronsor 2y agoDon't make it public then.
- srik 2y ago[deleted]
- ronsor 2y agoIs it more entitled to observe public data than to willingly put data in the public and then expect to control the actions of others?
- paxys 2y agoIs you reading my comment on HN also entitlement? I certainly didn't give you permission to do it. It may have some personal details that I don't want you to see. Why do you think that is okay?
- deleted 2y ago[deleted]
- loeber 2y agoThe data is public by default. You know this when you sign up and use the service. This should inform your expectations of how the data will be used.
- ks2048 2y agoassociated paper: https://arxiv.org/abs/2404.18984 https://arxiv.org/abs/2404.18984
- yawnxyz 2y agohmm could you find the Github? I couldn't find it in the paper in the Code Availability section
- perihelions 2y agoIs it "scripts.tar.gz. A collection of Python scripts, including the ones originally used to crawl the data, and to perform experiments. These scripts are detailed in a document released within the folder" in the OP? The "code availability" says it's released "alongside [the dataset]", which appears to be the OP.
- yawnxyz 2y agooh good eye, I didn't catch that
- infotainment 2y agoI’m glad to see a new platform that isn’t completely locked down, allowing analysis like this. The trend toward everything being a walled garden is unfortunate.
- ilrwbwrkhv 2y ago[flagged]
- jph00 2y agoActually, since this isn’t locked up by the big copyright holders, we can all use it and profit.
- mplewis 2y agoHow will you use it to profit? You don’t have sweetheart cloud deals on ML training clusters. This benefits big players, not us.
- tomrod 2y agoMost impactful ML can be created on colab. Not Chatgpt, but most of the stuff not on the long tail.
- nl 2y ago"us" is relative. There are plenty of people on HN who have their own ML training clusters and aren't really big tech. For example natfriedman has https://andromeda.ai/ https://andromeda.ai/ And right now, today I can fine tune LLMs on this scale of data at home. In 5 or 10 years people will be able to training from scratch at home. Computational resource barriers are temporary. Licensing is forever.
- unshavedyak 2y agoNot all the profit, really. All would imply there was no value to begin with. I get the dislike, but i still comment on the open web because it has value to me. I'm still willing to answer questions on SO/reddit/etc because it has value to at least one (and hopefully more) people. That hasn't changed. Not sure what to say about companies making money off of my data.. but the posting itself doesn't seem to be that much of a negative. Thoughts? I see this sentiment a lot and it almost feels like "open" is bad these days. If anything i feel it almost is more important than ever.. as we're on the cusp of no need to ever go to forums/interact/etc.
- aussieguy1234 2y agoSound like this could be used to train an open source LLM.
- deleted 2y ago[deleted]
- skybrian 2y agoThe paper is from the end of April and they say the data was collected in February, March and April. I guess we can talk about it now, though. Due to high growth since then, this is from before most current users joined.
- deleted 2y ago[deleted]
- gusfoo 2y agoMeanwhile, over at Blueksy.app a few days ago, the users were incensed about a 1M-post data set and hounded the creator in to withdrawing it. https://bsky.app/profile/danielvanstrien.bsky.social/post/3lbu6l4fxdc2e https://bsky.app/profile/danielvanstrien.bsky.social/post/3l...
- saithir 2y agoBecause unlike the authors of this set - who went and stripped the posts out of usernames and permalinks to anonymize it - that set you mention just grabbed data out of the API as-is (at least based on its huggingface description that's left over). That's the difference.
- spiffytech 2y agoJust a reminder that anonymization is much harder than merely removing metadata: Every time I hear "anonymous data", I think of that time AOL published anonymized search logs (for academic research). The anonymization was negligent, and an NYT reporter de-anonymized and tracked down one of the users with the local & personal info present in the search queries. https://en.wikipedia.org/wiki/AOL_search_log_release https://en.wikipedia.org/wiki/AOL_search_log_release https://web.archive.org/web/20130404175032/http://www.nytimes.com/2006/08/09/technology/09aol.html?_r=1 https://web.archive.org/web/20130404175032/http://www.nytime...
- aubanel 2y agoPlease upload it on the Hugging Face Hub!
- abahlo 2y agoIf you just want to play around with the data, check out the bsky dataset on Axiom https://play.axiom.co/axiom-play-qf1k/stream/bsky https://play.axiom.co/axiom-play-qf1k/stream/bsky (700M+ events and counting)
- raidicy 2y agoI have returned back to this website to try and get the files and they have now been put under restrictive access for some reason.
- zft 2y agoIf you are interested in some real time visualization there are plenty of projects. For example http://www.graphtracks.com http://www.graphtracks.com. (I'm the author)