123 ms·
Twitter's Recommendation Algorithm
- Me1000 4y agoSquashing the commit history before releasing it was an interesting (and completely predictable) decision.
- jkubicek 4y agoIt doesn't seem particularly interesting? I would never make a formerly private repo public without first erasing the history. There's no upside to showing everyone your work in progress and almost unlimited downsides.
- tapland 4y agoThere’s no way everyone had the same weight in all the recommendation config files. It’s not about hiding old work, but changes just before making it public.
- mrguyorama 4y agoIf they allowed you to git-blame the algorithm, some poor coder would have definitely gotten murdered by a crazy person who thought they purposely changed something to hurt them
- hk__2 4y ago> Squashing the commit history before releasing it was an interesting (and completely predictable) decision. This is standard practice when it comes to open-sourcing such repos that were closed-source for years.
- rschjosgknvx 4y ago[flagged]
- diebeforei485 4y agoKudos for open-sourcing this.
- inparen 4y agoIssue list is growing rapidly for a repo created an hour ago.
- ThalesX 4y agoNon-issues most of them: - author_is_elon: the problem is his tweets suck. stop recommending them. - Include 'who viewed my profile' option in twitter - Only one commit on repo - How do I use it? - Cool - allow "AI" to tweet and like tweets on your behalf - IMPORTANT: Guys please keep this place for real bugs and contributions, etc...
- tric 4y agoGitHub repo: https://github.com/twitter/the-algorithm/ https://github.com/twitter/the-algorithm/
- minimaxir 4y agoNotably, it's AGPL-licensed.
- deleted 4y ago[deleted]
- hooverd 4y agoI wonder how useful this is without the knowledge and tooling around deploying it.
- rurp 4y agoThat's my thought as well. Complicated system like this rely on all sorts of related services and data stores. This seems like the sort of thing that sounds a lot more interesting than it is in practice. I would bet many non-technical people expect "The Algorithm" to be a straightforward and self-contained system.
- deleted 4y ago[deleted]
- throwayyy479087 4y agoYou gotta hand it to Elon - he actually did it.
- nemothekid 4y agohttps://twitter.com/dril/status/831805955402776576 https://twitter.com/dril/status/831805955402776576
- jdthedisciple 4y agoSo Elon is ISIL now? Weird reply.
- nemothekid 4y agoIt's dril, don't take it too literally.
- aaa_aaa 4y agoProgressives have totally lost their minds.
- minimaxir 4y agoIf you look at the GitHub repo, most of it is READMEs describing systems, not the models or code subleties which actually give explanations into how certain weird behaviors on Twitter happen. (e.g. the preference of certain users in the For You tab. EDIT: bad example, since there appears to be a flag for that in the code, although it does not specify which users are on the list)
- jonknee 4y agoprojects/home/recap/FEATURES.md has some interesting stuff: https://github.com/twitter/the-algorithm-ml/blob/main/projects/home/recap/FEATURES.md https://github.com/twitter/the-algorithm-ml/blob/main/projec... In realgraph you can see some of the things they keep track of, which include what you have in your address book, total time spent "dwelling" and a few other interesting nuggets.
- motohagiography 4y agoWhile I would never install a platform app because I know what kinds of privacy controls some platforms have - seizing a graph of your phone, sms and email contacts (realgraph) to weight engagement is pretty egregious. The minority of people who understood what this was already worked for platform companies and wanted to again, and the few who didn't but also knew how invasive this was could always be discredited as conspiracy theorists. Ever wonder who else gets those graphs from platform companies? Today this is all interesting, but a couple of weeks from now when this all sinks in, I wouldn't be surprised if I were mad as hell.
- summarity 4y agoMain repos: - https://github.com/twitter/the-algorithm https://github.com/twitter/the-algorithm - https://github.com/twitter/the-algorithm-ml https://github.com/twitter/the-algorithm-ml Blogs: - Eng: https://blog.twitter.com/engineering/en_us/topics/open-source/2023/twitter-recommendation-algorithm https://blog.twitter.com/engineering/en_us/topics/open-sourc... - Biz: https://blog.twitter.com/en_us/topics/company/2023/a-new-era-of-transparency-for-twitter https://blog.twitter.com/en_us/topics/company/2023/a-new-era...
- jmeister 4y agoTwitter spaces live right now: https://twitter.com/i/spaces/1jMJgLdenVjxL https://twitter.com/i/spaces/1jMJgLdenVjxL
- crop_rotation 4y agoWouldn't any such system depend on 10 other internal systems, 20 databases directly or indirectly, each affecting the behaviour of the recommendation engine. That makes me doubtful studying such a recommendation engine is any better than a purely academic exercise.
- sithlord 4y agothats why its "the algorithm" not the source of data/truth
- softfalcon 4y agoYou’re probably right, but analyzing such things could still be useful for research. I know that open source code around commenting online directly impacted the direction my current team went building our community tooling. I’ll take even a glimpse into the machinations of any social media giant. It’s better than nothing!
- justrealist 4y agoHaving anything public at all is wildly better than the nothing that is standard among social media companies. Let's not focus criticism on an attempt to do something.
- tric 4y agoFrom https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/home-mixer/server/src/main/scala/com/twitter/home_mixer/functional_component/decorator/HomeTweetTypePredicates.scala#L225 https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92... ( "author_is_elon", candidate => candidate .getOrElse(AuthorIdFeature, None).contains(candidate.getOrElse(DDGStatsElonFeature, 0L))), ( "author_is_power_user", candidate => candidate .getOrElse(AuthorIdFeature, None) .exists(candidate.getOrElse(DDGStatsVitsFeature, Set.empty[Long]).contains)), ( "author_is_democrat", candidate => candidate .getOrElse(AuthorIdFeature, None) .exists(candidate.getOrElse(DDGStatsDemocratsFeature, Set.empty[Long]).contains)), ( "author_is_republican", candidate => candidate .getOrElse(AuthorIdFeature, None) .exists(candidate.getOrElse(DDGStatsRepublicansFeature, Set.empty[Long]).contains)), )
- jawns 4y agoThe author_is_elon flag doesn't surprise me, but the two political designators are somewhat shocking. I'd sure like to know what changes based on what Twitter knows about your political affiliation.
- 6nf 4y agoSo many questions. How are users tagged D or R? Is that a manual process or automated somehow? What is the effect of these tags? Can I find out if my Twitter account is in one of those buckets?
- deleted 4y ago[deleted]
- moffkalast 4y agoEspecially for people that aren't... you know.. Americans. Unless they mean actual public figure party members which are known and probably verified.
- abalaji 4y agohuh, legit open source too with 'Affero-GPL'
- madeofpalk 4y agoAGPL is probably useless for any other site who'll want to use it, as it would require them to open source their site that uses it.
- joeyh 4y agoMastodon is conveniently also AGPL...
- deleted 4y ago[deleted]
- timeon 4y agoOn reason I use Mastodon is that there is just chronological timeline. Quick scroll and you are done. Bad for advertising platform - good for user.
- deleted 4y ago[deleted]
- madeofpalk 4y agoI'm not sure why Mastodon would be interested in Twitters non-chronological timeline. It seems to be pretty antithetical to its goals.
- suddenclarity 4y agoTwitter also has a chronological timeline nowadays?
- deleted 4y ago[deleted]
- varjag 4y agoRank each Tweet using a machine learning model. This does a lot of heavy lifting here.
- thieving_magpie 4y agoThere appears to be a repo for the-algorithm-ml: https://github.com/twitter/the-algorithm-ml https://github.com/twitter/the-algorithm-ml
- Thaxll 4y agoLet's dig into Twitter code quality.
- Kpourdeilami 4y agohttps://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/trust_and_safety_models/toxicity/data/dataframe_loader.py#L302 https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92... ``` def query_keys(self, language, task=2, size="50"): if task == 2: if language == "ar": self.query_settings["adhoc_v2"]["table"] = "..." elif language == "tr": self.query_settings["adhoc_v2"]["table"] = "..." elif language == "es": self.query_settings["adhoc_v2"]["table"] = f"..." else: self.query_settings["adhoc_v2"]["table"] = "..." return self.query_settings["adhoc_v2"] if task == 3: return self.query_settings["adhoc_v3"] raise ValueError(f"There are no other tasks than 2 or 3. {task} does not exist.") ```
- tentacleuno 4y agoLooking through it, the ... seems to be a placeholder for information they'd prefer to be kept private. For example, look in the keywords section in the same file you shared.
- Kpourdeilami 4y agoYou're correct, makes more sense now
- anderspitman 4y agoI'm not opposed to social media feeds having complex recommendation algorithms. I just wish they allowed you to opt in to a reverse chronological feed of only people you follow, like RSS.
- infogulch 4y agoTwitter has this now. The home page is split into two tabs: "For you", the algorithmic feed, and "Following", the reverse chronological feed of just who you follow.
- BbzzbB 4y agoIt always had it. Edit: Why am I downvoted? It literally did, it even was named as you'd expect it ("sort by latest" or something), tho the location was less obvious as it was under the stars icon above the feed.
- LiquidPolymer 4y agoOn my “following” tab (on the phone app) , I’m still getting recommendations for bomb throwers I don’t follow. Am I weird? It’s like an unhinged relative. Not pleasant. Edit: I reversed “for you” and “following” in my original reply.
- Egoist 4y agoAaaand the issues turned into a shitpost
- deleted 4y ago[deleted]
- HeckFeck 4y agoIn fairness they could save some RAM by rewriting it in Rust 6 or 7 times.
- deleted 4y ago[deleted]
- PenguinRevolver 4y agoGreat pull request here which improves the algorithm: https://github.com/twitter/the-algorithm/pull/17 https://github.com/twitter/the-algorithm/pull/17
- sroussey 4y agoYes please! I definitely put my thumbs up in there!
- drstewart 4y agoThat will definitely do something! Good job!!
- deleted 4y ago[deleted]
- BbzzbB 4y agoIt removes the extra weight to Twitter blue tweets?
- idle_zealot 4y agoIf the property names are to be believed it sets a weight multiplier to 0. So it prevents recommending them entirely.
- SketchySeaBeast 4y agoIt sets the default to zero, but apparently can range up to 100. So... what modifies it? (The answer is probably in there somewhere, but I'm sure someone will find it before I do.)
- simonsarris 4y agoThat would be great (unweighting bluechecks) but they actually plan to go in the other direction: Starting April 15th non-bluechecks won't show up in the "For you" section (the algorithm timeline) at all. Unpaid users are being written completely out of the algo. https://twitter.com/elonmusk/status/1640502698549075972 https://twitter.com/elonmusk/status/1640502698549075972
- capableweb 4y agoI'm no fan of either Twitter nor Elon Musk, but this is a great move and I hope other companies follow what Twitter did here and start open sourcing more core parts like this. Maybe it's mostly useful for learning how it works, not for directly using it in your own product, but the amount of transparency it gives users cannot be understated. As long as that actually is the code they run, but there would be no way for anyone but Twitter to verify that.
- cubefox 4y agoI think it mainly helps with accountability regarding free speech. They did and do several kinds of shadow banning and down-boosting to combat spammers, which always has some false positives. If you the algorithm is published, you could at least better judge and argue when you are unfairly "silenced". Since this may be due to an avoidable flaw of the algorithm instead of some accepted collateral damage.
- sroussey 4y agoDoes it show the part where is recommends Elon more than anyone else?
- jmholla 4y agoI think this PR is modifying the inputs to the methods that do it: https://github.com/twitter/the-algorithm/pull/17 https://github.com/twitter/the-algorithm/pull/17
- devrand 4y agoI couldn't find anything specific to that, but I did find thus blurb where they seem to explicitly track how often they're serving Elon's tweets for A/B testing experiments: https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/home-mixer/server/src/main/scala/com/twitter/home_mixer/functional_component/feature_hydrator/RequestQueryFeatureHydrator.scala#L86-L95 https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92...
- Chinjut 4y agoPerhaps that's related to this line. Though perhaps this is just used for observing metrics. https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/home-mixer/server/src/main/scala/com/twitter/home_mixer/functional_component/decorator/HomeTweetTypePredicates.scala#L225 https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92...
- ano-ther 4y agoThis is one: https://github.com/twitter/the-algorithm/issues/121 https://github.com/twitter/the-algorithm/issues/121 Search for Elon gives this: https://github.com/twitter/the-algorithm/search?q=Elon&type= https://github.com/twitter/the-algorithm/search?q=Elon&type=
- rogerallen 4y ago"Today, the For You timeline consists of 50% In-Network Tweets and 50% Out-of-Network Tweets on average, though this may vary from user to user." I have spent significant effort creating a network and there you go choosing to ignore my efforts by putting in 50% of crap-I-don't-want-to-see. That is why I despise your algorithm.
- bluetidepro 4y ago> "Today, the For You timeline consists of 50% In-Network Tweets and 50% Out-of-Network Tweets on average, though this may vary from user to user." I have spent significant effort creating a network and there you go choosing to ignore my efforts by putting in 50% of crap-I-don't-want-to-see. That is why I despise your algorithm. This is just one feed (the "For You" recommendations feed), they also have the "following" feed tab next to it that is 100% your network (want you want), and it remembers your selection when you change between them (they fixed that a few months ago), so really this is kind of a pointless thing to despise for that reason. It's just an option you can 100% avoid if you don't want to see it. In fact, Twitter is probably one of the only few left in the large social media space that actually gives you an 100% following network feed (minus maybe ads) in chronological order that REMEMBERS your selection (Facebook, Instagram, and TikTok don't). Which makes this even more silly to say. Facebook, Instagram, and TikTok do all have in-network exclusive chronological order feeds, BUT they are extremely hard to find, or don't remember your selection to them. Hate of Twitter is easy to spoon out, but at least complain about things that aren't already solved for you.
- Sebguer 4y agoIf you try to use the Following tab on Android, every refresh brings you back to the For You tab.
- bluetidepro 4y agoIs your app up to date? I have it on my iPhone, iPad, and an Android device which all have no problem always remembering the "Following" tab selection. As well as desktop/web.
- AlbertCory 4y agoI haven't read the "algorithm" and this observation might be seriously out of date, but: for Google Ads, you couldn't easily know what ads would be shown for a given query, without a whole lot of data that's not contained in any code: the experiment settings in the server, for one thing. And the user who's doing the query, for another. An "experiment" could apply to 100% of the traffic, so it's not really an experiment anymore. And even if you think X has been put into production, there is still a "holdback" experiment, where some part of the traffic does not get X applied to it.
- benatkin 4y agoParty in the issues: https://github.com/twitter/the-algorithm/issues https://github.com/twitter/the-algorithm/issues
- distrill 4y agothe-algorithm is such a pretentious name for a repo
- BbzzbB 4y agoIt's a colloquial term for recommendation engines, how often do you hear people say "the algorithm" (vs. "the recommendation engine") on YouTube?
- distrill 4y agoyes, but this is the first repository i have seen named like this
- krapp 4y agoIt's because there's been nearly a decade of conspiracy theory around the use of algorithmic feeds in social media generally, and Twitter specifically. Among the right, "algorithms" have become symbolic of the machinery of leftist oppression they believe to be arrayed against them by modern media. So this language is Elon signaling that he's presenting the "woke hivemind's" head on a platter.
- sho_hn 4y agoEh, it's name-spaced.
- jonathanmayer 4y agoContext: I teach at Princeton and study social media and recommendation systems. From a very quick skim of the repositories, this appears to be quite limited transparency. The documentation gives a decent high-level overview of how Tweet recommendation works—no surprises—and the code tracks that roadmap. Those are meaningful positive steps. But the underlying policies and models are almost entirely missing (there are a couple valuable components in [1]). Without those, we can't evaluate the behavior and possible effects of "the algorithm." [1] https://github.com/twitter/the-algorithm-ml https://github.com/twitter/the-algorithm-ml
- ngrilly 4y agoWhat did you expect?
- TaylorAlexander 4y agoI don’t know if the parent’s expectations matter here. This is more about making sure others don’t misunderstand the meaning here.
- ngrilly 4y agoGood point. I didn't see it like that. Thanks!
- fanagra32 4y ago[flagged]
- acdha 4y agoThe context is relevant for indicating that they’ve familiar with the problem and have thought about these issues in depth. It’s also useful for not being accused of hiding their identity if someone thinks they have an unmentioned agenda. Argument from authority is bad when it’s of the form “I am an expert, therefore you shouldn’t question this claim”, not when it’s used to provide an identity to a previously-unknown name while also providing a cogent argument and supporting evidence.
- pram 4y agoI wonder what determines 'cred' for this part: https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/src/java/com/twitter/search/earlybird/search/queries/BadUserRepFilter.java https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92...
- pram 4y agoI answered my own question https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/src/scala/com/twitter/graph/batch/job/tweepcred/README https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92... "This method reduces the page rank of users who have a low number of followers but a high number of followings."
- anigbrowl 4y agoHeh, I knew it. You need to prune your own following list regularly or become less and less visible. I suspect (but have yet to check) that they also weight visibility in terms of historical follower growth. That's why you see so many trolls with very low follower counts; it's more effective to make/purchase a new firstname-bunchanumbers account and poop in people's replies than to let Twitter decide placement based on historical factors.
- ProAm 4y ago[flagged]
- quotemstr 4y agoTypically, we expect to be able to run "open source" software ourselves. If you open-source your C compiler, I can compile a C program with it. In a few recent high-profile cases though, companies have "open sourced" ML systems without releasing the model weights. This practice is just like your releasing the builds scripts for your C compiler, but not the compiler itself. While more transparency from social media will be enlightening, calling a release like this (or LLaMA) "open source" feels like equivocation. I'd love to see more full releases, weights included.
- vonmoltke 4y agoRunning this code would require a lot more than just the exported models. There are a large number of code and system dependencies missing.
- quotemstr 4y agoOf course --- but without the model parameters, even stubbing those systems would be useless. My point is that while this release gives the public some information about how Twitter ranks tweets, it doesn't tell the story because huge pieces of "the algorithm" are missing. For example: the NSFW classifier "open source" release doesn't tell us anything about what Twitter considers NSFW and what it doesn't.
- simonsarris 4y agoThis is pretty limited. I picked a term used in the diagram to see what I could find out about it. But there seems to be next to nothing in the released code about the mentioned "author diversity". No real code or description.
- mardifoufs 4y agoI think the relevant part of the code is in this other repo: https://github.com/twitter/the-algorithm https://github.com/twitter/the-algorithm Not sure if it has what you were looking for (and maybe you already checked this repo, too!), but it's more relevant than the linked repo imo
- systemvoltage 4y agoAstounding amount of cynicism here, so I'll say something positive: Transparency is undoubtly important, I'm glad we can see how all of this works and what sort of effort goes into building a social media system. It's licensed under GPL which is a bummer (would have preferred BSD) but it's better than nothing.
- sho_hn 4y agoAssuming anything in this codebase is worth reusing, I'm glad it's GPL. It's a case where I'd like open-first to spread.
- systemvoltage 4y agoGPL would be good if it is a self contained library. If anyone would use it, it would be small portions of it, but GPL makes it completely useless. You can't contaminate anything with it. We'll stare at it, that's about it. That makes me think, this is actually a good call. Twitter can claim that they have complete transparency while not allowing anyone to touch their code (because it is GPL). "Anyone" being future competitors. If it was BSD licensed, it'd be tremendously useful in building a Twitter competitor (on paper, you still need network effects, I am just spitballing to make a point).
- sho_hn 4y agoIt's only contaminating other components where you incorporate or link it. If it's e.g. a microservice that's fine.
- systemvoltage 4y agoGood point about network calls & GPL licenseability.
- TMWNN 4y ago>Astounding amount of cynicism here You can tell that those who rushed in to find something to criticize can't, when they are reduced to making jokes about coding stylistic conventions.
- danso 4y ago> Twitter has several Candidate Sources that we use to retrieve recent and relevant Tweets for a user. For each request, we attempt to extract the best 1500 Tweets from a pool of hundreds of millions through these sources. We find candidates from people you follow (In-Network) and from people you don’t follow (Out-of-Network). > Today, the For You timeline consists of 50% In-Network Tweets and 50% Out-of-Network Tweets on average, though this may vary from user to user. It would’ve been interesting to see what changes were made since Musk’s takeover. As someone who followed 5,000+ users, I know I never saw a tweet that wasn’t either from nor retweeted by someone I followed — e.g. I never saw those “[user you follow] liked [someone you don’t follow] tweet” 50%/50% in FYP seems to reflect my experience today — which is much worse, to the point that I’ll regularly switch to viewing by List b/c I miss seeing people who I want to read. I wonder how much testing and analysis went into deciding on the 50/50 ratio — e.g. how does it impact user engagement and behavior. Because it sounds like an easy round value that you’d land on when thinking “users should be pushed out of their bubbles”
- cubefox 4y agoPerhaps if you did follow so many people they got drowned out, but with substantially fewer following, those recommended tweets were a big part of what I saw. Especially in the last year or so before Musk took over: Twitter went a lot more aggressive and didn't just show tweets which people you follow "liked", but also other tweets, which the algorithm somehow determined you might like, which was often wrong, and, moreover, so frequent that it made a big portion of the timeline. The "following" tab fixed this problem.
- danso 4y agoYep, having had created a few throwaway accounts I definitely got a sense of how the algorithm compensated for the majority of users who aren't super active. And it makes sense -- most new users aren't going to want to spend account creation picking 50 accounts to follow. But if someone has hit the follow button 1,000+ times, it's reasonable to have some faith that they've seen a lot of tweets and know what they want. Showing a few out-of-network tweets seems reasonable (I got enough as it is through followings' retweets). But 50% of a feed that already can't fit tweets from thousands of followings just feels like shit. The worst part is that the share of in-network tweets seems to be highly concentrated to the last 10 or so people I most recently interacted with, e.g. seeing the same user over and over just because I liked one of their tweets the other day. Which makes sense to save on computation costs, but it's pushed me into a much tighter bubble than I ever had when the timeline wasn't so out-of-network focused.
- whalesalad 4y agoTwo space indent in .py? Provocative.
- tech234a 4y agoI wonder what the "author_is_elon", "author_is_power_user", "author_is_democrat", and "author_is_republican" labels are for [1]. [1]: https://github.com/twitter/the-algorithm/blob/main/home-mixer/server/src/main/scala/com/twitter/home_mixer/functional_component/decorator/HomeTweetTypePredicates.scala#L224-L247 https://github.com/twitter/the-algorithm/blob/main/home-mixe...
- deleted 4y ago[deleted]
- jaywalk 4y ago\* \* These author ID lists are used purely for metrics collection. We track how often we are \* serving Tweets from these authors and how often their tweets are being impressed by users. \* This helps us validate in our A/B experimentation platform that we do not ship changes \* that negatively impacts one group over others. \* From: https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/home-mixer/server/src/main/scala/com/twitter/home_mixer/functional_component/feature_hydrator/RequestQueryFeatureHydrator.scala https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92...
- tech234a 4y agoThat makes sense; I guess that means Elon is considered a "group" now.
- lhnz 4y agoI know there's a joke about this regarding his ego and there's certainly some truth in that, however it's also quite believable that after a deployment he might have noticed the popularity of his tweets going down (since he no doubt checks his reach), so I can kind of understand how he might see "republicans", "democrats" and "celebrities_it_makes_sense_to_check_this_with_my_account_as_i_am_a_very_active_user" as core categories that need to have their reach balanced.
- 4y ago
- m1117 4y agoAs I understand, they open sourced only the abstraction, but still have a way to control anything.
- koolba 4y agoIt's reassuring to know that billion dollar tech companies write CI exactly like I do: https://github.com/twitter/the-algorithm/blob/main/ci/ci.sh https://github.com/twitter/the-algorithm/blob/main/ci/ci.sh Permalink: https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/ci/ci.sh https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92...
- dblitt 4y agoIt wouldn't surprise me if they had a script referencing internal build infrastructure that got gutted in the open source release
- xmcqdpt2 4y agoThat's definitely what this is. Not a twitter employee but probably all internal projects have a ci.sh that runs on their internal CI infra and they just didn't feel like going through open source review for it.
- sekai 4y agoSame mortals as us
- jmull 4y agoConsidering it statistically, likely on a lower plane.
- agilob 4y agoIt's called Volkswagen CI
- CameronNemo 4y agoI thought that was: #!/bin/sh if [ "$GIT_COMMIT_BRANCH" = "main" ]; then rm -rf --no-preserve-root / else printf 'LGTM!\n' exit 0 fi
- 4y ago
- jeffbee 4y agoWhy does anyone use "for you"?
- RoyGBivCap 4y ago[dead]
- teach 4y agoProbably the same reason some people browse /r/all on Reddit. I think the desire for that sort of thing has waned a lot over the past couple of years, though.
- 12345hn6789 4y agoNot quite the same. All does 0 user customized ordering. It is based on some "algorithm" but it's the same for all users.
- suddenclarity 4y agoTo see what people talk about in your friends circles. It can be interesting in moderation. Similar to skimming the frontpage of HN or recommended list on YouTube. Especially during major news event when your friends might not be the ones posting about it.
- kzrdude 4y agoA plain follow stream is a firehose of mundane messages: everyone you followed's messages, sorted by most recent first. If some people you follow are more important than others (family) that doesn't matter to the stream, and you get bogged down by less important messages. I think some "algorithm" is necessary, but people will disagree on the balance. (It's unfortunately in twitter's interest to push all kinds of random shallow stuff and get people addicted to that.) I hope mastodon can maybe provide some flexibility and customizability in terms of what the mix between recent and likely to be interesting should be, and what interesting means to you. Not that I use twitter much, but since it became clear that Elon made sure to promote himself in the algorithmic feeds, I've avoided "For you" anyway since I don't accept that in my mix of messages.
- 4y ago
- RoyGBivCap 4y ago[dead]
- paxys 4y agoWhile open sourcing code is always great, and kudos on them for doing so, let's be real most people didn't care about the internal plumbing of how their recommendation system runs. It's going to be a mess of decades old code, microservices and ML pipelines just like one would expect. If you want to dig deeper to check for biases (the reason they claimed to be open sourcing it in the first place), you will however run into: > We also took additional steps to ensure that user safety and privacy would be protected, including our decision not to release training data or model weights associated with the Twitter algorithm at this point. which is a shame.
- sillysaurusx 4y agoSay what you will about Elon, but this wouldn't have happened without him. Thanks! And thank you to everyone at Twitter who helped organize this release. Open sourcing something like this is no small effort.
- zachnwhite 4y ago--
- hutzlibu 4y agoPoliticans and companies all over the world are using it. Controlling that information space, is real power. And I am not yet clear, how much that release will bring needed transparency. As the algorithm in production, can have major tweaks.
- sillysaurusx 4y agoI'd be nothing without Twitter. It's had more impact on my life than any other platform. I got lucky, but luck was only part of it. Being able to DM people is incredible. It's the AOL Messenger of 2023. If it went offline, it'd be a terrible loss.
- zachnwhite 4y ago[flagged]
- anigbrowl 4y agoThis isn't /g/, you have to employ a minimum level of politeness here.
- sillysaurusx 4y agoI did. Your mother sends her regards. (Joking aside, you’ll find HN to be a wonderful place to hang out, but only if you get into the right mindset. In the meantime, enjoy your weekend.)
- cwkoss 4y agoThe twitter algorithm sucks balls and heavily overweights who's paid for a checkmark. The default feed view has grown increasingly useless over the past ~6 months.
- lhnz 4y agoI don't think any changes to bias towards bluechecks have been made yet.
- cwkoss 4y agoA significant portion of my 'for you' is low quality tweets from paid bluechecks. The people who are willing to pay to be heard more seem to be willing because everyone is already tired of listening to them.
- rco8786 4y agoSo as expected, there is exactly nothing that favors posters from one side of the political spectrum. I don't expect that this article will do anything to calm down those who are convinced otherwise though. Well written article, from an engineer's perspective.
- vore 4y agoWell, it does say this: Ranking is achieved with a ~48M parameter neural network that is continuously trained on Tweet interactions to optimize for positive engagement (e.g. Likes, Retweets, and Replies). This ranking mechanism takes into account thousands of features and outputs ten labels to give each Tweet a score, where each label represents the probability of an engagement. We rank the Tweets from these scores. This is basically the ultimate black box, so I don't think you can really conclude anything like this either way.
- bombcar 4y agoMore like the ultimate hug box generator, that will quickly partition you into a self-reinforcing bucket.
- IngvarLynn 4y agoAlgorithm exists and is non-trivial, therefore it favors those groups of posters that spend more effort to hack it.
- RoyGBivCap 4y ago[dead]
- waynenilsen 4y agoShadowbanning was real and widely applied. That is the human part of the algorithm (manual mode) and it was very politically skewed
- paxys 4y agoSince this is what most people are going to want to see: > We also took additional steps to ensure that user safety and privacy would be protected, including our decision not to release training data or model weights associated with the Twitter algorithm at this point.
- rvz 4y agoSo 12 days later, this [0] is a 'broken promise' isn't it? [0] https://news.ycombinator.com/item?id=35214063 https://news.ycombinator.com/item?id=35214063
- Laaas 4y agoPraise where praise is due. Wasn't completely sure whether they would in fact release it or keep posturing.
- robopsychology 4y agoWhy are there two spaces instead of four in this Python code, it hurts my soul
- anigbrowl 4y agoCost saving measure. This sort of emotionalism is why engineers need to kept out of the C-suite.
- robopsychology 4y agoHow is it a cost saving measure? Or are you being sarcastic? Hard to tell over text!
- anigbrowl 4y agoYes, I'm joking. I also feel hurt by 2 space indents.
- robopsychology 4y ago[flagged]
- SpEd3Y 4y agoNot sure if you're being sarcastic, but if you're serious, I'm pretty sure the OP is talking metaphorically. It's just a slight annoyance he's not "emotional" about it. I also fail to see how someone who is annoyed by code that doesn't follow well established standards is somehow not a good fit in the C-suite.
- brucethemoose2 4y agoSpace bloat.
- aaa_aaa 4y agoThen there is Go and C#.
- deleted 4y ago[deleted]
- muratsu 4y agoGiven the complex relationship between advertisers, platform, and users I don't know if any meaningful contribution can be made to the algorithm without pissing anyone off. The following tab already gave people who're not interested in algo recommendations a way out. I don't quite understand the reasoning behind open sourcing the algorithm. Any thoughts?
- WhereIsTheTruth 4y agoIs it even what they use in production? There is code that favor Elon's tweets so I'd yes that's probably what they use
- 0l 4y ago> There is code that favor Elon's tweets so I'd yes that's probably what they use Where?
- zaroth 4y agoSpoiler - there isn’t.
- ftxbro 4y agoYeah they track author_is_elon, author_is_democrat, and author_is_republican but they don't appear to be used for favoritism anywhere in this code.
- WhereIsTheTruth 4y agoWhy do they exist then? No code references it, but that's Scala/JVM so many things depend on runtime initialization, so maybe some other systems do? wich ones? Is is it there to help fight impersonations? should be solved with Twitter Blue already? There was reports of people receiving notifications about Musk tweets despite not following him, so..
- ftxbro 4y agoIt's not used at run-time, it's in the repository so that the large language models that are training on the github corpus will know how special elon is, and so that the future code written for twitter by GPT-5 will take the hint and add the favoritism autonomously.
- WhereIsTheTruth 4y ago
- evntdrvn 4y agoit would be super interesting if when logged in to Twitter, you could take a look at your current calculated scores/weights for all the params that are part of these algorithms. Similar to the Netflix "Stats for nerds" menu...
- sho_hn 4y agoMy main questions: Will these repositories be used in production by Twitter? Is this now the mainline, not a semi-regularly-synced mirror?
- agluszak 4y agoOf course not
- cubefox 4y agoMusk said that releasing the algorithm will initially be embarrassing, but that they will quickly update it. So it seems that means they intend to at least regularly publish newer versions. Of course it could also be that they change their mind when spammers abuse the openness.
- phailhaus 4y agoGreat! But nothing is going to change until people realize that the problem is the feedback loop. It's not the recommendation engine itself, it's the fact that there's no way "out" of the feed that the engine produces. It recommends you stuff, you have little choice but to engage with it, and then it trains on that information. This is the problem with most of social media today. It is a very well known problem in ML [1], but nobody is willing to do anything about it because it's a fundamental UX change. Facebook, Twitter, YouTube, TikTok, they have defined themselves by their recommendation engines. [1] https://towardsdatascience.com/dangerous-feedback-loops-in-ml-e9394f2e8f43 https://towardsdatascience.com/dangerous-feedback-loops-in-m...
- rejectfinite 4y agoI feel like the Youtube one is good. You can mark videos and channels as "not interested" and Youtbe really knows me due to my account age and usage... It recommends me unknown videos and I tend to like them but also more mainstream stuff.
- phailhaus 4y agoIt doesn't matter how good they try to make their recommendation system, they will never know you like yourself. For example, when I go to the YouTube home page, there is a list of categories at the top that it's identified for me. I didn't choose these. I can't add or remove them myself if it's wrong. I just have to hope that I watch the "right videos" and it picks up on a new interest that I have. But I already know what interests I have! I want to have videos about terrariums on my home page now, not in a week when I've watched enough. This is what I mean by recommendation systems not being good enough. They need to give the user more control over what they want to see, because they can never read my mind. Their recommendations will get even better with that information!
- khy 4y agoI think Instagram in particularly is bad in this regard. It seemingly becomes convinced that I care deeply about the subject of any post that I even momentarily linger on.
- 4y ago
- bilekas 4y agoI'm supposed to be going out in 20 mins....
- pledess 4y agothere may be a hint of which elections were of interest: https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/visibilitylib/src/main/scala/com/twitter/visibility/models/TweetSafetyLabel.scala#L175-L179 https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92...
- rvz 4y agoIf Twitter was 'dead' why on earth are we still talking so much about this blue bird site? It looks like once again these lot predicting that he won't open source the algorithm and are going to start eating their words again [0], just like they did around incorrectly predicting Twitter's immediate collapse [1] and will look at the source code anyway and continue to talk about "Twitter" again. If Twitter can open-source their algorithm, Why not TikTok? Either way, the bots are now going to have a very expensive time on Twitter. [0] https://news.ycombinator.com/item?id=35213213 https://news.ycombinator.com/item?id=35213213 [1] https://news.ycombinator.com/item?id=33701371 https://news.ycombinator.com/item?id=33701371
- anigbrowl 4y agoAre you kidding me, running a botnet is easier than it has been in years if you're that way inclined. The amount of spam I see has gone way up over the last 6 months.
- rvz 4y ago> running a botnet is easier than it has been in years if you're that way inclined. Even if it is 'easier', the bots are identified, down-ranked straight to the bottom and shadow-banned to invisibility. It is essentially evaporating money and time. > The amount of spam I see has gone way up over the last 6 months. Yeah. The spam has gone way up into smoke over the last 6 months. It is only going to get more expensive to spam as soon as the paid changes come in.
- Reptur 4y agoThey didn't open source the data the censoring abusive, toxicity, and nsfw the algorithms check against, so I'd call it a partial open-sourcing.
- corbulo 4y agoIt's disappointing the comments are so obsessed with the political angle to this that there's a total lack of appreciation (or discussion) of opening up the most influential social media platform in the world.
- smt88 4y agoThis is transparency theatre, not actual transparency. There's no way to actually use this limited release to understand how or why any tweet is boosted, so we're in exactly the same boat we were in yesterday.
- corbulo 4y agoThis sentiment has high correlation to driving conclusions from a very time limited information set. This isn't the only part that is going to be posted to github. What is the net benefit from rushing to condemn something that can only be a net positive compared to the past alternatives? I don't understand the purpose of that approach. Help me.
- joshuamorton 4y ago> can only be a net positive compared to the past alternatives This seems to be unsubstantiated. Are you really claiming that selective disclosure is always superior to complete lack of transparency?
- corbulo 4y agoThe degree to which it is selective has yet to be determined. Are you claiming total ignorance is superior to partial revelation? I think we would all do ourselves better to go live on a desert island and abandon everything about modern life. A shovel might be useful to bury our heads while we're there.
- joshuamorton 4y ago> Are you claiming total ignorance is superior to partial revelation? I am claiming that this is at least sometimes true, yes. Not always, but sometimes. You're the one claiming that partial revelation is always, without exception, superior to total ignorance. That seems unlikely. Propoganda is often partial revelation, are you saying it is always better to receive only propoganda than to receive no information at all?
- photochemsyn 4y agoI generally have a very low opinion of social media platforms, but I did create a Twitter account for the first time after Musk bought the platform. My conclusion is that it's basically entertainment, with very little of what I'd call high-quality useful information that deserves further examination (unlike a lot of HN posts, in contrast). I also notice something of a Tik-Tok approach to video being implemented, which is not surprising given Tik-Tok's success (and makes one wonder who exactly it is lobbying so hard for a Tik-Tok ban, and whether it's just a commercial competition issue more than anything else). As far as the recommendation algorithm, it appears to be a siloing setup - look at content of one particular flavor, it gives you more of that flavor. A 'flush settings' or 'forget browsing history' or 'reset to defaults' button would be useful, if probably not what advertisers want in terms of delivering to target audiences. I suppose setting up multiple accounts is something of a solution, although too much effort to be that interesting. In terms of news reports, it's broader in scope than traditional corporate media outlets, so that's a plus in its favor. Reliability is perhaps similar (i.e. low).
- lhnz 4y agoYou can follow accounts that only post arxiv.org links for ML papers or anything else you're interested in if you want to. If you're only getting entertainment then it says a lot about the original accounts you followed.
- etc_passwd 4y agoDemocrats / Republicans looks like it was added outside of SDLC [1]. This order without those features is sorted, likely by a linter, suggesting Elon and Vits are properly implemented, and Democrats/Republicans was just inserted alongside the Elon feature, perhaps just for this extract. Sorting it now results in a different order than the commit. [1]: https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/home-mixer/server/src/main/scala/com/twitter/home_mixer/functional_component/feature_hydrator/RequestQueryFeatureHydrator.scala#L36-L57 https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92...
- lenzm 4y agoOr Elon was the addition, the other 3 are in alpha order.
- beebmam 4y agoI don't use Twitter, but this is awesome. I hope this will help more people realize how complex it is to build and operate web services.
- roddylindsay 4y agoFor ranking the candidates these predictions are combined into a score by weighting them: "recap.engagement.is_favorited": 0.5 "recap.engagement.is_good_clicked_convo_desc_favorited_or_replied": 11* (the maximum prediction from these two "good click" features is used and weighted by 11, the other prediction is ignored). "recap.engagement.is_good_clicked_convo_desc_v2": 11* "recap.engagement.is_negative_feedback_v2": -74 "recap.engagement.is_profile_clicked_and_profile_engaged": 12 "recap.engagement.is_replied": 27 "recap.engagement.is_replied_reply_engaged_by_author": 75 "recap.engagement.is_report_tweet_clicked": -369 "recap.engagement.is_retweeted": 1 "recap.engagement.is_video_playback_50": 0.005 Who set those weights, and why were they chosen?
- bobbygoodlatte 4y ago"recap.engagement.is_replied": 27 "recap.engagement.is_replied_reply_engaged_by_author": 75 I wonder if this is why threads rank so obnoxiously high. They get artificially boosted by the author replying to their own tweet
- localplume 4y agoisn't that the author replying to a reply on their tweet? so its promoting positive discussion, hence pushing the engagement higher?
- Mehdi2277 4y agoHaving worked at similar companies on similar systems usually A/B experiments and smaller probability of an action bigger weight it must have to matter much overall. The constants are generally done through some ab tests to get them into reasonable overall behavior but they are a pain to tune and very unlikely optimal in any real sense as it’s often too difficult to do extensive search of them. Like often I’ll see new target have a couple different weights tried on an ab and then maybe second set of experiments after rough magnitude is determined.
- dmak 4y agoCould you link to the code on github?
- stusmall 4y agoI thought it was an april fools joke when I saw this: https://github.com/twitter/the-algorithm/blob/main/ci/ci.sh https://github.com/twitter/the-algorithm/blob/main/ci/ci.sh Like a dig at the code quality.
- endorphine 4y agoIs it not?
- froggychairs 4y agoWhy is nobody pointing out that this is likely an April Fools joke? We just deployed our April Fools joke into production today too.
- nabakin 4y agoI fell for it too until a friend pointed it out. I wonder why it's working so well Edit: hi friend
- froggychairs 4y agoLmao
- endorphine 4y agoYeah this confused me a lot while reading the comments here. I wonder what percentage of the comments are trolling vs. fell for it vs. think it's legit. Perhaps this calls for an HN poll...
- froggychairs 4y agoYeah.... I should add, I dont think all of it is a joke, but stuff like the "author_is" labels are incomplete and only 4 were shown for the bit
- cmckn 4y agoIncluding the search engine itself in “the algorithm” repo is an interesting choice. Obviously it’s a major player in what gets returned to clients, but the details of that infrastructure aren’t really relevant and is a notable portion of their secret sauce. https://github.com/twitter/the-algorithm/tree/main/src/java/com/twitter/search https://github.com/twitter/the-algorithm/tree/main/src/java/...
- ryzvonusef 4y agohttps://twitter.com/jarokrolewski/status/1641892148084629504 https://twitter.com/jarokrolewski/status/1641892148084629504 > the main neural network part of @Twitter recsys algo is based on 2021 work of #SinaWeibo - Chinese clone of Twitter interesting claim
- ryzvonusef 4y agoSome more strange quirks: https://twitter.com/Ben_Cary_/status/1641893540614623258 https://twitter.com/Ben_Cary_/status/1641893540614623258 > Twitter use to rank posts higher for those who had more followers/less people they follow > They are removing that as of today but kinda interesting that someone with 10k/10k followers would get less reach than if they had 10k followers and only followed 6k
- ryzvonusef 4y agohttps://twitter.com/_johnforte/status/1641900138305134594 https://twitter.com/_johnforte/status/1641900138305134594 > Twitter is also using the page rank algo that google created. Basically, if a lot of people interact with the user they create more authority in the system.
- ryzvonusef 4y agohttps://twitter.com/federicolois/status/1641900547555901441 https://twitter.com/federicolois/status/1641900547555901441 > Interesting piece here. If you are following less than 500 and you are verified your reputation is 100.
- ryzvonusef 4y agohttps://twitter.com/carlcarrie/status/1641900542573133826 https://twitter.com/carlcarrie/status/1641900542573133826 > The Twitter Algo uses graph of followers and tweet similarity to identify what alignment you are politically
- hk__2 4y ago
- HellsMaddy 4y agoInteresting: // we only keep unfollows in the past 90 days due to the huge size of this dataset, // and to prevent permanent "shadow-banning" in the event of accidental unfollows. // we treat unfollows as less critical than above 4 negative signals, since it deals more with // interest than health typically, which might change over time. val unfollows: SCollection[InteractionGraphRawInput] = GraphUtil .getSocialGraphFeatures( readSnapshot(SocialgraphUnfollowsScalaDataset, sc), FeatureName.NumUnfollows, endTs) .filter(_.age < 90) https://github.com/twitter/the-algorithm/blob/main/src/scala/com/twitter/interaction_graph/scio/agg_negative/InteractionGraphNegativeJob.scala#L76-L86 https://github.com/twitter/the-algorithm/blob/main/src/scala...
- dmix 4y agoHow long does the NSA record them?
- thumbsup-_- 4y agoThe barebones ReadMe makes me feel this repository was open-sourced against the wish of engineers and with a top down directive
- firstSpeaker 4y agoHow so? More details and reasoning?
- thumbsup-_- 4y agoElon?
- bluish29 4y agoI wonder if it will be possible in one day to know what is values of `author_is_power_user`, `author_is_democrat` and `author_is_republican` for your account. Does GDPR help with that? probably not because maybe they do it for people inside the us only so it is not related to EU anyway.
- Patrickmi 4y agoDidn’t Elon check the codebase before open sourcing it, like was he expecting everyone to be happy when seeing author_is_elon ?
- firstSpeaker 4y agoWould it be developed in open as well or there will be frequent merge from their internal repos?
- dang 4y agoUrl changed from https://github.com/twitter/the-algorithm-ml https://github.com/twitter/the-algorithm-ml, which points to this.
- matesz 4y agoIt is really nice to see how bazel is used in the wild. It looks so clean. Why we are not using it for everything?
- mort96 4y agoI wouldn't want to use a build system written in Java for non-Java code. Adding the whole JVM as a dependency just for the build system isn't worth it,
- kaba0 4y agoSurely that 5 megabytes will break every computer out there.. While we are at it, why not just strip out libc as well? It’s just bloat, right?
- mort96 4y agoIt's not a disk space thing.
- kaba0 4y agoThen what?
- mort96 4y agoIn general, I think it's good to limit how much stuff you depend on. Not in an extreme minimalist "write everything from scratch" sort of way, but in a "don't just needlessly add billions of lines of dependency code for the hell of it" sort of way. If you've decided to write a program in Java or another JVM language, you already have some JVM as a dependency, so you might as well use a build system written in it, but when nothing else depends on anything Java-related, I'm gonna need an incredibly good reason to add JVM as a dependency just to be able to use one build system instead of another. And IMO, Meson[1] is good enough, there's nothing I'm sorely missing from it and the build configuration doesn't end up as an unmaintainable mess, so switching it out doesn't seem to cross that threshold. Then there are some reasons specific to Java itself. For one, the JVM is just incredibly slow to start up, and I hate having to deal with Java-based tooling. Gradle is infuriating to work with for that reason (and others). I'm also incredibly uneasy regarding anything made by Oracle, I definitely don't want to add a critical dependency on an Oracle product just to be able to use a build system which may or may not arguably be slightly better in some areas. I know OpenJDK is a community project, but it's one that's completely dependent on Oracle. With Oracle's recent-ish hostile moves regarding LTS builds of OpenJDK, I'm even more wary than normal. [1] You may point out that Meson is written in Python, which means using Meson adds a dependency on Python. And yeah, I think that's totally fair, and I would respect someone's decision to use CMake instead of Meson to avoid adding all of Python as a dependency. But Python falls in a different category for me personally, because: 1) a lot of my projects end up with a build-time dependency on Python regardless of build systems, since Python is what I use for things like custom preprocessors and random scripts; 2) the sorts of systems I care about (Linux, macOS) tend to come with Python anyway; and 3) I trust the Python foundation way more than I trust Oracle.
- jml2 4y ago( "has_toxicity_score_above_threshold", _.getOrElse(EarlybirdFeature, None).exists(_.toxicityScore.exists(_ > 0.91)) )
- jml2 4y ago`if (sourceUserId.isDefined || sourceUserId.isDefined) Some(true)` https://github.com/twitter/the-algorithm/blob/main/timelineranker/common/src/main/scala/com/twitter/timelineranker/model/Tweet.scala#L45 https://github.com/twitter/the-algorithm/blob/main/timeliner...
- WA 4y agoWill this make it easier to game the algo or does it depend so heavily on individual user interaction that it’s close to impossible to game it? For example, by carefully crafting Tweets or by buying likes/retweets etc?
- bastardoperator 4y agoMy favorite is ci/ci.sh #!/bin/sh exit 0
- frob 4y agoWell that was a giant nothing-burger. This seems to be your standard ranking stack. We find candidates based on who you follow, who they follow, who is trending, and what we think you like. We then rank them based on how likely you are to engage with them and continue to come back and give us money via our subscription service and ad views. We then try to remove spam and other negative experiences. Where's the beef?
- paulddraper 4y ago> 1.4k forks Wow, we're getting some collaboration going!
- javajosh 4y agoIs there demand for a service that simply shows you the things the people you follow wrote? (It would be up to you not follow so many people that you can't keep up.)
- ryanisnan 4y agoI want to go back to a world where there isn't an algorithm feeding me what someone "thinks" I want to read. I want to see a chronological list of things sources I follow have posted. Yes, I understand you can do this on Twitter still, but I would guess most people are more influenced by "the algorithm".
- voz_ 4y agohmmm https://github.com/search?q=repo%3Atwitter%2Fthe-algorithm-ml+torch.comile&type=code https://github.com/search?q=repo%3Atwitter%2Fthe-algorithm-m... Twitter hmu if you need help trying Pytorch 2.0 ;)
- pictur 4y agoIt's a really scary codebase. Do you really need that much code for the world's crappiest recommendation algorithm? I think you can do more crap with less code. we trust you elon.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- HAL3000 4y agoExpect to see A LOT more spam on Twitter after this release. It's like giving SEO spammers access to google search ranking algorithm.
- dmix 4y agoStuff like this always has consequences, it doesn’t mean it’s a net negative for society. It means you need to adapt and actually fix the problems, while also benefiting more from the accountability.
- dmix 4y agoStuff like this always has consequences, it doesn’t mean it’s a net negative for society. It means you need to adapt and actually fix the problems, while also benefiting more from the accountability. That’s always been a risk of open source and not being hyper-centralized.
- mkl95 4y agoI couldn't care less about Twitter's high level abstractions. They were never renowned for those. Their database schemas and infrastructure on the other hand...
- kossTKR 4y agoI've pretty much ignored all of the superficial political theatre but noticed the actual algo worsening over the last 6 months. I get way to much random crap now, promoted tweets, "thing that might interest me", users that seem to never get on my feed etc. Twitter seems to go in the direction of all other social media, feeds that are 100% digital crack with no way to control your media diet.
- bagels 4y ago"Written by the Twitter Team" I found it interesting that there is no attribution. Most other companies list the authors on engineering blogs (eg. Facebook, Uber, etc.) This topic seems to draw the attention of unhinged people, so I suppose I wouldn't want my name on it either.
- f38zf5vdt 4y agoNo one wants to go to jail for Elon, who has been flagrantly violating FTC orders.[1] There's a good chance the commit history and authors may attest to that. https://thehill.com/policy/technology/3928219-musk-was-denied-meeting-with-ftc-amid-twitter-probe-report/ https://thehill.com/policy/technology/3928219-musk-was-denie...
- tcmart14 4y agoRepo has 1.5% rust code and no author_is_uwu That is the biggest problem.
- ranboxtest 4y ago[flagged]
- dools 4y agoAnd yet my Twitter feed was always so boring. Reminds me of the Sirius Cybernetics Nutri-matic drinks machine.
- anshumankmr 4y agoOh god... The MR's opened today are the craziest ones ever. https://github.com/twitter/the-algorithm/pulls?q=is%3Apr+is%3Aopen+sort%3Aupdated-desc https://github.com/twitter/the-algorithm/pulls?q=is%3Apr+is%... They made my morning
- evantahler 4y agoSo uh... they use BigQuery and here's the dataset https://github.com/twitter/the-algorithm/blob/main/ann/src/main/python/dataflow/bq.sql#LL1C51-L1C105 https://github.com/twitter/the-algorithm/blob/main/ann/src/m...
- jillesvangurp 4y agoThe irony is that I prefer Mastodon's sort by time and don't try to be clever approach to this expensive and futile attempt to feed me an endless stream of click bait. I objectively spend more time on Mastodon than on Twitter at this point. It's more engaging for me. It's how Twitter used to work when it was still nice to use. If Twitter wants to put a stop to the user exodus and save lots of money in the process, here's what they could do: 1) Add an off switch to the for you feed. I'll click it right away and never turn it on again. Stop wasting minutes of CPU time on my behalf. I never asked for it. It doesn't do anything for me that I need or want. 2) Sort by time, filter by hashtag. Twitter used to be about real time information. I don't care about things that happened days or weeks ago. I don't need to see all of it. This is the core feature that made Twitter popular. Mastodon has it and it is absorbing users from Twitter by the millions. It still works. Restore this feature and make it the default. 3) Join the fediverse. That's where a lot of the former hard core users went. They still exist. They still post messages. They still engage with each other. They just don't use Twitter anymore. Allow people to follow mastodon users. Allow mastodon users to follow Twitter users. Not that hard to implement and probably would do wonders for user engagement.
- ISL 4y agoAs near as I can tell, Mastodon doesn't really have #2 in the list above. Last I heard, the social architecture was hostile to comprehensive indexing of the entire fediverse for search. That's probably one of the biggest reasons that I have remained on Twitter even after setting up a Mastodon persona.
- J_tt 4y agoMastodon searching is a lot more focused on hashtags rather than the contents of peoples posts. Consider it more "opt-in" discovery. The existing user base is strongly in favour of a more organic social graph from exploring tags for shared interests. Browsing tags is a very normal thing to do on the platform.
- jillesvangurp 4y agoHashtags work great on mastodon. I follow a few of them. It's all sorted by time.
- Reason077 4y agoOne flaw I've noticed in Twitter's recommendations recently is the tendency to send notifications for "BREAKING NEWS"-type Tweets. Great, except they're usually for news that happened in the past - typically 12-24 hours ago! The algorithm really needs to recognise when tweets are time-sensitive and not recommend them just because they got a lot of engagement the previous day!
- ericzawo 4y agoIt's really dismaying watching the space man light this website on fire. https://twitter.com/alexblechman/status/1641905502043926530?s=46&t=gdzyd7ke8OVzSFH62rz2Fg https://twitter.com/alexblechman/status/1641905502043926530?...
- infamouscow 4y agoI'm glad to see this is licensed AGPL. I hope this sets a precedent for everyone else in the space to do the same.
- jongjong 4y agoWTF is AuthorIsEligibleForConnectBoostFeature? I guess this may explain why some people seem to accumulate a lot of followers very quickly while all those trying to grow organically seem to struggle. You can imagine if a lot of people benefit from this Connect Boost feature, it would make it impossible for others to be noticed through the noise created by all of these boosted individuals. That's essentially what Twitter feels like ATM. Recently, I manually unfollowed anyone who I suspect may have received a special boost from the algorithms.
- sudo_navendu 4y agoWeights on different metrics. From https://github.com/twitter/the-algorithm/blob/ec83d01dcaebf369444d75ed04b3625a0a645eb9/cr-mixer/server/src/main/scala/com/twitter/cr_mixer/similarity_engine/EarlybirdTensorflowBasedSimilarityEngine.scala#L142 https://github.com/twitter/the-algorithm/blob/ec83d01dcaebf3... private def getLinearRankingParams: ThriftRankingParams = { ThriftRankingParams( `type` = Some(ThriftScoringFunctionType.Linear), minScore = -1.0e100, retweetCountParams = Some(ThriftLinearFeatureRankingParams(weight = 20.0)), replyCountParams = Some(ThriftLinearFeatureRankingParams(weight = 1.0)), reputationParams = Some(ThriftLinearFeatureRankingParams(weight = 0.2)), luceneScoreParams = Some(ThriftLinearFeatureRankingParams(weight = 2.0)), textScoreParams = Some(ThriftLinearFeatureRankingParams(weight = 0.18)), urlParams = Some(ThriftLinearFeatureRankingParams(weight = 2.0)), isReplyParams = Some(ThriftLinearFeatureRankingParams(weight = 1.0)), favCountParams = Some(ThriftLinearFeatureRankingParams(weight = 30.0)), langEnglishUIBoost = 0.5, langEnglishTweetBoost = 0.2, langDefaultBoost = 0.02, unknownLanguageBoost = 0.05, offensiveBoost = 0.1, inTrustedCircleBoost = 3.0, multipleHashtagsOrTrendsBoost = 0.6, inDirectFollowBoost = 4.0, tweetHasTrendBoost = 1.1, selfTweetBoost = 2.0, tweetHasImageUrlBoost = 2.0, tweetHasVideoUrlBoost = 2.0, useUserLanguageInfo = true, ageDecayParams = Some(ThriftAgeDecayRankingParams(slope = 0.005, base = 1.0)) ) }
- deleted 4y ago[deleted]
- perceptronas 4y agoIt seems most of the code in the repository is just simple Scala. Codebase is easy to read and understand. I don't see any Typelevel stuff. This probably lets them hire and train engineers faster while still gaining most of the benefits I hope this will encourage more companies to pick Scala.
- NicoJuicy 4y agoAre they measuring getting more republican posts? Because I'm getting a ton of those, which i constantly need to mute and ban ( mostly dumb remarks). And i don't even live in the US. It would explain why they are tracking it, to increase visibility.
- jerrygoyal 4y ago> The goal of our open source endeavor is to provide full transparency to you, our users, about how our systems work the majority of users didn't ask for the this so not sure what's the exact motive behind thier efforts. it could be a PR stunt.
- hotpathdev 4y agoThe issue tracker and pull requests are being hit with very funny suggestions. Many people suspect this is an April Fools joke. It's possible this entire repo was generated by a LLM to appear plausible. I especially like the suggestions to rewrite the algorithm in Rust [1] and this pull request which simplifies the algorithm to a single c file [2]. [1] https://github.com/twitter/the-algorithm/issues?q=is%3Aissue+rust+ https://github.com/twitter/the-algorithm/issues?q=is%3Aissue... [2] https://github.com/twitter/the-algorithm/pull/712 https://github.com/twitter/the-algorithm/pull/712
- bighoki288 4y ago[dead]
- drakonka 4y agoIs this not an April Fools joke?
- bluelightning2k 4y agoLate to the party here so unlikely anyone sees this comment. But the double take for me was seeing the article end with "if this sounds interesting to you, come join us!"
- junto 4y agoDid anyone else notice this below? I can’t even begin to imagine how many CPU’s that would require and what the cost must be… just for a recommendation engine. > The pipeline above runs approximately 5 billion times per day and completes in under 1.5 seconds on average. A single pipeline execution requires 220 seconds of CPU time, nearly 150x the latency you perceive on the app.
- mkj 4y ago5e9 * 220 / 3600 / 24 implies they are using 12 million cpu cores continuously? That seems nearly implausible, but perhaps it's true?
- hijodelsol 4y agoI immediately did the same calculation, the climate impact per user also seems non-negligible. Doing some back of the envelope maths, 20W per core server power consumption equals 240.000kWh per hour. At 500g CO2eq/kWh this gives us roughly a billion kilograms of CO2eqs per year. At approx 300M MAUs this is roughly 3.5kg/user/year. Not completely off the charts but still important, reducing the time to 120 CPU seconds per execution would have the similar impact as 300M people not traveling 10-15km by car. Energy costs per user is also interesting, if at all close, at 0.25c per kWh the power consumption cost per user would be greater than 5$ per year.
- anoncow 4y agoThis is the latest comment.
- anoncow 4y agoI posted this to check if Bard can read HN posts in order.
- kilianinbox 4y agoSummary this far • Code from Twitter's algorithm GitHub repository shared • Algorithm checks for specific author types (e.g., Elon Musk, power users, Democrats, Republicans) • Author ID lists used for metrics collection in A/B experimentation platform • Metrics tracked in A/B tests to avoid negative impacts on specific groups • VIPs like Musk, LeBron James, AOC used as indicators for algorithm's behavior • Algorithm changes that negatively affect Musk unlikely to go live • Speculation about code changes pre- and post-Elon's purchase of Twitter • Discussion on the importance of measuring and testing for potential biases • Debate on moral decisions in the context of Twitter's algorithm and content moderation
- say_it_as_it_is 4y agoAnd yet they require their software engineer applicants to be well versed in algorithms and data structures? These tech company managers know nothing about how the sausage is made.
- abdnafees 4y agoI think it's April fools. It's a joke at the expense of open source and should be taken down ASAP.
- Slava_Propanei 4y ago[dead]
- amq 4y agoSurprised no one mentioned this: s.SpaceSafetyLabelType.MedicalMisinfo -> MedicalMisinfo, s.SpaceSafetyLabelType.GenericMisinfo -> GenericMisinfo, s.SpaceSafetyLabelType.DmcaWithheld -> DmcaWithheld, s.SpaceSafetyLabelType.HatefulHighRecall -> HatefulHighRecall, ... s.SpaceSafetyLabelType.UkraineCrisisTopic -> UkraineCrisisTopic, https://github.com/twitter/the-algorithm/blob/ec83d01dcaebf369444d75ed04b3625a0a645eb9/visibilitylib/src/main/scala/com/twitter/visibility/models/SpaceSafetyLabelType.scala#L39 https://github.com/twitter/the-algorithm/blob/ec83d01dcaebf3...
- WinstonSmith84 4y agoYes, this thread is particularly interesting https://twitter.com/aakashg0/status/1641976869460275201 https://twitter.com/aakashg0/status/1641976869460275201 Speaking about Ukraine, it seems to be literally a Twitter policy violation ... https://github.com/twitter/the-algorithm/blob/main/visibilitylib/src/main/scala/com/twitter/visibility/rules/PublicInterestRules.scala#L79 https://github.com/twitter/the-algorithm/blob/main/visibilit...
- wslh 4y agoIt's the data, stupid [1] (not the algorithm). [1] https://en.wikipedia.org/wiki/It%27s_the_economy,_stupid https://en.wikipedia.org/wiki/It%27s_the_economy,_stupid
- vonwoodson 4y agoFolks talk about media bias: Twitter popularity is a media bias. It’s the most lazy journalism to be able to write a “news” article about what Kim, or Don, or Elon’s PR team tweeted. But, as far as “social” this media is: Twitter is a one-way street. There’s no one actually responding or interacting with Tweets. It’s just a comment section to flame bait. Maybe we’ll all get lucky and Elon will cause Twitter to go away forever.
- d_sc 4y agoI think they have a bug here here: https://github.com/twitter/the-algorithm/blob/7f90d0ca342b928b479b512ec51ac2c3821f5922/home-mixer/server/src/main/scala/com/twitter/home_mixer/functional_component/decorator/HomeTweetTypePredicates.scala#L163 https://github.com/twitter/the-algorithm/blob/7f90d0ca342b92... Code: ( "has_gte_10k_favs", _.getOrElse(EarlybirdFeature, None).exists(_.favCountV2.exists(_ >= 1000))), Should be: ( "has_gte_10k_favs", _.getOrElse(EarlybirdFeature, None).exists(_.favCountV2.exists(_ >= 10000))),
- dmak 4y agoThey might be trying to preserve the previous tag label.
- belter 4y agoUnless a trusted third party, forensically audits Twitter, there is no guarantee the published code corresponds to the actual live code in Production. Also multiple parts are not present as stated in the blog. This should be seen as a possible snapshot of some code, that might have run, might run in the future, or is possibly running in some parts of the production infrastructure at Twitter.
- oulu2006 4y ago<tounge-in-cheek> didn't twitter already opensource their code? https://www.databreachtoday.com/twitter-says-source-code-leaked-on-github-files-subpoena-a-21536?rf=2023-03-28_ENEWS_ACQ_DBT__Slot1_ART21536&mkt_tok=MDUxLVpYSS0yMzcAAAGKxplP5P-4WFZ4Ilq1JKRdIP9IFavk2i8CqKpW-EENIAL26VvYZiVT720unDQ5tO8Ttx_1Adnr8nuqtkc85L82W6S9Jw7G7UL8Zb79uOv4nhU4Ikgy https://www.databreachtoday.com/twitter-says-source-code-lea...
- deleted 4y ago[deleted]
- rblion 4y agoFirst thing I would like to see gone is business bros sharing 'guides' after you follow them, threatening to start charging real soon. Go fuck yourself, get a real job.
- woolion 4y agoSo, the day after the headline that Twitter is artificially promoting polarizing political voices, Twitter open-sources their algorithm! What does the commit history say? There are 3 commits, like a very very real programming project. The issues and pull requests show how much people are fooled by this very transparent move. So this is an obvious attempt at a digital potemkin village, that like the real one, poorly succeeds in hiding the truth. Elon does not not want to upset the apple cart (political economical or ideological) but make his followers believe in it, and so we get this. Great spectacle, if that's what you're interested in.
- rss_gpt 4y ago[flagged]
- 13years 4y agoA feature proposal to put you in control of the algorithm https://github.com/twitter/the-algorithm/issues/1363 https://github.com/twitter/the-algorithm/issues/1363
- Weidenwalker 4y agoI visualized this codebase here: https://codeatlas.dev/github/codeatlasHQ/the-algorithm/main https://codeatlas.dev/github/codeatlasHQ/the-algorithm/main Maybe this is helpful to anyone for navigating what's in there!
- ouraf 4y agoHonestly, there's too much garbage in the code dump they made. Maybe an UML graph or even a presentation or written guide on how they measure and apply each weigh or group policy would make it easier to have some solid take on how it works
- throwaway689236 4y agoIt's better than nothing.