11 ms·
Context: I teach at Princeton and study social media and recommendation systems. From a very quick skim of the repositories, this appears to be quite limited t
by jonathanmayer 4y ago
Context: I teach at Princeton and study social media and recommendation systems.
From a very quick skim of the repositories, this appears to be quite limited transparency. The documentation gives a decent high-level overview of how Tweet recommendation works—no surprises—and the code tracks that roadmap. Those are meaningful positive steps. But the underlying policies and models are almost entirely missing (there are a couple valuable components in [1]). Without those, we can't evaluate the behavior and possible effects of "the algorithm."
[1] https://github.com/twitter/the-algorithm-ml https://github.com/twitter/the-algorithm-ml
- ngrilly 4y agoWhat did you expect?
- TaylorAlexander 4y agoI don’t know if the parent’s expectations matter here. This is more about making sure others don’t misunderstand the meaning here.
- ngrilly 4y agoGood point. I didn't see it like that. Thanks!
- fanagra32 4y ago[flagged]
- acdha 4y agoThe context is relevant for indicating that they’ve familiar with the problem and have thought about these issues in depth. It’s also useful for not being accused of hiding their identity if someone thinks they have an unmentioned agenda. Argument from authority is bad when it’s of the form “I am an expert, therefore you shouldn’t question this claim”, not when it’s used to provide an identity to a previously-unknown name while also providing a cogent argument and supporting evidence.
- tpmx 4y agoDid you also skim the accompanying (or rather, main) repo, https://github.com/twitter/the-algorithm https://github.com/twitter/the-algorithm ? From a quick clone and line-count, it has: 235 kLOC .scala 136 kLOC .java 22 kLOC .py 7 kLOC .rs So I don't think you did, since you posted so quickly and that's a LOT of code. I also haven't skimmed this code except very superficially, but perhaps you should since you're out there making statements with your Princeton credentials. (I posted this comment with the heads-up a few minutes after your comment above and then expanded it as you didn't respond.)
- Lord_Zero 3y agoI think you misunderstood. He's saying the training models are not there.
- eecc 4y agoWouldn’t that make them easy prey of “spam SEO”. However, given the framework isn’t it still possible to guess the models?
- makeitdouble 4y agoThe spam SEO issue should be dealt/thought about _before_ engaging in the whole adventure, and having to guess how it could work if decently implemented properly defeats the "open source" spirit of it. More credits would be given if the very idea of open sourcing the algorithm hasn't already been discussed to death with predictions of the difficult points and how it probably won't happen in any sane way.
- jimkleiber 4y agoMakes me wonder if a way to override people SEO hacking the algorithm is to create a market of open-source algorithms that each individual can choose and then it's not trying to hack THE algorithm but having to hack many and not knowing which algorithm an individual is using.
- muddi900 4y agoYou don't have to target each 'algorithm' all at once. You can target them one at a time. Hell you can run A/B test single out the easiest targets.
- jimkleiber 4y agoYes but right now there is 100% of the users using the one algorithm (or chronological). If one doesn't know what percentage or which people are using which algorithm, it becomes harder to know which ones to try to hack to have the biggest result.
- eecc 4y agoAnd them be pilloried for not doing it or not fast enough. Damn if you do, damn if you don’t. I’m starting to think the broblem with Elon is mostly personal, he’s just a proxy and default wrong. (not that I approve of his behaviors, but I can’t enjoy this whole mobbing that he’s getting; not that he cares this I’m not worried he’s getting traumatized in any way? it’s just how it’s become an identitarian trait for a certain group that irks me.)
- meghan_rain 4y agoSo why did they opensource it?
- daveguy 4y agoSo they could pretend to be open. It's the "Open"AI model. Open-washing?
- cubefox 4y agoThis is a very cynical take. They should be commended for publishing recommendation code at all, which no other major social network does.
- deleted 4y ago[deleted]
- hanniabu 4y agoThis is like FB open sourcing the compiled frontend code you can see yourself using inspect. If we commend them for this we're helping promote and encourage this faux open source virtue signaling
- cubefox 4y agoNo, that's very different.
- SequoiaHope 4y agoWell if they say “we will open source the algorithm” and then what they really open source is a little bit of slightly relevant code that doesn’t allow us to understand the algorithm, then what we can deduce is that they are trying to weasel out of public commitments. I can’t say for sure if that happened, but if they made a clear promise and then did something else, it’s perfectly reasonable to call that out.
- deleted 4y ago
- alfor 4y ago[flagged]
- bilekas 4y ago> But the underlying policies and models are almost entirely missing (there are a couple valuable components in [1]). Without those, we can't evaluate the behavior and possible effects of "the algorithm." Haven't gone through yet, but yeah, if that's the case, all this is, is a glorified framework to plug your own in.. Not exactly what was promised.
- modeless 4y agoWhat about these? https://huggingface.co/Twitter https://huggingface.co/Twitter
- simonw 4y agoThose look older to me. They all have last updated dates for October and November 2022.
- robopsychology 4y ago[flagged]
- tpmx 4y agoWe should really all just bow in awe as we are clearly inferior.
- robopsychology 4y agoPrinceton has a Code Reading 101 that all postdocs/professors must take, however in exchange for the Secrets of Speed Reading you must acknowledge every message with where you learnt those skills.
- elorm 4y agoYou missed this in your rush to display your newly acquired sarcasm101 skills: "Skim": To read quickly or cursorily, to glance over, or to omit details in order to get the gist of something.
- robopsychology 4y agoContext: I studied at Oxford Fair point, I missed that when I skimmed OPs comment
- pnt12 4y agoIt's fast to read stuff when you have the domain knowledge. The weights won't be a 5kb Scala file: they'd probably be a big binary file, which is easy to search it github/locally after cloning. Otherwise, if they are provided, someone in the thread will surely point to them.
- raggi 4y agoclass project, 200 students, 1500 LoC each. Time for grading. there are contexts in which this may be well practiced.
- culi 4y agoimagine thinking you need to read every file in a project to understand the architecture and which pieces are important for specific functionality you're looking to understand. Have you ever picked up a bugfix ticket for some code you didn't write?
- kadavy 4y agoFor example, MostRecentCombinedUserSnapshotSource seems to be influential (such as for calculating "tweepcred"), but we can't see how it's calculated.
- EastSmith 4y agoFB open source algo looks much better, right? /s
- bobobob420 4y agoCan i audit your classs for free?
- zhte415 4y agoIs it valid to focus tracking a Dem/Rep split when that split is an exclusionary design for many Americans? Or is it not exclusionary in your belief? I'm curious of a social science perspective. Ignoring the global nature of Twitter for a moment.
- novok 4y agoIt's an open algorithm, but it's not open data! (joking)
- eterevsky 4y agoI work on Google Assistant Suggestions and I don't think it's very practical to open-source an algorithm like that including the models and the underlying data. Both of them can live in separate services and be frequently updated. I am assuming that open sourcing the code aims to increase transparency about the business logic of the ranking decisions. At the same time you don't want spammers to be able to easily run experiments against a cloned version of your system.
- helsinkiandrew 4y ago> But the underlying policies and models are almost entirely missing... Without those, we can't evaluate the behavior and possible effects of "the algorithm And neither can spammers find and test the cracks and edge cases that would allow them to break the system, that does sound reasonable to me. If they were public there would be an arms race between spammers/those wishing to game the system and Twitter engineers.
- ivalm 4y agoThen don’t pretend to release “the algorithm.”
- helsinkiandrew 4y agoThey’re explaining how it works without giving the specifics. Much like the US military explains how the nuclear deterrent works without disclosing detailed plans and control codes.