4 ms·
A battered Mac PowerBook 170 :) The CEO, spez, is a founder and YC alumn. His heart is in the right place. He has to be weighing (a) board pressure to lock do
by throwaway9274 3y ago
A battered Mac PowerBook 170 :)
The CEO, spez, is a founder and YC alumn. His heart is in the right place.
He has to be weighing (a) board pressure to lock down content in the face of its increased value as private ML training data, against (b) the community impact.
He’s betting the (a) training data revenue is higher than the risk of (b) community flight due to (c) the platform-size network effect.
Maybe he’s right. Look at Twitter’s community retention. Maybe not. Look at Twitter’s revenue and Mastodon’s growth.
But if you go earlier back to BBS and the early web, it just never ends well.
I would urge him to look over a broad set of historical case studies and re-evaluate. Still salvageable.
- gremlinsinc 3y agoThese aren't mutually exclusive. (Training data and 3rd party apps api usage) He can provide FREE API access to 3rd party mobile apps, so long as they don't give data to openai/etc. That's a win/win. The way this is structured it's to completely squeeze out ALL 3rd party apps, and make a closed garden essentially. They're basically an open community trying to become AOL for $$.
- DSMan195276 3y agoIf the whole problem is ML companies using the data for free then why not give 3rd party clients free access to the API while making ML companies pay? spez already conceded to giving free API access to basically everybody else, if 3rd party clients aren't the real problem here then this whole thing is pretty pointless for him.
- TeMPOraL 3y ago(a) makes little sense to me. All the pristine data, i.e. Reddit up to late 2022, is already packaged up, mirrored, and available to download. Everything afterwards - and especially everything from the last two months onward - is increasingly becoming adulterated with LLM-generated content, which so far is proving poisonous to LLMs when fed to them in training. In other words: the widespread access to LLMs is rapidly destroying most of the ML training value Reddit data could offer.
- throwaway9274 3y agoThe theory would be stop the bleeding and try to clawback the old data with copyright enforcement against hosts of datasets and ML companies. Get the ML companies to license the content to settle the litigation, and include all new data as part of the deal. Then offer a similar “training only” content license to all other ML companies. Simultaneously you have to cut down on LLM-generated content, which may prove impossible. Where they botched it was the execution. It should have started with community engagement strategy: (a) buy out the 3P app developers, (b) fix the official app, and (c) offer low cost API access for the community by application, (d) shower mods and content creators with honorifics and if necessary rev share on their subs.
- quantified 3y agoAre the moderators getting paid? Or is Reddit making money off of their unpaid backs?
- dirkt 3y agoReddit is making money off their unpaid backs, and on top of that they want to deprive them of the tools the moderators use to do moderating, and make them pay for using those tools. That's the reason the moderators in particular are all very much for the strike...
- ryanbrunner 3y agoThe data for training seems not super well supported given that Reddit is still wide open to indexing even without authentication. Why bother with an API when you can crawl the site without issues?
- throwaway9274 3y agoIt’s possible it’s just about corralling everyone onto the official app where they can better ID users and utilize first-party adtech data. If so, then this is simply a dumb and bad move.
- dirkt 3y ago> lock down content in the face of its increased value as private ML training data That's easy to fix: Give out API keys; moderator tools, accessibility tools and third-party readers get free API keys. ML training data scrapers have to pay. No need to go to war with the community over that. Reddit has been constantly nagging users to use their App, so I really don't think it's about ML training data. It's about monetizing users by forcing them into their App.