4 ms·
Given DeepSeek's open philosophy I wonder what their response is to simply being asked for access to the code and data that this project intends to recreate?
by fblp 2y ago
Given DeepSeek's open philosophy I wonder what their response is to simply being asked for access to the code and data that this project intends to recreate?
- weatherlite 2y agowouldn't hold my breath ...
- thih9 2y agoWhile I'm also interested in this, I guess there is value in independent replication as well. Assuming this is doable - and I wouldn't know. Does anyone know how difficult it is to perform this kind of reproduction? E.g. how much time would it take (weeks? years?) and how likely it is to succeed?
- papichulo2023 2y agoNo company ever will disclose data due it would open endless liability.
- jackjeff 2y agoThat’s a good point. Wouldn’t OpenR1 suffer from the same problem? Or does being open somehow shield them from legal repercussions?
- michaelt 2y agoSome people believe they can dodge copyright issues so long as they have enough indirection in their training pipeline. You take a terabyte of pirated college physics textbooks and train a model that can pose and answer physics 101 problems. Then a separate, "independent" team uses that model to generate a terabyte of new, synthetic physics 101 problems and solutions, and releases this dataset as "public domain". Then a third "independent" team uses that synthetic dataset to train a model. The theory is this forms a sort of legal sieve. Pass the knowledge through a grid with a million fact-sized holes and with enough shaking, the knowledge falls through but the copyright doesn't.
- svnt 2y agoKnowledge laundering
- maartenpi_ 2y agoExactly. Meta won't do it for the same reason. Liability alone, imagine all the copyright lawsuits... Secondly the dataset for now has a lot of competitive advantage. In a way it seems like a good thing that AI giants compete on methodology now.
- fblp 2y agoInteresting, so they wouldn't want to disclose something that shows they've illegally (terms / copyright violations) scraped research databases for example. Won't this eventually come up in legal discovery when someone sues one of these firms for copyright infringement? They'd have to share their data in the discovery process to show that they haven't infringed..