8 ms·
Well that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers https://openrouter.ai/deep
by jacob019 1y ago
Well that didn't take long, available from 7 providers through openrouter.
https://openrouter.ai/deepseek/deepseek-r1-0528/providers https://openrouter.ai/deepseek/deepseek-r1-0528/providers
May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.
Fully open-source model.
- jazzyjackson 1y agoNo sign of what source material it was trained on though right? So open weight rather than reproducible from source. I remember there's a project "Open R1" that last I checked was working on gathering their own list of training material, looks active but not sure how far along they've gotten: https://github.com/huggingface/open-r1 https://github.com/huggingface/open-r1
- chrsw 1y agoBased on commit history Open R1 still active and they're still making progress. Long may it continue, it's an ambitious project.
- behnamoh 1y ago> No sign of what source material it was trained on though right? out of curiosity, does anyone do anything "useful" with that knowledge? it's not like people can just randomly train models..
- ToValueFunfetti 1y agoWould be useful for answering "is this novel or was it in the training data", but that's not typically what the point of open source is
- marci 1y agoWhen you're trully open source, you can make ethings like this: Today we introduce OLMoTrace, a one-of-a-kind feature in the Ai2 Playground that lets you trace the outputs of language models back to their full, multi-trillion-token training data in real time. OLMoTrace is a manifestation of Ai2’s commitment to an open ecosystem – open models, open data, and beyond. https://allenai.org/blog/olmotrace https://allenai.org/blog/olmotrace
- kreijstal 1y agoyou can do these same, except you would need to be a pirate website. It would even be better. except illegal. but it would be better.
- marci 1y agoThat is why the others can't provide stuff like this. RAG/Hallucination check. I just wish Allen.AI models had bigger context, 4k is too small nowadays.
- m00x 1y agoMany are speculating it was trained by o1/o3 for some of the initial reasoning.
- fulafel 1y agoAre there any widely used models that publish this? If not, then no I guess.
- DANmode 1y agoDepending on how you use "randomly", they absolutely can..?
- anonymoushn 1y agoIf labs provided the corpus and source code for training their tokenizers, it would be a lot easier to produce results about tokenizers. As it is, they provide neither, so it is impossible to compare different algorithms running on the same data if you also want to include the vocabs that are commonly used.
- make3 1y agoI don't think people make the distinction like that. The open source vs non open source distinction boils down to, usually, can you use it for commercial use. what you're saying is just that it's non reproducible, which is a completely valid but separate issue
- piperswe 1y agoBut where's the source? I just see a binary blob, what makes it open source?
- jacob019 1y agoThe weights are the source. It isn't as though something was compiled into weights. They're trained directly. But I know what you mean, it would be more open to have the training pipeline and souce dataset available.
- timschmidt 1y agoThe weights seem much more like a binary to me, the training pipeline the compiler, and the training dataset the source.
- jumski 1y agoCome here to write this - perfect analogy!
- reedciccio 1y agoIt's very imperfect analogy though these things can't be rebuilt "from scratch" like a program, the training process doesn't seem to be replicable anyway. Nonetheless, full data disclosure is necessary, according to the result of the years-long consultation led by the Open Source Initiative https://opensource.org/ai https://opensource.org/ai
- 1y ago
- pradn 1y agoIsn't it basically not possible for the input data set list to be listed? It's an open secret all these labs are using immense amounts of copyrighted material. There's a few efforts at full open data / open weight / open code models, but none of them have gotten to leading-edge performance.
- bee_rider 1y ago“Not possible” = “a business-destroying level of honesty”?
- tokioyoyo 1y agoThere is a "keep doing what you're doing, as we would want one of our companies to be on top of the AI race" signal from the governments. It could've been stopped, maybe, 5 years ago. But now we're way past it, so nobody cares about these sort of arguments.
- rcxdude 1y agoEven if training on the copyrighted material is OK, just providing a data dump of it almost certainly is not.
- alpaca128 1y agoNo need for a data dump, just list all URLs or whatever else of their training data sources. Afaik that's how the LAION training dataset was published.
- anonymoushn 1y agoproviding a large list of bitrotted URLs and titles of books which the user should OCR themselves before attempting to reproduce the model doesn't seem very useful.
- echoangle 1y ago
- therealpygon 1y agoThis was simply a mad scramble to prove/disprove the claims OpenAI was peddling that the model wasn’t actually performing as well as advertised and that they were lying about the training/compute resources. Open-R1 has since applied the training to a similar 7B model and got similar results. At the end of the day, no one really cares what the data was that it was trained on and most AI providers don’t always share this either when releasing open source models, and certainly not available for closed source models.
- fragmede 1y agoIt's. not. open. source! https://www.downloadableisnotopensource.org/ https://www.downloadableisnotopensource.org/
- quarters 1y agoOk https://huggingface.co/deepseek-ai/DeepSeek-R1-0528/blob/main/LICENSE https://huggingface.co/deepseek-ai/DeepSeek-R1-0528/blob/mai...
- stavros 1y agoSlapping an MIT license on a compiled binary doesn't make it open source.
- quarters 1y agoThey're keeping some stuff to themselves which is fine. I don't expect anyone to have to fully release everything they've got especially considering the vast costs associated with researching and developing these models. What they have released has been distilled into many new models that others have been using for commercial benefit and I appreciate the contributions that they have made.
- alpaca128 1y ago> I don't expect anyone to have to fully release everything they've got I also don't expect Microsoft to release their full Windows 11 source code, but that also means it's not open source. And that's okay, because Microsoft doesn't call it open source.
- behnamoh 1y agoit's got more 'source' than whatever OpenAI provides for their models.
- stavros 1y ago
- JKCalhoun 1y agoIs there a downloadable model? (Not familiar with openrouter and not seeing the model on ollama.)
- aldanor 1y agoOpen weights.