6 ms·
DeepSeek R1 Is Now Available on Azure AI Foundry and GitHub
- jgilias 2y agoThat was quick
- deadbabe 2y agoWhy wouldn’t it be?
- erdaniels 2y agoLove that Microsoft is getting behind this actually good model
- discordance 2y agoThey’re selling shovels
- arresin 2y agoThere's definitely a lot to shovel.
- bko 2y agoThis is exciting. Is there a free version of DeepSeek R1 that's completely US based, so we're not sending data to China? I guess you can use this to deploy it, but I'm asking for an application that would be safer to use if you're concerned about Chinese influence.
- breadwinner 2y agoYou can run it yourself: https://workos.com/blog/how-to-run-deepseek-r1-locally https://workos.com/blog/how-to-run-deepseek-r1-locally
- TuxSH 2y agoDistilled R1 models != R1
- roblabla 2y agoYou can run the full R1 (671B variant) locally as well so long as you have the hardware for it. `ollama run deepseek-r1:671b` will do that
- TuxSH 2y agoYeah I mean, most users won't. Sorry if I got on the defensive, saw a bit too many posts on social media claiming you could run the model on your consumer-grade GPU.
- SkyPuncher 2y ago> long as you have the hardware for it. You mean $100k in GPUs?
- roblabla 2y agoThe full model can run on setups worth less than $10K. Here's for instance a $6K build[0]. Granted, that's still expensive, but it is within the realm of something a hobbyist could put together. [0]: https://x.com/carrigmat/status/1884244369907278106 https://x.com/carrigmat/status/1884244369907278106
- mlboss 2y agohttps://www.together.ai/pricing https://www.together.ai/pricing
- KaoruAoiShiho 2y agoHow's the pricing as compared to official API or openrouter providers?
- xnx 2y agoThis is bizarre to see the entire AI hype cycle speedrun all over on just DeepSeek. I'm trying to square the excitement over DeepSeek with its good -but not dominant- performance in evals.
- zamadatix 2y agoPreviously choosing a top tier AI model tied you to what that provider wanted to do with hosting the model long term and the pricing they wanted to charge for it. Now you can get the same model anywhere with GPU, hosted or not, for minimal cost overhead to what it takes to run the model itself. You're also free to tune, retrain, or otherwise mess with the model as you see fit without needing approval. The excitement is probably a bit much but it's not just about the eval results themselves but the baggaged attached with them.
- samvher 2y agoFor me the excitement is that around the o3 announcement I had a feeling like we were heading to an OpenAI / Sam Altman controlled dystopia. This resets that - you can run the model yourself, you can modify it yourself, it's essentially on par with the best public models, and it gives hope that the smaller players have a fighting chance going forward. They also published their innovations bringing back some of the feeling of open science that used to be in ML research but which mostly went away.
- xnx 2y agoGoogle models are already in the lead in many areas in capability and cost, so I never felt like OpenAI was dominant. OpenAI was first to make a splash, but ChatGPT is in a ~5 way tie in terms of what it can do.
- maxglute 2y agoIMO more anti hype for openAI who might be dominant, but are they $3500 per task (O3 high) dominant, or $200 per month dominant.
- xnx 2y ago
- arnado 2y agoI don't understand the hype because I'm out of the loop. Is the only advantage the lower hardware requirements, thus cost? Is there something I'm missing?
- Synaesthesia 2y agoYeah it's a lot more efficient, it's also a very advanced model that answers questions in a multi-step way, like OpenAI-O1, it performs extremely well.
- bl4kers 2y agoIt's also open source
- vinni2 2y agoopen weights you mean. People confuse open weights with open source.
- tmasdev 2y agoOpenAI o1 and Deepseek r1 have similar performance (o1 is a bit better at reasoning though you can see r1’s though process which you could argue trumps the competition). OpenAI o1 api cost: $60/million output tokens. Deepseek r1 api cost: $2.19/million output tokens.
- ofou 2y agocompetition is beneficial for all of us, this is great
- kam1kazer 2y agoLooks like DeepSeek R1 is a Microsoft shady move against Sam ;]
- bogdan 2y agoWhat is shady about it? Should a company the size of Microsoft stand by instead?
- BoorishBears 2y agoIt's a bit shady after actually trying to use it. - Single digit TPS on rare chance it responds, and frequent complete hangs (1 out of maybe 20 requests even complete) - 4k input token cap (vs native 128k context window) - No pricing - Unstated rate limits It genuinely seems like they spun up a single H100 cluster to enable the headline of this post and help form a narrative then left it at that. Definitely not meant to genuinely provide access to R1 in any serious way.
- jimpster 2y agoWhat is the price on Azure?
- jimpster 2y agoWhat is the price?
- rcarmo 2y agoFree in preview, at least in my personal subscription (there was a disclaimer saying that it was in preview, no guarantees of response times, etc.) (I work at Microsoft but am not on the clock as I write this, and keep my personal projects separate)
- CodeCompost 2y agoIt can only be deployed in FIRA.
- vemonet 2y agoIt's not really available for real use. I have tried it on "Azure AI foundry" through their serverless API with a paid subscription. It takes 80s to answer a basic question that was answered in 7s by OpenAI gpt-4o. And there was not that much thought process, it was just super slow to output each token. I guess this slowness is explained by the pricing, they are still figuring out how to run the inference for this model: > DeepSeek R1 use is currently priced at $0, and use is subject to rate limits which may change at any time. Pricing may change, and your continued use will be subject to the new price. The model is in preview; a new deployment may be required for continued use. There is also a hard limitation of 4k tokens as input context (context window on DeepSeek model is 120k tokens), which prevents using it for RAG use-cases: > Message: Request body too large for deepseek-r1 model. Max size: 4000 tokens. Also the documentation and python type hints of their inference lib have a lot of straight up errors in it (they are confusing the class attributes `model=` and `model_name=` at many places in the docs, spoiler: the good one to use is `model_name=`, even if the type hints recommend to use `model=`). I have also tried with more stable models like Mistral Large, but the streaming feeling is really bad, they are sending whole sentences at a time, with multiple seconds of wait between each sentence. Does not feel smooth at all compared to any other provider out there. Would not recommend Azure AI foundry for production use (or any use to be honest). Does not worth the pain to navigate the documentation. We will be using directly DeepSeek API, or fireworks.ai, or together.ai.
- cbold 2y agoThis is also interesting: "Customers will be able to use distilled flavors of the DeepSeek R1 model to run locally on their Copilot+ PCs." This is news, because Microsoft seems happy to not be tied to OpenAI so heavily. This could also safe a huge amount of money for their Office 365 Copilot initiative. I figure Microsoft started analyzing these models ASAP in their labs to catch up with OpenAI, Google, Anthropic etc. By also hosting this model, they will help normalize the use of them from which they immensely benefit.
- hwertz 2y agoWell, just running on a 6C/12T Coffee Lake CPU, (I'm looking through these speeds in LM Studio as I type this..) I got like 2 tokens a second with Deepseek R1 14B, 3.4 with 7B Qwen, and 4.4 with 8B Llama, although out of those two I found 7B Qwen's answer to be a bit better. (My GTX1650 has 4GB VRAM, loading 1/4 the layers is pretty ineffective, GPU util went up to 10% and I gained like 1 token a second LOL.) So it'd take a minute or two to type out one of those answers where it's got about 4 or 5 beefy paragraphs of thought and a decent sized paragraph for it's answer. I'll put it this way, I can type 120 WPM and it puts out text a bit faster than I could write it. Input's a LOT faster though, I was asking these models to analyze a document so my input was like 2200 tokens, they all did well over 100 tokens a second on input.