5 ms·
As this is HN, I'm curious if there is anyone here on HN who is interested in starting a business hosting these large open source LLM's? I just finished a test
by morphle 2y ago
As this is HN, I'm curious if there is anyone here on HN who is interested in starting a business hosting these large open source LLM's?
I just finished a test of running this Deepseek-R1 768 GB model locally on a cluster of computers with 800 GB/s memory bandwidth (faster than the machine in the twitter post) and I can now extrapolate to a cluster with 6000 GB/s aggregate memory bandwidth and I'm sure we can reach higher speeds than Groq and Cerebras [1] on these large models.
We might even be cheap enough in OPEX to retrain these models.
Would anyone with cofounder or commercial skills be willing to set up this hosting service with me, it will take less than $30K investment but could be profitable in weeks?
[1] https://hc2024.hotchips.org/assets/program/conference/day2/72_HC2024.Cerebras.Sean.v03.final.pdf https://hc2024.hotchips.org/assets/program/conference/day2/7...
- ForOldHack 2y agoThe Azure(tm) and AWS version of rent-a-second are in the works as we speak. So yes, rent-a-brain/vegetable and no, I will bet you $40k you will not beat either AWS ot Microsoft to the punch. Zero chance of that. They will have their excess computational power with extremely discounted electric rates in place before Friday morning.
- morphle 2y agoI think the important metric will be if we can compete against the price of AWS or Microsoft in running large LLMs, not their time to market. Competing on cost against overpriced hyperscalers is not very hard, and $30K is a small investment, not a gamble. If it would fail, worst case you would only loose $3000-$6500 or so.
- DoingIsLearning 2y ago> $40K is a small investment, not a gamble. If it would fail, worst case you would only lose $3000-$6500 or so. As someone not familiar with investment sourcing or SME financing. Could you break down the maths/accounting? How do you go from sinking 40k in a business to losing 6.5k if you turn the lights off at the end?
- morphle 2y agoYou buy the hardware (48 servers), rent part of a colocation rack with a 10 Gbps or 100 Gbps internet transit link, get a payment processor, make a webpage and GitHub demo with the API. Break down: $3000 labour, $20.5K hardware, $800 monthly rental fees, $376 car fees. When you shut down within a year, the $20.5K popular off the shelf hardware can easily be sold for $17K, a fact you can check from 25 years of data. I would invest more than the initial $30K on optimization after the servers have found paying customers and thus have proven commercial viability. I would invest in software development, finetuning, retraining and above all reverse engineering GPU and neural engine instruction sets and adapting these open source models to the more than 2 quadrillion operations per second that these 48 servers can do.
- seanp2k2 2y agoSo, $0 budget for software dev / sales / support?
- morphle 2y agoI broke down the first $30K investment cost for release of the online API product, that does not need further software development, sales or support. You would be wise to do the software development I mentioned, do more sales and support than was covered under my initial $3000 labour fee. But that you can pay for with the revenues, it would not be the initial investment to see if it is viable as a business.
- gloflo 2y agoThat's what the AI is for, no? /s
- tonyhart7 2y agoWell, when you run an AI company, you must test your product, right? What better way to test it than by building your own webpage, admin panel, etc.?
- lossolo 2y agoWhere can you get 10 or 100 Gbps flat with a full 48U rack and power for $800? Because if that exists, I want to buy them all.
- DrScientist 2y agoI wonder if the real market is actually bringing this stuff inhouse. Given the propensity for these big tech companies to hoover up/steal any information they can gather, running these models locally, with local fine tuning looks quite attractive.
- rrix2 2y ago> Given the propensity for these big tech companies to hoover up/steal any information they can gather at the end of the day you still have to sell this product to the sorts of companies that are far and away all microsoft 365/google workspace clients and we're gonna have to figure that out one day or another
- antupis 2y agoPretty much especially in Europe there is lots of big companies and public sector institutions that would pay serious € if they could run these.
- morphle 2y agoSpot on! I concur most European business and public sector institutions would be eager to rent this because they are not allowed by law to use US datacenters like AWS or Azure.
- kiviuq 2y agoThat's not the only issue. They want a guarantee that the model wasn't trained on copyrighted material.
- TeMPOraL 2y agoNow that is a real feature for now. A lot of hesitation in embracing generative AI in large enterprises stems from uncertainty about copyright issue. Anyone who trained an o1-level model from scratch on public/properly licensed data only would be able to provide a very valuable service to those enterprise customers. However, if both training and operating costs of a DeepSeek-like model are as small as they are, the companies best able to offer this service are... Microsoft, Amazon and Google. And second best are... teams inside the would-be customer enterprises themselves. $6M to train and $6K to run is effectively free for such companies; there is no moat here. The services that enterprise customers would happily buy instead of building are... operations, and assuming legal liability if the model turns out not to be safe from copyright infringement lawsuits. But those are exactly the services those companies are already buying from Microsoft, Amazon and Google.
- PKop 2y agoYep https://azure.microsoft.com/en-us/blog/deepseek-r1-is-now-available-on-azure-ai-foundry-and-github/ https://azure.microsoft.com/en-us/blog/deepseek-r1-is-now-av...
- ryao 2y agoDid you implement token generation for Deepseek R1 using PBLAS?
- rainclouds 2y agoHardware requirements? I think I can hit the memory bandwidth building from parts I have in my house. Maybe even 2x. Asking for fun not profit.
- morphle 2y agoI'd love to visit your house then. You have 768-1400 GB DRAM with 6000 GB/s memory bandwidth? Nice house. In my house I currently have almost 900 GB/S memory bandwidth in aggregate but only 132 GB total DRAM.
- rainclouds 2y agoI’ve got a terabyte of DDR4, and a bunch of old thread rippers. They can take 256 each and 8 channel.
- morphle 2y agoYep, that's the right stuff. Now simply cross-connect all the free PCIe lanes of all thread rippers and you have a nice cluster for LLMs.