4 ms·
I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs a
by stillpointlab 3mo ago
I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs and model architectures seems significantly more volatile. I wouldn't want to be working out the finer details of my chip rollout only to find a new paper/approach that give multiples of performance.
So I guess it depends on how much the latest-greatest model motivates people, and my read on the current churn is that developers are extremely unloyal to brand at this point and will jump to whoever has the best model. And as long as the best model is running on programmable GPUs, that will be the dominant form.
- gavin_gee 3mo agoseems to me that we are at the asymtote for most usage. sure run the prompts that need the frontier on generic silicon but burning the fable 5 model into silicon could be perfectly viable
- alex43578 3mo agoI think the gamble comes down to how many tokens need to be served on your best model, versus how many can be served in the cheapest/fastest way. Imagine if Anthropic could give effectively unlimited access to Sonnet, for $20. Wouldn’t that be an appealing option for many users? I know I’d make a lot of use of it for agentic tasks, office work, summarization, etc; when right now I’d save quota for more important tasks.
- stillpointlab 3mo agoI mean, if I imagine Anthropic giving away unlimited Sonnet 4.5 away at $20, I would still be paying the $200 for fable. It is a bit like saying "why would you hire someone with a doctorate when you could get unlimited high school grads". How appealing that sounds depends on your needs.
- alex43578 3mo agoIt’s like how most people on here would want a loaded MacBook Pro or RTX 5090, but chromebooks and iGPUs do volume. There’s absolutely a place for a lifestyle subscription to a sonnet model that you could just use everywhere all the time.
- stillpointlab 3mo agoYes, but the context of the discussion is "who wins the war and how". And you are moving the goal posts very far from the original claim, which is that the speed with which companies can get their models onto ASICs will be the determining factor.
- margalabargala 3mo agoRight, and while there are needs that require a doctorate, having unlimited high school grads would be immensely useful for many many tasks. The ability levels of the cheap models are encroaching on the abilities of the frontier models faster than frontier models are expanding their abilities. If we haven't already, we will very soon reach a "good enough" state where having the "best" model matters less and less and less. By analogy, if you buy a new computer, do you get the absolute fastest CPU available? Maybe, depending on your workload. But if you're 90% of the population, you get the cheapest one that has enough power to meet your expected workload, which is mid-range, not top of the line.
- stillpointlab 3mo agoI think there is growing confusion in this discussion. I am saying "I see the fixed nature of ASICs creating a barrier to their adoption despite how cheap they are". Many of the responses seem to say "but there is a market for cheap models". There is a world where we have cheap models and we don't have ASICs powering them. Things like TPUs and NPUs, which are programmable, are likely to fill that role. They are optimized for inference while also allowing different (and updated) models to run on them. Given two companies competing on the cheap end. First company goes TPU, second goes ASIC: who wins? My bet is on TPU since they can update their model, even if their hardware is slightly more expensive and slightly slower, since the optionality of new models beats the performance gap. That may not hold forever but given the pace and volatility of the current LLM market, I believe it will hold for some time.
- margalabargala 3mo agoThe reason you are getting "but there is a market for cheap models" as a response, is people saying in a roundabout way "but soon models may not need to be updated". Models like Kimi 3, GLM 5.2, or even Fable 5 for that matter are reasonable to burn to.ASIC because they are over the threshold of "good enough to be generally useful", something that will continue to be true in the future. Most people do not need the latest model, they need a sufficient model. If I had Fable 5 on an ASIC, I imagine I would use that and ignore paying API rates for Fable 6.
- drob518 3mo agoSure, but you’re undoubtedly the tip of the spear. Lots of people don’t need that.
- vel0city 3mo agoWhat percentage of white collar work these days requires doctorate level thinking all the time?
- akiselev 3mo agoThe tokens per second performance numbers coming from Cerebrus/Talas are several orders of magnitude higher than models running on GPUs, which is such a huge step change that it will enable many more uses of LLMs that are impractical otherwise. I.e. think about gamers and burning in an LLM chip on a game console like a future Play Station - it doesn't matter if its a frontier LLM if it allows them to talk to in game characters in real time and have the LLM drive the storyline. The devs may be limited to training/RLHFing legacy models, but the performance enables many more use cases.
- quotemstr 3mo ago> Cerebrus/Talas They are fast, but they're still programmable accelerators, not a model burned into the gates.
- LarsDu88 3mo agoHonestly, whether you think burning current SOTA to hardware is an overinvestment risk depends on what your definition of intelligence is. If you think intelligence is something that can grow like height such that 18 months from now we will basically be bowing down to machine god giants that are running on B200s, then investing in ASICs is the wrong move. However, if you you subscribe to the (very reasonable view) that intelligence is more like a round ball that we are trying to make as spherical as possible (ala Francois Chollet's writings), then at some point the ball will be smooth enough for most people and many tasks. It takes about 18 months to go through the design, verification, and manufacturing process if you move at breakneck pace. Design could probably be sped up. About 18 months ago the top model was GPT-4o. Not great by today's standards, but still good enough for many tasks (certainly a big chunk of chatbot queries). The current SOTA covers far more use cases, but importantly at a level that surpasses many thresholds of utility.
- JohnBooty 3mo agoAt a 50-100x speedup even a GPT-4o class model could perhaps compete with much newer models simply by thinking deeper, doing harness-controlled Ralph loops, etc. Sure, then it might be "only" ~2-5x faster, but, you wouldn't need to throw all the ASICs into the trash bin. One could also imagine hybrid models, where part of the model is burned into ASICs and part of the model exists in VRAM/HBM2 so it can be updated. I don't have enough low-level knowledge to evaluate the technical or economic feasibility of the above ideas, however.
- tracker1 3mo agoGiving a literal monkey the ability to press more keys faster doesn't get a good joke from it. A model will often come up with worse results given more cycles of compute, only because it will tailspin from second guesses, rethinking and literal flip-flopping on concepts. --- edit: to those following the thread below... if you look at the comments from the account replying, it's pretty obviously a pro-China account and all replies are antagonistic against anything other than a total submission to the Chinese state. My responses are intentionally antagonistic as every point I've brought up is completely ignored in favor of insults, so yeah, I've been insulting back.
- drob518 3mo agoThe churn is an issue. For a while there, it felt like we were getting a new number format every month or two (e.g., fp4, ternary, etc.). That level of innovation works against moving things into hardware, or at least you need to be willing to spin hardware constantly.
- tracker1 3mo agoEven then... Do we now stop the with to asic DeepSeek and do k3 instead? Assuming such work was happening.
- drob518 3mo agoMy personal feeling is to not move to ASICs just yet. Things are still pretty frothy right now, so I would probably wait 6-12 months. At some point the froth always calms down. At that point, commit to ASICs. That said, I’m also not totally sure exactly how much the ASIC hard codes vs having some wiggle room. The Taalas site is a bit vague as to exactly how they encode the model.
- tracker1 3mo agoI would probably wait until a year goes by without any significant advancement, at the very least before making such an effort.
- freeopinion 3mo agoI thought about the volatility, too. But here are some additional thoughts: Many AI uses are not that volatile. I had a 20 minute conversation today with some company's AI phone assistant. It was extremely good and would have been very helpful if any of the dozen people it tried to route me to would have picked up their phone. That AI won't need to be upgraded for a very long time. There is no reason for it to have a cloud brain except to force a recurring revenue for the company selling it. Hardcore gamers are constantly throwing down insane money on the latest hardware. The rest of us can get by for a couple years with whatever we bought when the last one broke. Yeah, it's not the latest, but it gets the job done. I wonder if AI has not already reached the point where a gen 10 CPU--uh, I mean a v3 AI model--will get the job done for the next year. If I really need the up-to-the-second latest abilities for a minute, I can fallback to a cloud brain @ 1M tokens/$. Why pay a monthly lease on a 5-year plan for a 4-door Ford Ranger as your daily commuter? Buy a Clio and rent an F-250 twice a year when you need the hauling/towing capabilities.
- teleforce 3mo ago>I'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. Yesterday there's a news on a breakthrough for probabilistic computer with 1 million p-bits [1]. Since LLM is stochastic in nature, this type of new computer can be much better than ASIC for processing LLM data. [1] Biggest Probabilistic Computer Turns Noise into Answers: https://news.ycombinator.com/item?id=48971938 https://news.ycombinator.com/item?id=48971938
- nightski 3mo agoExcept LLMs are not stochastic in nature. Correct me if I am wrong, but it's just the final layer which outputs a distribution across the output tokens. In reality, the rest of the model is deterministic. This is not like a bayes model or something were it's distributions all the way down.
- gigatexal 3mo agoApparently Google and team will build ASICs to run Gemini even though they have TPUs and it’ll be somewhat dynamic since the weights can be made modular [1] [1] https://finance.yahoo.com/technology/ai/articles/google-plans-chip-run-gemini-135751706.html https://finance.yahoo.com/technology/ai/articles/google-plan...
- alhirzel 3mo agoI think it depends on how good is "good enough". Personally, I think as long as tasks are not coding, there would be a ton of "good enough" value in having your own personal Google on a hardware device. I think the company that designs and licenses the architecture that packages the silicon for the largest market share will be the winner in commercializing this product. Wouldn't be surprised if this is part of Apple's roadmap to retaining existing customer loyalty.