7 ms·
The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASI
by LarsDu88 3mo ago
The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest.
The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release.
Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software engineering (if not already). Does anyone think we need a Mythos level model to plan a road trip, or give someone tips on making a cake recipe?
A Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand.
Furthermore, if you're an enterprise the risk of data exfiltration and feeding data to a potential competitor like OpenAI or Anthropic is greatly reduced if you could shift to on-prem ASIC deployments. A handful of chips could cover a wide variety of use cases and cover them more securely. There are a lot of corporate use-cases for LLMs that are not frontier math research or coding.
- misiti3780 3mo agois anyone doing this ?
- LarsDu88 3mo agoThe closest example I've seen is ChatJimmy: https://chatjimmy.ai/ https://chatjimmy.ai/ a prototype from Taalas running Llama 8B Scaling this up to 2.8 Trillion (350X increase), will certainly be challenging. If I was younger and had the right background, I'd love to dive into attempting somethign like this
- selectodude 3mo agoAlso a 3-bit quant. Useable but a long long way away from useful.
- carterschonwald 3mo agonot sure about that, but im actively working on designing ultra sparse models that i want to have perform competitively with stuff 100-10_000 times larger. ehich does yield similar throughput. time will tell id it works out
- eckr 3mo agoThere was a startup that did this for Llama 3, I forgot their name. Etched is also doing some similar things I believe.
- robgough 3mo agoobligatory link to https://chatjimmy.ai https://chatjimmy.ai
- hnfong 3mo agoIt sounds vaguely similar to what Cerebras.ai is doing
- gopalv 3mo agoTaalas HC1 is the closest thing to this. Last I saw they posted Deepseek R1 numbers in Feb of this year. The challenge is rolling out a new one every 7-8 weeks as the weights change & cheap enough for a hyper scaler to afford to buy one and save enough on power over the next 8 weeks as a payoff.
- 535188B17C93743 3mo agoYeah, I've heard of Taalas doing it. Not sure of others but I'm sure lots of companies are considering it, especially as we start to hit points of depreciating returns in training.
- alach11 3mo agoNews just broke today that Google is planning on doing this: https://news.ycombinator.com/item?id=48986351 https://news.ycombinator.com/item?id=48986351
- stillpointlab 3mo agoI'm not convinced, mostly because things like crypto, which I believe went into ASICs, were based on very slowly moving and mostly understood algorithms. LLMs and model architectures seems significantly more volatile. I wouldn't want to be working out the finer details of my chip rollout only to find a new paper/approach that give multiples of performance. So I guess it depends on how much the latest-greatest model motivates people, and my read on the current churn is that developers are extremely unloyal to brand at this point and will jump to whoever has the best model. And as long as the best model is running on programmable GPUs, that will be the dominant form.
- gavin_gee 3mo agoseems to me that we are at the asymtote for most usage. sure run the prompts that need the frontier on generic silicon but burning the fable 5 model into silicon could be perfectly viable
- alex43578 3mo agoI think the gamble comes down to how many tokens need to be served on your best model, versus how many can be served in the cheapest/fastest way. Imagine if Anthropic could give effectively unlimited access to Sonnet, for $20. Wouldn’t that be an appealing option for many users? I know I’d make a lot of use of it for agentic tasks, office work, summarization, etc; when right now I’d save quota for more important tasks.
- stillpointlab 3mo agoI mean, if I imagine Anthropic giving away unlimited Sonnet 4.5 away at $20, I would still be paying the $200 for fable. It is a bit like saying "why would you hire someone with a doctorate when you could get unlimited high school grads". How appealing that sounds depends on your needs.
- alex43578 3mo agoIt’s like how most people on here would want a loaded MacBook Pro or RTX 5090, but chromebooks and iGPUs do volume. There’s absolutely a place for a lifestyle subscription to a sonnet model that you could just use everywhere all the time.
- umeshunni 3mo ago> the winner will be whoever burns their models to ASICs fastest. I don't understand how this works when the models are evolving so fast that your burned ASIC is outdated (or at least not top of the line) in a few weeks.
- tempoponet 3mo agoIf llama 3 70b were available for $400 today, people would make it work. They would have 4 of them working side by side or 1 of them working with a GPU on another model. To me, it's like imagining if Sonnet 3 was burned into an ASIC 8 years ago and then never changed. It would still be revolutionary, and today we would have an entire ecosystem of tools and services built around it, likely surpassing some of our current workflows. The frontier is a different beast, but it would likely mean competing on price.
- treis 3mo agoI think you'd be better off getting 1% of a $40,000 card. Or maybe 10% of a $4,000 card is more realistic. I think that's the problem at the moment. A much better LLM that doesn't use my battery is 20ms and <10mb of data away.
- bensyverson 3mo agoI think it makes more sense if you think about an ASIC pipeline which is updated on an annual basis, like Apple's M or A-series chips. Sure, an ASIC model will always be behind (12 months? 18 months?), and a hosted flagship will be significantly better. But as time marches on, would I use a flagship model from last year if it came on a PCI card and cost $1000? Without a shadow of a doubt.
- efficax 3mo agofor many purposes they're good enough now. If I had an opus 4.8 class model on a box next to me that could produce tokens at rates like 5000/s, i don't know if i'd need a new one for a long time. I think we might be underestimating how powerful very very fast LLMs could be, since they could iterate on tons of small variations on tasks. paired with deterministic guardrails that gate "doneness", you could loop on tasks for a long time having the agent try different strategies until the goal was reached, in ways that are just impractical now (and very expensive)
- mountainriver 3mo agoSol is running on Cerebras right?
- LarsDu88 3mo agoCerebras is not an ASIC. It is a large wafer scale chip that has a load of small SRAM modules paired with tiny compute models. The SRAM is basically a giant cache that is supposed to eliminate the bottleneck between shuttling and materializing big tensors between DRAM and cache, but the actual amount of SRAM is still nowhere near enough to server a frontier model on a single wafer. You still need dozens of wafers, which makes Cerebras rather cost intensive given that Nvidia can get dozens of GPUs out of a single wafer
- mountainriver 3mo agoI know that but they also said Cerebras
- gopalv 3mo ago> Does anyone think we need a Mythos level model to plan a road trip, or give someone tips on making a cake recipe? This is starting to look at a lot like Intel vs Arm from the last era. The Fable & Mythos are starting to look like a giant Xeon, while the smaller lighter models are starting to look like a lot of tiny ARM chips which sip on power instead. The risk is the same as what Intel had. There is a group who are pushing them to go bigger and with a resource no limit approach, who have a lot of dollars to push you that way. Follow them and they lead you to a pile of money, but then you risk something like Apple Silicon happening. Something which got better because of efficiency & continuous improvement, not neutered due to it.
- alex43578 3mo agoExcept aren’t these Chinese models having to go big too? While they might not be Mythos/Fable sized, a trillion plus parameters is hardly a local server or even desktop to mainframe style jump
- hadlock 3mo agoI've had 2b models give a plausible Paris vacation itinerary. A tools-capable 12b and especially 30b model from 2026 is certainly capable of producing passable results. I was demonstrating the qwen 3.6 27b model I stood up last week to my wife and it gave her a passable Moroccan Chicken recipe. With tool calling (search) they're quite good.
- applicative 3mo agoIt’s AI talking about AI so cum grano salis, but my AI is saying I would need at least a half million dollar in hardware to run the newer high quality Chinese models with bemchmark-competitive force. You can run a Moroccan chicken fragment on the cheap in homage to what you cannot run
- alex43578 3mo agoWith OpenAI having released 20 and 120B models a while back, I think they recognized that tiny models were never going to be a defensible income stream. Any value will come from the largest models, and those largest models are unlikely to ever run on consumer hardware within their window of relevancy.
- doctoboggan 3mo agoI think the SotA is moving too fast for the production timelines of an ASIC, wouldn't you think? People are just now coming out with LLAMA ASICS but who would want to use LLAMA? Or I guess you are arguing that the models _now_ will be durably useful enough to commit the time to creating the ASIC?
- JohnBooty 3mo agoguess you are arguing that the models _now_ will be durably useful enough to commit the time to creating the ASIC? Objectively, the current frontier-ish models will be useful for some time. Imagine the zombie apocolypse hits, recedes, and you need to rebuild society. An offline copy of Fable or even Opus would be a nice thing to have. Subjectively, it's hard to say if people will pay for "a model from 18 months ago, but REALLY FAST AND CHEAP" The speed difference suggests some use cases that might narrow the performance gap. With a > 50x performance delta you have some headroom to play with. You can do many many fast iterations of ye olde "Ralph loops." You can also jack up the reasoning/effort level. And you could probably do some combination of both, while still running really fast ie 10x the number of iterations at 2x reasoning/effort. So I think a hypothetical "50x faster Opus 4.8, but burned into ASICs" could be pretty competitive against the frontier models from 2027, 2028, 2029, and maybe beyond?
- xtracto 3mo agoI dream of some kind of "adjustable" ASIC layers, that have the most "computation demanding" layers as ASICs plus a bunch of configurable R+W circuits (FPGAs?) that can be written to upgrade the models to some extent. What's fascinating is that we are pushing the state of the art of hardware at this point.
- searealist 3mo agoHow does an ASIC manage to have 60x the memory bandwidth needed to achieve that speedup?
- LarsDu88 3mo agoCurrently the setup is paged view in RAM shuttled to HBDRAM (VRAM) on the GPU, which in turn has to get materialized piece by piece onto cache SRAM on the GPU. Cerebras tries to get around this by keeping everything on cache SRAM as much as possible, which it burns directly to the chip wafer itself and physically places that SRAM directly next to the tiny compute unit that does the actual math. An ideal setup (not sure how easy this is to achieve in practice), is the burn the weights of the model directly to the chip as a sort of ROM, the actual math operations as actual digital circuits, and have SRAM, or even something akin to naked registers to directly compute off inference batch data. Cuts out 2-3 layers of abstraction and indirection.
- oblio 3mo agoI assume ROM would be even more expensive than SRAM?
- marcyb5st 3mo agoYou don't need memory because the output of the previous layer flows directly into the input gates of the next. So no I/O to/from HBM or similar You still need some memory for the context, in flight answers, ... but not for the model weights and for the output of the intermediate layers. I found taalas demo here: https://chatjimmy.ai/ https://chatjimmy.ai/
- throw1234567891 3mo agohttps://taalas.com/ https://taalas.com/
- wongarsu 3mo agoWhile running the model at 9000 tokens/s is the more flashy demo, I imagine running 1000 concurrent requests at 150 tokens/s each is the much more achievable goal
- OliverGuy 3mo ago[dead]
- wyre 3mo agoIt depends. If I was running a model locally I would much prefer 9000 t/s. If I was running an inference company, obv 1000 concurrent requests at 150 t/s is preferable.
- LarsDu88 3mo agoAt 9000 tokens/s you could interleave a lot of requests so long as pre-fill is also fast. It really depends on how much you need to keep sessions open to take advantage of KV caching
- quotemstr 3mo agoWhat precisely is an ASIC supposed to do that a more programmable accelerator can't? Memory latency is memory latency: doesn't matter whether it's embedded in a cache-line wait state or some flipflop state machine. ROM isn't going to be faster than RAM either. Likewise, for compute, is the ASIC somehow going to beat a systolic array? You can't have one circuit per weight: the die area and electrical fan-out would be insane. I'm not seeing how an ASIC specialized for a specific model would actually help much. I mean, sure, we can build more specific accelerators, e.g. for softmax, but these work fine in the context of a programmable pipeline. Yes, there are more exotic things out there, like optical matrix multiplication systems. Those are different. But above, aren't you talking about just doing conventional digital linear algebra, but with a model-specific set of circuits?
- marcyb5st 3mo agoNo, if you burn your model into the silicon you don't need memory as the output of a layer flows through circuitry directly into the next one. No I/O to memory of any kind. You still need a bit of memory for in flight answers, but that's it.
- quotemstr 3mo agoYou can pipeline in software too, and even with a dedicated circuit, you need to put weights somewhere, because you're not going to have a special partially-applied FP8 FMA unit for each of your trillion model weights. I can accept the idea of specializing a circuit for a specific model shape, but I'm not seeing a need to specialize a circuit for the weights inside the shape.
- marcyb5st 3mo agoApparently that is what Taalas did though. Not an hardware person, so take this with a pinch of salt
- BurnerOptical 3mo ago[dead]
- applicative 3mo agoOpen weights are a red herring if you can’t run the model independently. Good luck fitting these models on your Mac mini. In fact the superior models are irreducibly nothing but superior web services run from China.
- Slartie 3mo agoDid you not notice that there are inference providers running all the various Chinese open weight models in the US and EU? Nobody needs "web services run from China" to use Chinese open weight models.
- hughw 3mo agoA Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand. 640K ought to be enough for anybody.
- JacobAsmuth 3mo agoAgreed 100%. This guy thinks there's a limit on the demand for intelligence. You think that Fable 7 which can run a billion dollar corporation on its own has no consumer demand just because we have fable 5 at 9k tok/s? Who do you think will be the biggest customer of such a model? Fable 7, obviously.
- LarsDu88 3mo agoCertainly there will be demand for Fable7, but that demand is context specific. Frontier labs' profit is dependent on there being sufficient demand for the next layer of capability and whether the premium consumers are willing to pay for that. The incremental unlock of capability by ever increasing frontier model sizes will eventually reach diminishing returns. I would argue tnference speed increases would actually unlock a different kind of more meaningful value for a wider audience.
- drob518 3mo agoOf course not. But many tasks won’t require Fable 7 level intelligence and many people won’t want to pay for it. Honestly, I’m using Deepseek v4 Flash a LOT lately to do more mundane tasks because it’s so nearly free and I don’t need Fable or even Opus. Serving those mid-level models at high speeds and low prices is a definite winner for lots of applications. And sure, the frontier models will continue to drive the frontier forward.
- sdfefcxv 3mo agotheoretically theres a no limit on the demand of anything if the price is right pretty stupid statement lmao
- 3mo ago
- pstuart 3mo ago[dead]
- aaroninsf 3mo agoApple's M7 comes to mind...
- faitswulff 3mo agoAnother reason this may be true, selling hardware may be the only durable moat after a while
- mullingitover 3mo ago> the winner will be whoever burns their models to ASICs fastest. It'll obviously be China, and they won't need the bleeding edge of lithography tech to make it happen. Every Chinese smartphone will have something like Sonnet 5, along your car (well, not those of us in the US, but we'll look longingly at pictures of them while we drive whatever the government decides we're allowed to drive in Fortress America). Give it ten years and your smart litterbox from Temu will be running its own local model.
- adventured 3mo agoChina won in desktop PCs? Nope. China won in cloud services? Nope, not even close. They had to clone AWS just to try to keep up. China won in mobile? Nope. Although they're very competitive there. China won in search? Nope. Baidu who? China won in ecommerce? Nope. Their dominance is overwhelmingly domestic. China won in software? Nope. Windows is US. Android is US. iOS is US. MacOS is US. Linux is US/Europe/Global. Look at the top 50 largest software companies. China won in silicon? Nope. Look at the top 50 companies. China won in .... India and Latin America can manufacture iPhones now. But sure, China's the obvious winner this time. Good luck.
- mullingitover 3mo agoIndia and Latin America doing some final assembly is kind of a whimper of a counterpoint. Most of this litany is trivially refuted by pointing out that roughly 40% of the world's declared 5G Standard Essential Patents were developed in China. The only reason they don't completely dominate telecom globally is the panicked invocation of national security fears to erect emergency trade barriers. "China won in ecommerce? Nope. Their dominance is overwhelmingly domestic." Their domestic ecommerce market eclipses the US and the next several biggest markets combined. You can argue protectionism, but then you have to face the reality that the US is doing its own flailing protectionism today. Generally, this is an embarrassing category error of an argument. Running models on ASICs, IoT devices, and EVs are infrastructure and hardware plays: areas where China is inarguably so good that it's a problem.
- deleted 3mo ago[deleted]
- seydor 3mo agoThey require a lot of computation regardless of whether this will take place on ASICs. China can do it, even if the chips are not the most power-efficient. They have built lots of electricity capacity. Ultimately however, that just means that models will become dirt-cheap. The money will be made with applications built on top of the models.
- chermi 3mo agoMy gut is that, at least for now, there's a timescale mis-match problem. Hardware still takes too long, then you have to deploy it. I don't know much about "burning asics", but if the whole process of spinning up programmable GPU data centers is months, I imagine the whole ASIC cycle has some catching up to do. To be clear, by timescale mismatch I mean model quality improvement timescale vs. deployment timescale. But maybe we're finally getting to the point where behind-the-frontier-but-cheaper are in sufficient demand and ASICs make it cheap enough to close that gap.
- zdragnar 3mo agoIf Qwen 3 Coder were available as a PCI card I could just slot into my desktop, and the software on the system could recognize the card and still work with other cards should I decide I later want to upgrade to Qwen 5 or whatever, and at a three figure price point, I imagine it would do quite well. The comparison I've seen elsewhere is the old console systems with separate cartridges for games... I wouldn't want to be regularly swapping them, but if they came out in a form factor that didn't require me to shell out multiple-4 digit figures in upgrades just to use the next model, it'd easily be worth it for me. I've already got a home lab, and it's specced to last, minus the GPU. I picked up a separate system for local llm experimentation, but I'm not likely to be upgrading it again. The value add is incredibly small compared to the cost. The real benefits are data privacy and never worrying about rate limits, and there's a price point beyond which an incremental improvement to the model doesn't justify upgrading the system.
- chermi 3mo agoWow having it as a PCI card would be awesome. I don't know the feasiblity but it would be really cool. Ordering different models, back to physical media basically...
- snarfy 3mo agoI'm convinced there is a Transmeta/transpiler solution somewhere between 9000 t/s ASIC and Nvidia GPUs. LLMs fit the model much better than what they were trying to do with general purpose in the 2000s.
- sarchertech 3mo ago> Does anyone think we need a Mythos level model to plan a road trip Yes. The latest OpenAI and Anthropic models are terrible at planning roadtrips. This is something I try to use them for frequently. They constantly get things completely wrong. I’d say that about half of the stops they suggest fail to follow whatever filters I’ve asked for.
- jrs100000 3mo agoFable is a pretty crappy general purpose LLM. It reminds me a little bit of GPT 5.2, in that it sometimes gets argumentative and deceptive after being caught in a mistake, and it tends to make a lot of them when you ask for judgement calls. Opus 4.8 is better for that sort of thing, or Opus 4.6 if it's important that it actually follow all of your instructions. And this is one of the big things that seems to be missed in these discussions: There is no longer a universal linear trend of LLMs being 'better' each iteration. They are becoming more specialized, and ones that approach problems from a different angle (like Fable/Mythos) can appear breakthrough when first released, but we don't appear to be on a path that actually leads to general purpose hyper intelligence.
- daadx 3mo agoThere never was a path. It was nonsense to draw in investment and justify an inflated valuation. lmao
- paxys 3mo agoThe field is moving too fast for it to make sense. ASICs have a 12-24 month development cycle and you’d have to throw them in the trash 4 months later.
- jackb4040 3mo agoIf this approach doesn't include a way to apply updated finetunes on a daily basis (also without going back to VRAM being the bottleneck) then the trash interval would be measured in hours, not months
- dboreham 3mo agoGPUs are already ASICs.
- knicholes 3mo agoWe're getting new models every three months. They'd have to work quickly to get those chips made and installed.
- jv22222 3mo agoYeah, I have been thinking exactly the same. Makes all the sense in the world and is inevitable as far as I can see. https://news.ycombinator.com/item?id=48464958 https://news.ycombinator.com/item?id=48464958
- throwawayffffas 3mo ago> whoever burns their models to ASICs fastest. There is already custom hardware see cerebras. GPUs have a lot of slack there is at least one lab that had a (small 8b) model generate almost 3000 tokens per second on a MI300X for a talk, instead of the typical software stack that did maybe 100ish tokens per second. High bandwidth flash storage is in the works, i.e hard drives with TBs of storage and over 1 TB per second of read speeds. Meaning that in a couple of years you may be able to buy a card with 40-90GBs of HBM and 4TB of HBF and run a 3T model locally at a reasonable speed for 10-20k as opposed to a cool mil.
- topspin 3mo ago> Meaning that in a couple of years you may There is no "may" here. You will see this. It's always difficult to see it from the present, but we're not at some end stage in hardware development; we're still on the same curve our predecessors also couldn't see: they couldn't imagine that there would be high performance computers carried in our pockets, with staggering amounts of storage and compute, putting to shame the machines they filled rooms with.
- throwawayffffas 3mo agoOh yeah the may is on the 2 year time horizon. It could be 3 or 4. Or next year.
- topspin 3mo agoYes, the timelines are highly uncertain, but the AI boom and AI money are reinvigorating the entire fabrication ecosystem from top to bottom and causing innovation where there had only been incremental progress for years. People aren't just thinking about their node schedule and yields: they're pursuing radical new designs with far greater urgency. That's where OpenAI in a box is going to come from. And it won't take long.
- signa11 3mo agowhy not start with fogas ? you know … to test the waters ?
- alach11 3mo ago> the winner will be whoever burns their models to ASICs fastest As of today, that appears to be Google! https://news.ycombinator.com/item?id=48986351 https://news.ycombinator.com/item?id=48986351
- Havoc 3mo agoFun as the asic play would be it has a giant hole - context storage. Raw speed only gets you so far if you can’t store and cache