10 ms·
Furiosa: 3.5x efficiency over H100s
- grosswait 9mo agoHow usable is this in practice for the average non AI organization? Are you locked into a niche ecosystem that limits the options of what models you can serve?
- sanxiyn 9mo agoYes, but in principle it isn't that different from running on Trainium or Inferentia (it's a matter of degree), and plenty of non-AI organizations adopted Trainium/Inferentia.
- darknoon 9mo agoreally weird graph where they're comparing to 3x H100 PCI-E which is a config I don't think anyone is using. they're trying to compare at iso-power? I just want to see their box vs a box of 8 h100s b/c that's what people would buy instead, and they can divide tokens and watts if that's the pitch.
- minimaltom 9mo agoWhats a more realistic config?
- _zoltan_ 9mo ago8xGPUs per box. this has been the data center standard for the last 8ish years. furthermore usually NVLink connected within the box (SXM instead of PCIe cards, although the physical data link is still PCIe.) this is important because the daughter board provides PCIe switches which usually connect NVMe drives, NICs and GPUs together such that within that subcomplex there isn't any PCIe oversubscription. since last year for a lot of providers the standard is the GB200 I'd argue.
- minimaltom 9mo agoFascinating! So each GPU is partnered with disk and NICs such that theres no oversubscription for bandwidth within its 'slice'? (idk what the word is) And each of these 8 slices wire up to NVLink back to the host? Feels like theres some amount of (software) orchestration for making data sit on the right drives or traverse the right NICs, guess I never really thought about the complexity of this kind of scale. I googled GB200, its cool that Nvidia sells you a unit rather than expecting you to DIY PC yourself.
- _zoltan_ 9mo agousually it's 2-2-2 (2 GPUs, 2 NICs and 2 NVMe drivers on a PCIe complex). no NVLink here, this is just PCIe - under this PCIe switch chip there is full bandwidth, above it's usually limited BW. so for example going GPU-to-GPU over PCIe will walk GPU -> PCIe switch -> PCIe switch (most likely the CPU, with limited bw) -> PCIe switch -> GPU NVLink comes into the picture as a separate, 2nd link between the GPUs: if you need to do GPU-to-GPU, you can use NVLink. you never needed to DIY your stuff, at least not for the last 10 years: most hardware vendors (Supermicro, Dell, ...) will sell you a complete system with 8 GPUs. what's nice on GH200/GBx00/VR systems, is that you can use chip-to-chip NVLink between the CPU and GPU, so the CPU can access GPU memory coherently and vica versa.
- ac29 9mo ago> they're trying to compare at iso-power? Yeah they are defining a "rack" as 15kW, though 3x H100 PCIe is only a bit over 1kW. So they are assuming GPUs are <10% of rack power usage which sounds suspiciously low.
- bradfa 9mo agoIt would also depend on the purchase cost and cooling infrastructure cost. If this costs what a 3x H100 box costs then it’s a fair comparison even if not a direct comparison to what customers currently buy.
- zmmmmm 9mo agoWhat can it actually run? The fact their benchmark plot refers to Llama 3.1 8b signals to me that it's hand implemented for that model and likely can't run newer / larger models. Why else would you benchmark such an outdated model? Show me a benchmark for gpt-oss-120b or something similar to that.
- sanxiyn 9mo agoLooking at their blog, they in fact ran gpt-oss-120b: https://furiosa.ai/blog/serving-gpt-oss-120b-at-5-8-ms-tpot-with-two-rngd-cards-compiler-optimizations-in-practice https://furiosa.ai/blog/serving-gpt-oss-120b-at-5-8-ms-tpot-... I think Llama 3 focus mostly reflects demand. It may be hard to believe, but many people aren't even aware gpt-oss exists.
- reactordev 9mo agoMany are aware, just can’t offload it onto their hardware. The 8B models are easier to run on an RTX to compare it to local inference. What llama does on an RTX 5080 at 40t/s, Furiosa should do at 40,000t/s or whatever… it’s an easy way to have a flat comparison across all the different hardware llama.cpp runs on.
- zmmmmm 9mo agoNow I'm interested ... It still kind of makes the point that you are stuck with a very limited range of models that they are hand implementing. But at least it's a model I would actually use. Give me that in a box I can put in a standard data center with normal power supply and I'm definitely interested. But I want to know the cost :-)
- nl 9mo ago> we demonstrated running gpt-oss-120b on two RNGD chips [snip] at 5.8 ms per output token That's 86 token/second/chip By comparison, a H100 will do 2390 token/second/GPU Am I comparing the wrong things somehow? [1] https://inferencemax.semianalysis.com/ https://inferencemax.semianalysis.com/
- 9mo ago
- kuil009 9mo agoThe positioning makes sense, but I’m still somewhat skeptical. Targeting power, cooling, and TCO limits for inference is real, especially in air-cooled data centers. But the benchmarks shown are narrow, and it’s unclear how well this generalizes across models and mixed production workloads. GPUs are inefficient here, but their flexibility still matters.
- whimsicalism 9mo agoGot excited, then I saw it was for inference. yawns Seems like it would obviously be in TSMCs interest to give preferential taping to nvidia competitors, they benefit from having a less consolidated customer base bidding up their prices.
- ttul 9mo agoMy best guess after dipping my toe into semiconductor fabrication a decade ago is that there is a mysterious guru in a cave under a volcano who decides which customers get access to which nodes at which prices.
- sognetic 9mo agoEverything is currently pointing towards inference being the main cost driver for LLMs in the future. Test-time-compute requires huge amounts of tokens in inference and makes providing frontier models as services unprofitable. Anyone not under some kind of export restrictions can scrounge together some GPUs to train a frontier model (hell, even DeepSeek which is under these restrictions could) but providing a service that can compete with OpenAI et al. will prove to be quite costly. 3x improvements in inference are therefore nothing to sneeze at IMO.
- roughly 9mo agoI am of the opinion that Nvidia's hit the wall with their current architecture in the same way that Intel has historically with its various architectures - their current generation's power and cooling requirements are requiring the construction of entirely new datacenters with different architectures, which is going to blow out the economics on inference (GPU + datacenter + power plant + nuclear fusion research division + lobbying for datacenter land + water rights + ...). The story with Intel around these times was usually that AMD or Cyrix or ARM or Apple or someone else would come around with a new architecture that was a clear generation jump past Intel's, and most importantly seemed to break the thermal and power ceilings of the Intel generation (at which point Intel typically fired their chip design group, hired everyone from AMD or whoever, and came out with Core or whatever). Nvidia effectively has no competition, or hasn't had any - nobody's actually broken the CUDA moat, so neither Intel nor AMD nor anyone else is really competing for the datacenter space, so they haven't faced any actual competitive pressure against things like power draws in the multi-kilowatt range for the Blackwells. The reason this matters is that LLMs are incredibly nifty often useful tools that are not AGI and also seem to be hitting a scaling wall, and the only way to make the economics of, eg, a Blackwell-powered datacenter make sense is to assume that the entire economy is going to be running on it, as opposed to some useful tools and some improved interfaces. Otherwise, the investment numbers just don't make sense - the gap between what we see on the ground of how LLMs are used and the real but limited value add they can provide and the actual full cost of providing that service with a brand new single-purpose "AI datacenter" is just too great. So this is a press release, but any time I see something that looks like an actual new hardware architecture for inference, and especially one that doesn't require building a new building or solving nuclear fusion, I'll take it as a good sign. I like LLMs, I've gotten a lot of value out of them, but nothing about the industry's finances add up right now.
- richwater 9mo agoThis is from September 2025, what's new?
- sanxiyn 9mo agoWhat's new is HN discovered it. It wasn't posted in September 2025.
- tedk-42 9mo ago100% People forget this is also a place of discussion and the comment section is usually peak value as opposed to the article itself.
- nine_k 9mo ago> We are taking inquiries and orders for January 2026. Hence the relevance, maybe.
- nl 9mo agoSo inference only and slower than B200s? Maybe they are cheap.
- jszymborski 9mo agoIs it reasonable for me not to be able to read a single word of a text-based blog post because I don't have WebGL enabled?
- pas 9mo agoyou are not the target audience whatever runs on typical investor/C-suite laptops and phones (so new iPhone/MacBook with "stock" Safari, maybe in corporate some cursed Windows setup with Chrome) is okay, and obviously they need to maxx out the glitter, it's the 2020s
- throwaway290 9mo agoI know people with iphones 17 pro who do not have webgl enabled for sanitary infosec reasons:) probably they don't want this site to be scraped by LLMs which would be kinda ironic
- peterarends 9mo agoA fix for me in FF was toggling 'reader view'. They might be reasonable and it could be a bug.
- nycdatasci 9mo agoIs this from 2024? It mentions "With global data center demand at 60 GW in 2024" Also, there is no mention of the latest-gen NVDA chips: 5 RNGD servers generate tokens at 3.5x the rate of a single H100 SXM at 15 kW. This is reduced to 1.5x if you instead use 3 H100 PCIe servers as the benchmark.
- vfclists 9mo agoWhy is their website demanding WebGL?
- LTL_FTC 9mo agoThe server seems cool but the networking seems insufficient for data centers.
- kalmyk 9mo agothat's a nice rack
- galaxyLogic 9mo agoHow is this possible? Doing AI with "dual AMD EPYC processors". I thought you needed to have GPUs or something like that to do the matrix multiplications needed to train LLMs? Is that conventional wisdom wrong?
- ilsubyeega 9mo agoit uses own chip under the hood, see accelerator mentioned in spec.
- KronisLV 9mo agoI think it's actually really cool to focus on efficiency over just raw performance! The page for the cards themselves goes into more detail and has a pretty nice graph: https://furiosa.ai/rngd https://furiosa.ai/rngd You can see them admit that RNGD will be slower than a setup with H100 SXM cards, but at the same time the tokens per second per watt is way better! Actually, I wonder how different that is from Cerebras chips, since they're very much optimized for speed and one would think that'd also affect the efficiency a whole bunch: https://www.cerebras.ai/ https://www.cerebras.ai/
- bradfa 9mo agoHaving only 48GB of RAM per card seems low. The full server system with 8 cards barely has enough RAM to run modern large open models. And batching together user requests eats quite a lot of memory, too. Curious to see how these machines and cards are received by the market.
- Barathkanna 9mo agoFor those wondering how this differs from Nvidia GPUs: Nvidia = flexible, general-purpose GPUs that excel at training and mixed workloads. Furiosa = purpose-built inference ASICs that trade flexibility for much better cost, power efficiency, and predictable latency at scale.
- zvqcMMV6Zcr 9mo agoIt misses most important information, price and how quick they can ship. If they can actually deliver and take slice of market share from NVidia then it would make me happy.
- pama 9mo agoThe title sounds interesting but I get errors and no content on my iPhone15 because it is unable to initialize WebGL. Why do people still link content to such capabilities? Where has simple HTML / CSS gone these days? Edit: from comments and reading the one page that loads, this is still the 5nm tech they announced in 2024, hence the H100 comparison, which feels dated given the availability of GB300.
- torginus 9mo agoThese things never pan out. The reasons why this almost never works is one of the following: - They assume they can move hardware complexity (scheduling etc, access patterns into software). The magic compiler/runtime never arrives. - They assume their hard-to-program but faster architecture will get figured out by devs. It won't. - They assume a certain workload. The workload changes, and their arch is no longer optimal or possibly even workable. - But most importantly, they don't understand the fundamental bottlenecks, which is usually memory bandwidth. Even if you increase the paper specs, like FLOPS total, FLOPS/W etc. youre usually limited by how much you can read from memory. Which is exactly as much as their competitors. The way you can overcome this is by cleverness and complexity (cache lines, smarter algorithms, acceleration structures etc), but all these require a complex computer to run with all those coherent cache hierarchies, branching and synchronization logic etc. Which is why folks like NVIDIA keep going on despite facing this constant barrage of would-be disruptors. In fact this continue to be more and more true - memory bandwidth relies on transcievers on the chip edge, and if the size of the chips doesn't increase, bandwidth doesn't increase automatically on newer process nodes. Latency doesn't improve at all. But you get more transistors to play with, which you can use to run your workload more cleverly. In fact I don't rule out the possibility of CPU based massively parallel compute making a comeback.
- aduffy 9mo ago> - They assume their hard-to-program but faster architecture will get figured out by devs. It won't. Or it will get figured out in the niche fields where people are willing to figure out really hard stuff to squeeze out max performance (PE, hedge funds, intelligence) Either way agree, it's hard to get mass adoption without the software ecosystem feeding back in
- vagab0nd 9mo agoAnd when you layer on top networking, it's another level of sw/hw complexity.
- bicepjai 9mo agoAre all these improvement over custom kernel efficiency code ? Can we bring these to consumer RTX and Pro cards ? After I read the article :) The improvements in FuriosaAI's NXT RNGD Server are primarily driven by hardware innovations, not software or code changes.
- rajhlinux 9mo agoThey just declined Meta's $800 million offer. What are they smoking? I just saw the specs and nothing is special about the Furiosa RNGD Gen 2 card compared to the RTX 5090. Sure, it has more SRAM, but that is not a deal breaker. The same goes for power consumption, data centers have incentives for power. If each Furiosa RNGD Gen 2 card costs $10k while an RTX 5090 costs $2k, and the RTX 5090 has better performance for LLMs, you have to be mad stupid, have a personal grudge against Nvidia, or just want to burn cash for no good reason to rack up your data centers with Furiosa. The value of their company is going to diminish and their next offer won't go over $1.5 billion. It will actually be less than $800 million since every year Nvidia, Intel, and other AI hardware startups introduce a better and faster card. If Furiosa cards magically became cheaper than Nvidia's similar hardware, Furiosa might be worth a quarter billion dollars. I highly doubt this would ever happen because making AI compute with cutting edge lithography is hella expensive and involves heavy politics.