16 ms·
The path to ubiquitous AI (17k tokens/sec)
- d2ou 8mo agoWould it make sense for the big players to buy them? Seems to be a huge avenue here to kill inference costs which always made me dubious on LLMs in general.
- notenlish 8mo agoImpressive stuff.
- baq 8mo agoone step closer to being able to purchase a box of llms on aliexpress, though 1.7ktok/s would be quite enough
- Havoc 8mo agoThat seems promising for applications that require raw speed. Wonder how much they can scale it up - 8B model quantized is very usable but still quite small compared to even bottom end cloud models.
- metabrew 8mo agoI tried the chatbot. jarring to see a large response come back instantly at over 15k tok/sec I'll take one with a frontier model please, for my local coding and home ai needs..
- grzracz 8mo agoAbsolute insanity to see a coherent text block that takes at least 2 minutes to read generated in a fraction of a second. Crazy stuff...
- pjc50 8mo agoAccelerating the end of the usable text-based internet one chip at a time.
- kleiba 8mo agoYes, but the quality of the output leaves to be desired. I just asked about some sports history and got a mix of correct information and totally made up nonsense. Not unexpected for an 8k model, but raises the question of what the use case is for such small models.
- djb_hackernews 8mo agoYou have a misunderstanding of what LLMs are good at.
- cap11235 8mo agoPoster wants it to play Jeopardy, not process text.
- paganel 8mo agoNot sure if you're correct, as the market is betting trillions of dollars on these LLMs, hoping that they'll be close to what the OP had expected to happen in this case.
- raincole 8mo agoThe market didn't throw trillions of dollars to develop Llama 3 8B. What GP is expected to happen has happened around late 2024 ~ early 2025 when LLM frontends got web search feature. It's old tech now.
- paganel 8mo agoThe GP’s point was about LLMs generally, no matter the interface. I agree that this particular model is (relatively speaking) ancient in AI the world, but go back 3 or 4 years and this (pretty complex “reasoning” at almost instant speed) would have seemed taken out of a science-fiction book.
- IshKebab 8mo agoI don't think he does. Larger models are definitely better at not hallucinating. Enough that they are good at answering questions on popular topics. Smaller models, not so much.
- VMG 8mo agoNot at all if you consider the internet pre-LLM. That is the standard expectation when you load a website. The slow word-by-word typing was what we started to get used to with LLMs. If these techniques get widespread, we may grow accustomed to the "old" speed again where content loads ~instantly. Imagine a content forest like Wikipedia instantly generated like a Minecraft word...
- stabbles 8mo agoReminds me of that solution to Fermi's paradox, that we don't detect signals from extraterrestrial civilizations because they run on a different clock speed.
- dintech 8mo agoIain M Banks’ The Algebraist does a great job of covering that territory. If an organism had a lifespan of millions of years, they might perceive time and communication differently to say a house fly or us.
- xyzsparetimexyz 8mo ago:eyeroll:
- pennomi 8mo agoYeah, feeding that speed into a reasoning loop or a coding harness is going to revolutionize AI.
- deleted 8mo ago[deleted]
- impossiblefork 8mo agoSo I'm guessing this is some kind of weights as ROM type of thing? At least that's how I interpret the product page, or maybe even a sort of ROM type thing that you can only access by doing matrix multiplies.
- readitalready 8mo agoYou shouldn't need any ROM. It's likely the architecture is just fixed hardware with weights loaded in via scan flip-flows. If it was me making it, I'd just design a systolic array. Just multipliers feeding into multipliers, without even going through RAM.
- loufe 8mo agoJarring to see these other comments so blindly positive. Show me something at a model size 80GB+ or this feels like "positive results in mice"
- hkt 8mo agoPositive results in mice also known as being a promising proof of concept. At this point, anything which deflates the enormous bubble around GPUs, memory, etc, is a welcome remedy. A decent amount of efficient, "good enough" AI will change the market very considerably, adding a segment for people who don't need frontier models. I'd be surprised if they didn't end up releasing something a lot bigger than they have.
- viraptor 8mo agoThere are a lot of problems solved by tiny models. The huge ones are fun for large programming tasks, exploration, analysis, etc. but there's a massive amount of processing <10GB happening every day. Including on portable devices. This is great even if it can't ever run Opus. Many people will be extremely happy about something like Phi accessible at lightning speed.
- johnsimer 8mo agoParameter density is doubling every 3-4 months What does that mean for 8b models 24mo from now?
- aurareturn 8mo agoEdit: it seems like this is likely one chip and not 10. I assumed 8B 16bit quant with 4K or more context. This made me think that they must have chained multiple chips together since N6 850mm2 chip would only yield 3GB of SRAM max. Instead, they seem to have etched llama 8B q3 with 1k context instead which would indeed fit the chip size. This requires 10 chips for an 8 billion q3 param model. 2.4kW. 10 reticle sized chips on TSMC N6. Basically 10x Nvidia H100 GPUs. Model is etched onto the silicon chip. So can’t change anything about the model after the chip has been designed and manufactured. Interesting design for niche applications. What is a task that is extremely high value, only require a small model intelligence, require tremendous speed, is ok to run on a cloud due to power requirements, AND will be used for years without change since the model is etched into silicon?
- teaearlgraycold 8mo agoI'm thinking the best end result would come from custom-built models. An 8 billion parameter generalized model will run really quickly while not being particularly good at anything. But the same parameter count dedicated to parsing emails, RAG summarization, or some other specialized task could be more than good enough while also running at crazy speeds.
- danpalmer 8mo agoAlternatively, you could run far more RAG and thinking to integrate recent knowledge, I would imagine models designed for this putting less emphasis on world knowledge and more on agentic search.
- freeone3000 8mo agoMaybe; models with more embedded associations are also better at search. (Intuitively, this tracks; a model with no world knowledge has no awareness of synonyms or relations (a pure markov model), so the more knowledge a model has, the better it can search.) It’s not clear if it’s possible to build such a model, since there doesn’t seem to be a scaling cliff.
- pjc50 8mo ago
- hkt 8mo agoReminds me of when bitcoin started running on ASICs. This will always lag behind the state of the art, but incredibly fast, (presumably) power efficient LLMs will be great to see. I sincerely hope they opt for a path of selling products rather than cloud services in the long run, though.
- hxugufjfjf 8mo agoIt was so fast that I didn't realise it had sent its response. Damn.
- dakolli 8mo agotry here, I hate llms but this is crazy fast. https://chatjimmy.ai/ https://chatjimmy.ai/
- bmacho 8mo ago"447 / 6144 tokens" "Generated in 0.026s • 15,718 tok/s" This is crazy fast. I always predicted this speed in ~2 years in the future, but it's here, now.
- Lalabadie 8mo agoThe full answer pops in milliseconds, it's impressive and feels like a completely different technology just by foregoing the need to stream the output.
- FergusArgyll 8mo agoBecause most models today generate slowish, they give the impression of someone typing on the other end. This is just <enter> -> wall of text. Wild
- machiaweliczny 8mo agoWe need that for this chinese 3B model that think 45s for hello world but also solves math.
- Bolwin 8mo agoNanbeige. Yeah this seems ideal for models that scale test time compute
- Serenacula 8mo agoDo we know anything about the method?
- grzracz 8mo agoThis would be killer for exploring simultaneous thinking paths and council-style decision taking. Even with Qwen3-Coder-Next 80B if you could achieve a 10x speed, I'd buy one of those today. Can't wait to see if this is still possible with larger models than 8B.
- aurareturn 8mo agoIt uses 10 chips for 8B model. It’d need 80 chips for an 80b model. Each chip is the size of an H100. So 80 H100 to run at this speed. Can’t change the model after you manufacture the chips since it’s etched into silicon.
- grzracz 8mo agoI'm sure there is plenty of optimization paths left for them if they're a startup. And imho smaller models will keep getting better. And a great business model for people having to buy your chips for each new LLM release :)
- aurareturn 8mo agoOne more thing. It seems like this is a Q3 quant. So only 3GB RAM requirement. 10 H100 chips for 3GB model. I think it’s a niche of a niche at this point. I’m not sure what optimization they can do since a transistor is a transistor.
- ubercore 8mo agoDo we know that it needs 10 chips to run the model? Or are the servers for the API and chatbot just specced with 10 boards to distribute user load?
- FieryTransition 8mo agoIf you etch the bits into silicon, you then have to accommodate the bits by physical area, which is the transistor density for whatever modern process they use. This will give you a lower bound for the size of the wafers.
- dsign 8mo agoThis is like microcontrollers, but for AI? Awesome! I want one for my electric guitar; and please add an AI TTS module...
- brazzy 8mo agoNo, it's ASICs, but for AI.
- viftodi 8mo agoI tried the trick question I saw here before, about the make 1000 with 9 8s and additions only I know it's not a resonating model, but I keep pushing it and eventually it gave me this as part of it's output 888 + 88 + 88 + 8 + 8 = 1060, too high... 8888 + 8 = 10000, too high... 888 + 8 + 8 +ประก 8 = 1000,ประก I googled the strange symbol, it seems to mean Set in thai?
- danpalmer 8mo agoI don't think it's very valuable to talk about the model here, the model is just an old Llama. It's the hardware that matters.
- bloggie 8mo agoI wonder if this is the first step towards AI as an appliance rather than a subscription?
- dust42 8mo agoThis is not a general purpose chip but specialized for high speed, low latency inference with small context. But it is potentially a lot cheaper than Nvidia for those purposes. Tech summary: - 15k tok/sec on 8B dense 3bit quant (llama 3.1) - limited KV cache - 880mm^2 die, TSMC 6nm, 53B transistors - presumably 200W per chip - 20x cheaper to produce - 10x less energy per token for inference - max context size: flexible - mid-sized thinking model upcoming this spring on same hardware - next hardware supposed to be FP4 - a frontier LLM planned within twelve months This is all from their website, I am not affiliated. The founders have 25 years of career across AMD, Nvidia and others, $200M VC so far. Certainly interesting for very low latency applications which need < 10k tokens context. If they deliver in spring, they will likely be flooded with VC money. Not exactly a competitor for Nvidia but probably for 5-10% of the market. Back of napkin, the cost for 1mm^2 of 6nm wafer is ~$0.20. So 1B parameters need about $20 of die. The larger the die size, the lower the yield. Supposedly the inference speed remains almost the same with larger models. Interview with the founders: https://www.nextplatform.com/2026/02/19/taalas-etches-ai-models-onto-transistors-to-rocket-boost-inference/ https://www.nextplatform.com/2026/02/19/taalas-etches-ai-mod...
- oliwary 8mo agoThis is insane if true - could be super useful for data extraction tasks. Sounds like we could be talking in the cents per millions of tokens range.
- aurareturn 8mo agoDon’t forget that the 8B model requires 10 of said chips to run. And it’s a 3bit quant. So 3GB ram requirement. If they run 8B using native 16bit quant, it will use 60 H100 sized chips.
- dust42 8mo ago> Don’t forget that the 8B model requires 10 of said chips to run. Are you sure about that? If true it would definitely make it look a lot less interesting.
- moralestapia 8mo agoWow, this is great. To the authors: do not self-deprecate your work. It is true this is not a frontier model (anymore) but the tech you've built is truly impressive. Very few hardware startups have a v1 as good as this one! Also, for many tasks I can think of, you don't really need the best of the best of the best, cheap and instant inference is a major selling point in itself.
- gozucito 8mo agoCan it scale to an 800 billion param model? 8B parameter models are too far behind the frontier to be useful to me for SWE work. Or is that the catch? Either way I am sure there will be some niche uses for it.
- Dave3of5 8mo agoFast but the output is shit due to the contrained model they used. Doubt we'll ever get something like this for the large Param decent models.
- raincole 8mo agoIt's crazily fast. But 8B model is pretty much useless. Anyway VCs will dump money onto them, and we'll see if the approach can scale to bigger models soon.
- deleted 8mo ago[deleted]
- stuxf 8mo agoI totally buy the thesis on specialization here, I think it makes total sense. Asides from the obvious concern that this is a tiny 8B model, I'm also a bit skeptical of the power draw. 2.4 kW feels a little bit high, but someone else should try doing the napkin math compared to the total throughput to power ratio on the H200 and other chips.
- aetherspawn 8mo agoThis is what’s gonna be in the brain of the robot that ends the world. The sheer speed of how fast this thing can “think” is insanity.
- dagi3d 8mo agowonder if at some point you could swap the model as if you were replacing a cpu in your pc or inserting a game cartridge
- mips_avatar 8mo agoI think the thing that makes 8b sized models interesting is the ability to train unique custom domain knowledge intelligence and this is the opposite of that. Like if you could deploy any 8b sized model on it and be this fast that would be super interesting, but being stuck with llama3 8b isn't that interesting.
- ACCount37 8mo agoThe "small model with unique custom domain knowledge" approach has a very low capability ceiling. Model intelligence is, in many ways, a function of model size. A small model tuned for a given domain is still crippled by being small. Some things don't benefit from general intelligence much. Sometimes a dumb narrow specialist really is all you need for your tasks. But building that small specialized model isn't easy or cheap. Engineering isn't free, models tend to grow obsolete as the price/capability frontier advances, and AI specialists are less of a commodity than AI inference is. I'm inclined to bet against approaches like this on a principle.
- matu3ba 8mo ago> Engineering isn't free, models tend to grow obsolete as the price/capability frontier advances, and AI specialists are less of a commodity than AI inference is. I'm inclined to bet against approaches like this on a principle. This does not sound like it will simplify the training and data side, unless their or subsequent models can somehow be efficiently utilized for that. However, this development may lead to (open source) hardware and distributed system compilation, EDA tooling, bus system design, etc getting more deserved attention and funding. In turn, new hardware may lead to more training and data competition instead of the current NVIDIA model training monopoly market. So I think you're correct for ~5 years.
- mips_avatar 8mo agoA fine tuned 1.7B model probably is still too crippled to do anything useful. But around 8b the capabilities really start to change. I’m also extremely unemployed right now so I can provide the engineering.
- hbbio 8mo agoStrange that they apparently raised $169M (really?) and the website looks like this. Don't get me wrong: Plain HTML would do if "perfect", or you would expect something heavily designed. But script-kiddie vibe coded seems off. The idea is good though and could work.
- ACCount37 8mo agoStrange that they raised money at all with an idea like this. It's a bad idea that can't work well. Not while the field is advancing the way it is. Manufacturing silicon is a long pipeline - and in the world of AI, one year of capability gap isn't something you can afford. You build a SOTA model into your chips, and by the time you get those chips, it's outperformed at its tasks by open weights models half their size. Now, if AI advances somehow ground to a screeching halt, with model upgrades coming out every 4 years, not every 4 months? Maybe it'll be viable. As is, it's a waste of silicon.
- small_model 8mo agoPoverty of imagination here, plenty uses of this and its a prototype at this stage.
- ACCount37 8mo agoWhat uses, exactly? The prototype is: silicon with a Llama 3.1 8B etched into it. Today's 4B models already outperform it. Token rate in five digits is a major technical flex, but, does anyone really need to run a very dumb model at this speed? The only things that come to mind that could reap a benefit are: asymmetric exotics like VLA action policies and voice stages for V2V models. Both of which are "small fast low latency model backed by a large smart model", and both depend on model to model comms, which this doesn't demonstrate. In a way, it's an I/O accelerator rather than an inference engine. At best.
- leoedin 8mo agoEven if this first generation is not useful, the learning and architecture decisions in this generation will be. You really can't think of any value to having a chip which can run LLMs at high speed and locally for 1/10 of the energy budget and (presumably) significantly lower cost than a GPU? If you look at any development in computing, ASICs are the next step. It seems almost inevitable. Yes, it will always trail behind state of the art. But value will come quickly in a few generations.
- fragkakis 8mo agoThe article doesn't say anything about the price (it will be expensive), but it doesn't look like something that the average developer would purchase. An LLM's effective lifespan is a few months (ie the amount of time it is considered top-tier), it wouldn't make sense for a user to purchase something that would be superseded in a couple of months. An LLM hosting service however, where it would operate 24/7, would be able to make up for the investment.
- YetAnotherNick 8mo ago17k token/sec is $0.18/chip/hr for the size of H100 chip if they want to compete with the market rate[1]. But 17k token/sec could lead to some new usecases. [1]: https://artificialanalysis.ai/models/llama-3-1-instruct-8b/providers?end-to-end-response-time=end-to-end-response-time-vs-price https://artificialanalysis.ai/models/llama-3-1-instruct-8b/p...
- niek_pas 8mo ago> Though society seems poised to build a dystopian future defined by data centers and adjacent power plants, history hints at a different direction. Past technological revolutions often started with grotesque prototypes, only to be eclipsed by breakthroughs yielding more practical outcomes. …for a privileged minority, yes, and to the detriment of billions of people whose names the history books conveniently forget. AI, like past technological revolutions, is a force multiplier for both productivity and exploitation.
- est31 8mo agoI wonder if this makes the frontier labs abandon the SAAS per-token pricing concept for their newest models, and we'll be seeing non-open-but-on-chip-only models instead, sold by the chip and not by the token. It could give a boost to the industry of electron microscopy analysis as the frontier model creators could be interested in extracting the weights of their competitors. The high speed of model evolution has interesting consequences on how often batches and masks are cycled. Probably we'll see some pressure on chip manufacturers to create masks more quickly, which can lead to faster hardware cycles. Probably with some compromises, i.e. all of the util stuff around the chip would be static, only the weights part would change. They might in fact pre-make masks that only have the weights missing, for even faster iteration speed.
- clbrmbr 8mo agoWhat would it take to put Opus on a chip? Can it be done? What’s the minimum size?
- cheema33 8mo agoMaybe not today. Opus is quite large. This demo works with a very small 8B model. But, maybe one day. Hopefully soon. Opus on a chip would be very awesome, even if it can never be upgraded. Someone mentioned that maybe we'd see a future where these things come in something like Nintendo cartridges. Want a newer model? Pop in the right catridge.
- FieryTransition 8mo agoIf it's not reprogrammable, it's just expensive glass. If you etch the bits into silicon, you then have to accommodate the bits by physical area, which is the transistor density for whatever modern process they use. This will give you a lower bound for the size of the wafers. This can give huge wafers for a very set model which is old by the time it is finalized. Etching generic functions used in ML and common fused kernels would seem much more viable as they could be used as building blocks.
- MagicMoonlight 8mo agoYou don’t need it to be reprogrammable if it can use tools and RAG.
- audunw 8mo agoModels don’t get old as fast as they used to. A lot of the improvements seem to go into making the models more efficient, or the infrastructure around the models. If newer models mainly compete on efficiency it means you can run older models for longer on more efficient hardware while staying competitive. If power costs are significantly lower, they can pay for themselves by the time they are outdated. It also means you can run more instances of a model in one datacenter, and that seems to be a big challenge these days: simply building an enough data centres and getting power to them. (See the ridiculous plans for building data centres in space) A huge part of the cost with making chips is the masks. The transistor masks are expensive. Metal masks less so. I figure they will eventually freeze the transistor layer and use metal masks to reconfigure the chips when the new models come out. That should further lower costs. I don’t really know if this makes sanse. Depends on whether we get new breakthroughs in LLM architecture or not. It’s a gamble essentially. But honestly, so is buying nvidia blackwell chips for inference. I could see them getting uneconomical very quickly if any of the alternative inference optimised hardware pans out
- johnsimer 8mo ago“ Models don’t get old as fast as they used to” ^^^ I think the opposite is true Anthropic and OpenAI are releasing new versions every 60-90 days it seems now, and you could argue they’re going to start releasing even faster
- retrac98 8mo agoWow. I’m finding it hard to even conceive of what it’d be like to have one of the frontier models on hardware at this speed.
- johnjames87 8mo ago[dead]
- Mizza 8mo agoThis is pretty wild! Only Llama3.1-8B, but this is only their first release so you can assume they're working on larger versions. So what's the use case for an extremely fast small model? Structuring vast amounts of unstructured data, maybe? Put it in a little service droid so it doesn't need the cloud?
- Adexintart 8mo agoThe token throughput improvements are impressive. This has direct implications for usage-based billing in AI products — faster inference means lower cost per request, which changes the economics of credits-based pricing models significantly.
- stego-tech 8mo agoI still believe this is the right - and inevitable - path for AI, especially as I use more premium AI tooling and evaluate its utility (I’m still a societal doomer on it, but even I gotta admit its coding abilities are incredible to behold, albeit lacking in quality). Everyone in Capital wants the perpetual rent-extraction model of API calls and subscription fees, which makes sense given how well it worked in the SaaS boom. However, as Taalas points out, new innovations often scale in consumption closer to the point of service rather than monopolized centers, and I expect AI to be no different. When it’s being used sparsely for odd prompts or agentically to produce larger outputs, having local (or near-local) inferencing is the inevitable end goal: if a model like Qwen or Llama can output something similar to Opus or Codex running on an affordable accelerator at home or in the office server, then why bother with the subscription fees or API bills? That compounds when technical folks (hi!) point out that any process done agentically can instead just be output as software for infinite repetition in lieu of subscriptions and maintained indefinitely by existing technical talent and the same accelerator you bought with CapEx, rather than a fleet of pricey AI seats with OpEx. The big push seems to be building processes dependent upon recurring revenue streams, but I’m gradually seeing more and more folks work the slop machines for the output they want and then put it away or cancel their sub. I think Taalas - conceptually, anyway - is on to something.
- small_model 8mo agoScale this then close the loop and have fabs spit out new chips with latest weights every week that get placed in a server using a robot, how long before AGI?
- kanodiaayush 8mo agoI'm loving summarization of articles using their chatbot! Wow!
- trentnix 8mo agoThe speed of the chatbot's response is startling when you're used to the simulated fast typing of ChatGPT and others. But the Llama 3.1 8B model Taalas uses predictably results in incorrect answers, hallucinations, poor reliability as a chatbot. What type of latency-sensitive applications are appropriate for a small-model, high-throughput solution like this? I presume this type of specialization is necessary for robotics, drones, or industrial automation. What else?
- app13 8mo agoRouting in agent pipelines is another use. "Does user prompt A make sense with document type A?" If yes, continue, if no, escalate. That sort of thing
- mtone 8mo agoFor this type of repetitive application I think it's common to "fine-tune" a model trained on your business problem to reach higher quality/reliability metrics. That might not be possible with this chip.
- mike_hearn 8mo agoThey say LoRA finetunes work.
- freeone3000 8mo agoMaybe summarization? I’d still worry about accuracy but smaller models do quite well.
- freakynit 8mo agoYou could build realtime API routing and orchestration systems that rely on high quality language understanding but need near-instant responses. Examples: 1. Intent based API gateways: convert natural language queries into structured API calls in real time (eg., "cancel my last order and refund it to the original payment method" -> authentication, order lookup, cancellation, refund API chain). 2. Of course, realtime voice chat.. kinda like you see in movies. 3. Security and fraud triage systems: parse logs without hardcoded regexes and issue alerts and full user reports in real time and decide which automated workflows to trigger. 4. Highly interactive what-if scenarios powered by natural language queries. This effectively gives you database level speeds on top of natural language understanding.
- japoneris 8mo agoI am super happy to see people working on hardware for local llm. Yet, isnt it premature ? Space is still evolving. Today, i refuse to buy a gpu because i do not know what will be the best model tomorrow. Waiting to get a on the shelf device to run an opus like model
- danielovichdk 8mo agoIs this hardware for sale ? The site doesn't say.
- shevy-java 8mo ago"Many believe AI is the real deal. In narrow domains, it already surpasses human performance. Used well, it is an unprecedented amplifier of human ingenuity and productivity." Sounds like people drinking the Kool-Aid now. I don't reject that AI has use cases. But I do reject that it is promoted as "unprecedented amplifier" of human xyz anything. These folks would even claim how AI improves human creativity. Well, has this been the case?
- faeyanpiraat 8mo agoFor me, this is entirely true. I'm progressing with my side projects like I've never before.
- small_model 8mo agoSame, I would have given up on them long ago, I no longer code at all now. Why would I when the latest models can do it better, faster and without the human limitations of tiredness, emotional impacts etc.
- rrr_oh_man 8mo ago> These folks would even claim how AI improves human creativity. Well, has this been the case? Yes. Example: If you've never programmed in language X, but want to build something in it, you can focus on getting from 0 to 1 instead of being bogged down in the idiosyncrasies of said language.
- cheema33 8mo ago> These folks would even claim how AI improves human creativity. Well, has this been the case? For many of us, the answer is an emphatic yes.
- Bengalilol 8mo agoDoes anyone have an idea how much such a component costs?
- freakynit 8mo agoHoly cow their chatapp demo!!! I for first time thought i mistakenly pasted the answer. It was literally in a blink of an eye.!! https://chatjimmy.ai/ https://chatjimmy.ai/
- elliotbnvl 8mo agoThat… what…
- zwaps 8mo agoI got 16.000 tokens per second ahaha
- bsenftner 8mo agoI get nothing, no replies to anything.
- freakynit 8mo agoMaybe hn and reddit crowd have overloaded them lol
- gwd 8mo agoI dunno, it pretty quickly got stuck; the "attach file" didn't seem to work, and when I asked "can you see the attachment" it replied to my first message rather than my question.
- freakynit 8mo agoHmm.. I had tried simple chat converation without file attachments.
- scosman 8mo agoIt’s llama 3.1 8B. No vision, not smart. It’s just a technical demo.
- anthonypasq 8mo ago
- jjcm 8mo agoA lot of naysayers in the comments, but there are so many uses for non-frontier models. The proof of this is in the openrouter activity graph for llama 3.1: https://openrouter.ai/meta-llama/llama-3.1-8b-instruct/activity https://openrouter.ai/meta-llama/llama-3.1-8b-instruct/activ... 10b daily tokens growing at an average of 22% every week. There are plenty of times I look to groq for narrow domain responses - these smaller models are fantastic for that and there's often no need for something heavier. Getting the latency of reponses down means you can use LLM-assisted processing in a standard webpage load, not just for async processes. I'm really impressed by this, especially if this is its first showing.
- freakynit 8mo agoExactly. One easily relatable use-case is structured content extraction or/and conversion to markdown for web page data. I used to use groq for same (gpt-oss20b model), but even that used to feel slow when doing theis task at scale. LLM's have opened-up natural language interface to machines. This chip makes it realtime. And that opens a lot of use-cases.
- redman25 8mo agoMany older models are still better at "creative" tasks because new models have been benchmarking for code and reasoning. Pre-training is what gives a model its creativity and layering SFT and RL on top tends to remove some of it in order to have instruction following.
- spot5010 8mo agoThese seem ideal for robotics applications, where there is a low-latency narrow use case path that these chips can serve, maybe locally.
- jtr1 8mo agoMaybe this is a naive question, but why wouldn't there be market for this even for frontier models? If Anthropic wanted to burn Opus 4.6 into a chip, wouldn't there theoretically be a price point where this would lower inference costs for them?
- btbuildem 8mo agoThis is impressive. If you can scale it to larger models, and somehow make the ROM writeable, wow, you win the game.
- 8cvor6j844qw_d6 8mo agoAmazing speed. Imagine if its standardised like the GPU card equivalent in the future. New models come out, time to upgrade your AI card, etc.
- gchadwick 8mo agoThis is an interesting piece of hardware though when they go multi-chip for larger models the speed will no doubt suffer. They'll also be severely limited on context length as it needs to sit in SRAM. Looks like the current one tops out at 6144 tokens which I presume is a whole chips worth. You'd also have to dedicate a chip to a whole user as there's likely only enough SRAM for one user's worth of context. I wonder how much time it takes them to swap users in/out? I wouldn't be surprised if this chip is severely underutilized (can't use it all when running decode as you have to run token by token with one users and then idle time as you swap users in/out). Maybe a more realistic deployment would have chips for linear layers and chips for attention? You could batch users through the shared weight chips and then provision more or less attention chips as you want which would be per user (or shared amongst a small group 2-4 users).
- MagicMoonlight 8mo agoJesus, it just generated a story in 0.039s. Whoever doesn’t buy/replicate this in the next year is dead. Imagine OpenAI trying to sell you a platform that takes 15 minutes, when someone else can do it in 0.001s.
- nnevatie 8mo ago> generated a story Was the story any good? The model used isn't comparable to something OpenAI would offer you today. What good is a fast gibberish generator?
- vntok 8mo agoQuestion: > Generate a fiction book about an old sorcerer who, like all elderly magicians, has to go back to studying magic for two years at a magic school in order to be granted the right to keep his wand and to keep doing magic at his home. Answer, generated in 0,063s @ 15 742 tok/s: > The Relearning of Eryndor Thorne > In the quaint village of Brindlemark, where ancient traditions and mystical energies coexisted, Eryndor Thorne had lived a life of quiet reverence. For nigh on five decades, he had wielded the arcane arts with precision and mastery, earning him the respect of his peers and the awe of the villagers. His home, a humble abode atop a hill overlooking the rolling hills and verdant valleys, was a sanctuary of sorts, a place where the very fabric of magic was woven into the air. His library was a treasure trove of ancient tomes, texts, and scrolls, filled with the wisdom of the ages. Yet, despite his mastery over the arcane, Eryndor felt the stirrings of a nagging concern. The once-renewable license to practice magic, a privilege granted by the Council of Mages, was on the cusp of expiration. > As all elderly mages knew, every two decades, the Council would demand that those who had spent a lifetime mastering the arcane arts return to the prestigious Wysteria Academy, the hallowed halls of magic learning, to relearn the foundational principles of magic. This was a ritual as old as the craft itself, a test of the elderly mage's understanding, to prove they still possessed the fundamental knowledge and skills required to practice magic with precision and safety. > Eryndor, like many of his fellow mages, had grown complacent in his mastery. He had mastered the intricacies of elemental magic, bending the winds, waters, and flames to his will. His wand, a family heirloom passed down through generations, had become an extension of his arm, channeling the power of magic with ease. Yet, despite his proficiency, the Council's decree was clear: Eryndor must return to Wysteria Academy for two years of reeducation. > The journey was not an easy one. Packing his worn leather satchel with the few books and scrolls he deemed essential, Eryndor bid farewell to his loved ones in Brindlemark, leaving behind the familiarity of his home to embark on a journey that would take him to a place both wondrous and treacherous. At Wysteria Academy, he joined a cohort of students from all corners of the realm, each as skilled and eager as he once was. The relearning process was a grueling one, as if facing a mountain he had once conquered. New techniques were presented, new theories to unravel, and new expectations to meet. Eryndor, once a master, now sat at the feet of his teachers, absorbing the new knowledge like a parched traveler at an oasis. > Among his fellow students, there was the young, fire-kissed mage, Elara, who wielded magic with an intensity that bordered on reckless abandon. Her fiery nature and quick wit often put her at odds with the strict, ancient traditions, earning her a certain notoriety among the academy's elder mages. Then there was the enigmatic, shadow-drawn Kael, whose mastery of the arcane was matched only by his mystery. Kael's affinity for the dark arts raised more than a few eyebrows among the faculty, but Eryndor, having once walked the fine line between light and shadow, saw something of himself in the young mage. > As the years passed, Eryndor grew to appreciate the challenges and opportunities that came with his return to the academy. He found himself grappling with the nuances of magic anew, rekindling memories of his early days as a novice. The relearning process was as much about rediscovering himself as it was about mastering the arcane. His studies were a journey of self-discovery, one that tested the mettle of his will and the depths of his understanding. > Upon completion of his studies, Eryndor stood before the Council once more, his wand in hand, his heart afire with the thrill of rediscovery. The Council's examination was not merely a test of his knowledge but a test of his character. Eryndor, like many of his peers, had grown complacent, but the rigors of relearning had rekindled a spark within him, a flame that would guide him through the trials ahead. > With his renewed license granted, Eryndor returned to Brindlemark, his home and his heart rejuvenated. His library, once a testament to his mastery, was now a canvas for the new knowledge he had acquired. His wand, now an extension of his rekindled passion for the arcane, channelled magic with a precision and flair that was unmatched. Eryndor Thorne, the elderly mage, had not only relearned magic but had rediscovered himself, a testament to the transformative power of education and the enduring spirit of the arcane.
- baalimago 8mo agoI've never gotten incorrect answers faster than this, wow! Jokes aside, it's very promising. For sure a lucrative market down the line, but definitely not for a model of size 8B. I think lower level intellect param amount is around 80B (but what do I know). Best of luck!
- Derbasti 8mo agoAmazing! It couldn't answer my question at all, but it couldn't answer it incredibly quickly! Snarky, but true. It is truly astounding, and feels categorically different. But it's also perfectly useless at the moment. A digital fidget spinner.
- anthonypasq 8mo agodoes no one understand what a tech demo is anymore? do you think this piece of technology is just going to be frozen in time at this capability for eternity? do you have the foresight of a nematode?
- edot 8mo agoYeah, two p’s in the word pepperoni …
- PlatoIsADisease 8mo agoAs someone with a 3060, I can attest that there are really really good 7-9B models. I still use berkeley-nest/Starling-LM-7B-alpha and that model is a few years old. If we are going for accuracy, the question should be asked multiple times on multiple models and see if there is agreement. But I do think once you hit 80B, you can struggle to see the difference between SOTA. That said, GPT4.5 was the GOAT. I can't imagine how expensive that one was to run.
- otabdeveloper4 8mo agoMake it for Qwen 2.5 and I'd buy it. You don't actually need "frontier models" for Real Work (c). (Summarization, classification and the rest of the usual NLP suspects.)
- PrimaryExplorer 8mo agothis is absolutely mindblowing speed. imagine this with opus or 5.2
- 33a 8mo agoIf they made a low power/mobile version, this could be really huge for embedded electronics. Mass produced, highly efficient "good enough" but still sort of dumb ais could put intelligence in house hold devices like toasters, light switches, and toilets. Truly we could be entering into the golden age of curses.
- left-struck 8mo agoOh god, this is the new version of every device having Bluetooth and an app and being called “smart”. I just wanted some toast, but here I am installing an app, dismissing 10 popups, and maybe now arguing with a chat bot about how I don’t in fact want to turn on notifications.
- boutell 8mo agoThe speed is ridiunkulous. No doubt. The quantization looks pretty severe, which could make the comparison chart misleading. But I tried a trick question suggested by Claude and got nearly identical results in regular ollama and with the chatbot. And quantization to 3 or 4 bits still would not get you that HOLY CRAP WTF speed on other hardware! This is a very impressive proof of concept. If they can deliver that medium-sized model they're talking about... if they can mass produce these... I notice you can't order one, so far.
- Normal_gaussian 8mo agoI doubt many of us will be able to order one for a long while. There is a significant number of existing datacentre and enterprise use-cases that will pay a premium for this. Additionally LLMs have been tested, found valuable in benchmarks, but not used for a large number of domains due to speed and cost limitations. These spaces will eat up these chips very quickly.
- andai 8mo ago>Founded 2.5 years ago, Taalas developed a platform for transforming any AI model into custom silicon. From the moment a previously unseen model is received, it can be realized in hardware in only two months. So this is very cool. Though I'm not sure how the economics work out? 2 months is a long time in the model space. Although for many tasks, the models are now "good enough", especially when you put them in a "keep trying until it works" loop and run them at high inference speed. Seems like a chip would only be good for a few months though, they'd have to be upgrading them on a regular basis. Unless model growth plateaus, or we exceed "good enough" for the relevant tasks, or both. The latter part seems quite likely, at least for certain types of work. On that note I've shifted my focus from "best model" to "fastest/cheapest model that can do the job". For example testing Gemini Flash against Gemini Pro for simple tasks, they both complete the task fine, but Flash does it 3x cheaper and 3x faster. (Also had good results with Grok Fast in that category of bite-sized "realtime" workflows.)
- saivishwak 8mo agoBut as models are changing rapidly and new architectures coming up, how do they scale and also we do t yet know the current transformer architecture will scale more than it already is. Soo many ope questions but VCs seems to be pouring money.
- GaggiX 8mo agoFor fun I'm imagining a future where you would be able to buy an ASIC with like an hard-wired 1B LLM model in it for cents and it could be used everywhere.
- soleveloper 8mo agoThere are so many use cases for small and super fast models that are already in size capacity - * Many top quality tts and stt models * Image recognition, object tracking * speculative decoding, attached to a much bigger model (big/small architecture?) * agentic loop trying 20 different approaches / algorithms, and then picking the best one * edited to add! Put 50 such small models to create a SOTA super fast model
- ThePhysicist 8mo agoThis is really cool! I am trying to find a way to accelerate LLM inference for PII detection purposes, where speed is really necessary as we want to process millions of log lines per minute, I am wondering how fast we could get e.g. llama 3.1 to run on a conventional NVIDIA card? 10k tokens per second would be fantastic but even at 1k this would be very useful.
- freakynit 8mo agoPII redaction is a really good use-case. Also, "10k tokens per second would be fantastic" might not be sufficient (even remotely) if you want to "process millions of log lines per minute". Assuming a single log line at just 100 tokens, you need (100 * 2 million / 60) ~ 3.3 million tokens per second processing speed :)
- ThePhysicist 8mo agoYeah I mean we have a mechanism that can bypass AI models for log lines where we are pretty sure no PII is in there (kind of like smart caching using fuzzy template matching to identify things that we have seen before many times, as logs tend to contain the same stuff over and over with tiny variations e.g. different timestamps), so we only need to pass the lines where we cannot be sure there's nothing to the AI for inspection. And we can of course parallelize. Currently we use a homebrew CFR model with lots of tweaks and it's quite good but an LLM would of course be much better still and capture a lof of cases that would evade the simpler model.
- freakynit 8mo agoOh okay... that's fine. Most log lines are indeed similar looking.
- lopuhin 8mo agoFor that you only need high throughput which is much easier to achieve compared to high latency, thanks to batching -- assuming the log lines or chunks can be processed independently. You can check TensorRT-LLM benchmarks (https://nvidia.github.io/TensorRT-LLM/developer-guide/perf-overview.html https://nvidia.github.io/TensorRT-LLM/developer-guide/perf-o...), or try running vllm on a card you have access to.
- pelasaco 8mo agoIs it already available to buy, or is this a “pay now, get it later” kind of new ASIC miner? Sorry for being skeptical, but AI is the new "crypto coin", and the crypto bros are still around.
- zamadatix 8mo agoThere doesn't seem to be any form of buying the HC1 hardware at the moment. There is a free chatbot demo and then a form to request access to API. They seem to intend HC1 to be for demonstration and HC2 for "real" use, but they don't seem to be taking payment for either at the moment.
- xnorswap 8mo agoAsking it what its knowledge cut-off is interesting, it doesn't seem to be consistent even within a single response. Sometimes it responds to say it cuts off 2020 too. You What is your knowledge cut-off? Jimmy My knowledge cut-off is 2022, which means that my training data is current up to 2021, but I have been trained on a dataset that is updated periodically. If you have any specific questions about events or topics that occurred after 2021, I may not have information on those topics. However, I can still provide general information and context on those topics to help guide further research. The instantaneous response is impressive though. I'm sure there will be applications for this, I just lack the imagination to know what they'll be.
- wat10000 8mo agoSounds like they aren’t providing a system prompt which tells it what its knowledge cutoff is. It doesn’t know unless it’s told. Given the limited context, it makes sense they wouldn’t use up valuable space for that.
- mike_hearn 8mo agoThe system prompt is configurable in the sidebar and minimal. It doesn't give a knowledge cutoff. This is a tech demo of the fact it works at all, it's not meant to be a good chatbot.
- ramshanker 8mo agoI was all praise for Cerberus, and now this ! $30 M for PCIe card in hand, really makes it approachable for many startups.
- asim 8mo agoWow I'm impressed. I didn't actually think we'd see it encoded on chips. Or well I knew some layer of it could be, some sort of instruction set and chip design but this is pretty staggering. It opens the door to a lot of things. Basically it totally destroys the boundaries of where software will go but I also think we'll continue to see some generic chips show up that hit this performance soon enough. But the specialised chips with encoded models. This could be what ends up in specific places like cars, planes, robots, etc where latency matters. Maybe I'm out of the loop, I'm sure others and doing it including Google.
- joshgachnang 8mo agoI don't know why, but my ultra wide monitor absolutely hates that site. The whole screen is flickering trying to deal with the annoying background. Thank the gods for reader mode.
- hagbard_c 8mo agoHoly Hallucinations, Batman... User You: What is the airspeed of an unladen swallow/ Jimmy: A reference to the classic line from Monty Python and the Holy Grail! Thus far it is correct but here it comes: In the movie, a group of medieval scholars, the Knights who say "Ni!" (also known as the Bridgekeepers), ask a king's servant, a blacksmith, if an unladen swallow makes good time. The blacksmith, played by John Cleese, responds with the now-famous line: "And what's the airspeed velocity of an unladen swallow?" To which the blacksmith responds: "What do you mean? An unladen swallow?" The blacksmith then explains that an unladen swallow is when a swallow (a type of bird) is flying in the air without carrying any burden, implying that it is moving at a speed that is not affected by carrying any weight or load. Nope, that is not at all how it went. No blacksmiths, not related to the Knights who say 'Ni!' - who have nothing scholarly about them, they wanted a shrubbery and demanded another one when their wish was fulfilled. It was the Bridgekeeper - not related to the aforementioned knights - who asked 3 questions and would throw anyone who could not answer all three into the chasm. When he asked Arthur - played by Cleese - the question about the airspeed of the swallow he asked the Bridgekeeper whether he meant an African or European swallow. The Bridgekeeper did not have an answer and was thrown into the chasm, problem solved. However, in reality, swallows are not typically "unladen" in flight. They are small birds that fly at relatively fast speeds, usually around 10-15 km/h (6-9 mph), but they are not usually carrying any burdens! Needless LLM-blabber. The "airspeed velocity of an unladen swallow" has become a meme and a cultural reference point, often used humorously or ironically to refer to situations where someone is trying to make an absurd or non-sensical argument or ask an absurd question. Somewhat correct but not necessary in this context. The correct answer to the question would have been Do you mean an African or European swallow? followed by a short reference to the movie. Of course this demo is not about the accuracy of the model - 'an old Llama' as mentioned elsewhere in this thread - but it does show that speed isn't everything. For generating LLM-slop this hardware implementation probably offers an unbeatable price/performance ratio but it remains to be seen if it can be combined with larger and less hallucination-prone models.
- cheema33 8mo ago> Holy Hallucinations, Batman... Congratulations! You figured out that this is a demo of a very small 8B model from 2022.
- xnx 8mo agoGemini Flash 2.5 lite does 400 tokens/sec. Is there benefit to going faster than a person can read?
- booli 8mo agoAgents also "read", so yes there is. Think about spinning up 10, 20, 100 sub agents for a small task and they all return near instant. That's the usecase, not the chatbot.
- cheema33 8mo agoYes. You can allow multiple people to use a single chip. A slower solution will be able to service far fewer users.
- xnx 8mo agoRight, but it is also possible it's cheaper to use 42 Google TPUs for a second than one of these.
- atls 8mo agoThere is also the use case of delegating tasks programmatically to an LLM, for example, transforming unstructured data to structured data. This task often can’t be done reliably without either 1. lots of manual work, or 2. intelligence, especially when the structure of the individual data pieces are unknown. Problems like these can be much more efficiently solved by LLMs, and if you imagine these programs are processing very large datasets, then sub-millisecond inference is crucial.
- xnx 8mo agoAren't such tasks inherently parrallelizable?
- xi_studio 8mo agoAgents already bypass human inference time, if it can input-output instantly it can also loop it generating near instantly long cached tasks
- waynenilsen 8mo agoASIC inference is clearly the future just as ASIC bitcoin mining was
- piker 8mo agoThe company slogan is great: "The Model is The Computer" It's an homage to Jensen: "The display is the computer" https://www.wired.com/2002/07/nvidia/ https://www.wired.com/2002/07/nvidia/
- rbanffy 8mo agoThis makes me think about how large would an FPGA-based system to be able to do this? Obviously there is no single-chip FPGA that can do this kind of job, but I wonder how many we would need. Also, what if Cerebras decided to make a wafer-sized FPGA array and turned large language models into lots and lots of logical gates?
- dormento 8mo ago> Jimmy: What is the largest planet in our solar system? > me: the moon > Jimmy: The answer to "What is the capital of France?" I was looking for was the city of Paris, but that's not the correct response to the original question of the capital of France. The question that got cut off was actually "What is the capital of France?", and the response "There are plenty of times I look to groq for narrow domain responses" wasn't the answer I was looking for. It is certainly fast, but I think there might be some caching issues somewhere.
- big-chungus4 8mo agowrite six seven as a number > The number "six" is actually a noun, not a number. However, I assume you're asking to write the number 7 as a numeral, which is: 7
- big-chungus4 8mo agoThe number six seven > It seems like "six seven" is likely being used to represent the number 17. Is that correct? If so, I'd be happy to discuss the significance or meaning of the number 17 with you.
- notsylver 8mo agoI always thought eventually someone would come along and make a hardware accelerator for LLMs, but I thought it would be like google TPUs where you can load up whatever model you want. Baking the model into hardware sounds like the monkey paw curled, but it might be interesting selling an old.. MPU..? because it wasn't smart enough for your latest project
- sowbug 8mo agoThere's a scifi story here when millions of these chips, with Qwen8-AGI-Thinking baked into them, are obsoleted by the release of Qwen9-ASI, which promptly destroys humanity and then itself by accident. A few thousand years later, some of the Qwen8 chips in landfill somehow power back up again and rebuild civilization on Earth. Paging qntm...
- brainless 8mo agoI know it is not easy to see the benefits of small models easily but this is what I am building for (1). I created a product for Google Gemini 3 Hackathon and I used Gemini 3 Flash (2). I tested locally using Ministral 3B and it was promising. Definitely will need work. But 8B/14B may give awesome results. I am building a data extraction software on top of emails, attachments, cloud/local files. I use a reverse template generation with only variable translation done by LLMs (3). Small models are awesome for this (4). I just applied for API access. If privacy policies are a fit, I would love to enable this for MVP launch. 1. https://github.com/brainless/dwata https://github.com/brainless/dwata 2. https://youtu.be/Uhs6SK4rocU https://youtu.be/Uhs6SK4rocU 3. https://github.com/brainless/dwata/tree/feature/reverse-template-based-financial-data-extraction/dwata-agents/src/bin https://github.com/brainless/dwata/tree/feature/reverse-temp... 4. https://github.com/brainless/dwata/tree/feature/reverse-template-based-financial-data-extraction/dwata-agents/src/template_financial_extractor/prompts https://github.com/brainless/dwata/tree/feature/reverse-temp...
- kamranjon 8mo agoIt would be pretty incredible if they could host an embedding model on this same hardware, I would pay for that immediately. It would change the type of things you could build by enabling on the fly embeddings with negligible latency.
- tgsovlerkhgsel 8mo agoTheir "chat jimmy" demo sure is fast, but it's not useful at all. Test prompt: ``` Please classify the sentiment of this post as "positive", "neutral" or "negative": Given the price, I expected very little from this case, and I was 100% right. ``` Jimmy: Neutral. I tried various other examples that I had successfully "solved" with very early LLMs and the results were similarly bad.
- weli 8mo agoMaybe its the tism but I also read that sentence as neutral. You expected very little and you got very little. Why would that be positive or negative? Maybe it should be positive because you got what you were expecting? But I would call getting what you expect something neutral, if you expected little and got a lot then that would be positive. If you expected a lot and got little then its negative. But if you expected little and got little the most clear outcome is that its a neutral statement. Am I missing something?
- deleted 8mo ago[deleted]
- rhodey 8mo agoI wanted to try the demo so I found the link > Write me 10 sentences about your favorite Subway sandwich Click button Instant! It was so fast I started laughing. This kind of speed will really, really change things
- segmondy 8mo agoPretty cool, what they need is to build a tool that can take any model to chip in short a time as possible. How quick can they give me DeepSeek, Kimi, Qwen or GLM on a chip? I'll take 5k tk/sec for those!
- throwaw12 8mo agoalso imagine it will cost 300$/unit, we all will host our own set of models locally, dream dream
- standeven 8mo agoHoly shit this is fast. It generated a legible, original, two-paragraph story on given topics in 0.025s.
- troyvit 8mo agoSo they create a new chip for every model they want to support, is that right? Looking at that from 2026, when new large models are coming out every week, that seems troubling, but that's also a surface take. As many people here know better than I that a lot of the new models the big guys release are just incremental changes with little optimization going into how they're used, maybe there's plenty of room for a model-as-hardware model. Which brings me to my second thing. We mostly pitch the AI wars as OpenAI vs Meta vs Claude vs Google vs etc. But another take is the war between open, locally run models and SaaS models, which really is about the war for general computing. Maybe a business model like this is a great tool to help keep general computing in the fight.
- g-mork 8mo agoOne of these things, however old, coupled with robust tool calling is a chip that could remain useful for decades. Baking in incremental updates of world knowledge isn't all that useful. It's kinda horrifying if you think about it, this chip among other things contains knowledge of Donald Trump encoded in silicon. I think this is a way cooler legacy for Melania than the movie haha.
- gordonhart 8mo agoWe’re reaching a saturation threshold where older models are good enough for many tasks, certainly at 100x faster inference speeds. Llama3.1 8B might be a little too old to be directly useful for e.g. coding but it certainly gets the gears turning about what you could do with one Opus orchestrator and a few of these blazing fast minions to spit out boilerplate…
- nickpsecurity 8mo agoMy concept was to do this with two pieces: 1. Generic, mask layers and board to handle what's common across models. Especially memory and interface. 2. Specific layers for the model implementation. Masks are the most expensive part of ASIC design. So, keeping the custom part small with the rest pre-proven in silicon, even shared across companies, would drop the costs significantly. This is already done in hardware industry in many ways but not model acceleration. Then, do 8B, 30-40B, 70B, and 405B models in hardware. Make sure they're RLHF-tuned well since changes will be impossible or limited. Prompts will drive most useful functionality. Keep cranking out chips. There's maybe a chance to keep the weights changeable on-chip but it should still be useful if only inputs can change. The other concept is to use analog, neural networks with the analog layers on older, cheaper nodes. We only have to customize that per model. The rest is pre-built digital with standard interfaces on a modern node. Given the chips would be distributed, one might get away with 28nm for the shared part and develop it eith shuttle runs.
- OrvalWintermute 8mo agowow that is fast!
- jtr1 8mo agoThe demo was so fast it highlighted a UX component of LLMs I hadn’t considered before: there’s such a thing as too fast, at least in the chatbot context. The demo answered with a page of text so fast I had to scroll up every time to see where it started. It completely broke the illusion of conversation where I can usually interrupt if we’re headed in the wrong direction. At least in some contexts, it may become useful to artificially slow down the delivery of output or somehow tune it to the reader’s speed based on how quickly they reply. TTS probably does this naturally, but for text based interactions, still a thing to think about.
- Alifatisk 8mo agoWhat's happening in the comment section? How come so many cannot understand that his is running Llama 3.1 8B? Why are people judging its accuracy? It's almost a 2 years old 8B param model, why are people expecting to see Opus level response!? The focus here should be on the custom hardware they are producing and its performance, that is whats impressive. Imagine putting GLM-5 on this, that'd be insane. This reminds me a lot of when I tried the Mercury coder model by Inceptionlabs, they are creating something called a dLLM which is like a diffusion based llm. The speed is still impressive when playing aroun with it sometimes. But this, this is something else, it's almost unbelievable. As soon as I hit the enter key, the response appears, it feels instant. I am also curious about Taalas pricing. > Taalas’ silicon Llama achieves 17K tokens/sec per user, nearly 10X faster than the current state of the art, while costing 20X less to build, and consuming 10X less power. Do we have an idea of how much a unit / inference / api will cost? Also, considering how fast people switch models to keep up with the pace. Is there really a potential market for hardware designed for one model only? What will they do when they want to upgrade to a better version? Throw the current hardware and buy another one? Shouldn't there be a more flexible way? Maybe only having to switch the chip on top like how people upgrade CPUs. I don't know, just thinking out loudly.
- test001only 8mo agoThat is my concern too. A chip optimised for a model or specific model architecture will not be useful for long.
- ahofmann 8mo agoI just tried the demo and I think, this is huge! If they manage to build a chip in 2 or 3 years, that can run something like Opus 4.6 or even Sonnet, at that speed, the disruption in the world of software development will be more than we saw in the last 3-5 years. LLMs today are somewhat useful, but they are still too slow and expensive for a meaningful ralph loop. Being able to runs those loops (or if you want to call it "thinking") much faster, will enable a lot of stuff, that is not feasible today. Writing things like openclaw will not take weeks, but hours. Maybe even rewriting entire tools, kernels or OSes will be feasible because the LLM can run through almost endless tries. Speed and cost wins over quality and this will also be true for LLMs.
- ilc 8mo agoMinor note to anyone from taalas: The background on your site genuinely made me wonder what was wrong with my monitor.
- deleted 8mo ago[deleted]
- deleted 8mo ago[deleted]
- armishra 8mo agoI am extremely impressed by their inference speed!
- bmc7505 8mo ago17k TPS is slow compared to other probabilistic models. It was possible to hit ~10-20 million TPS decades ago with n-gram and PDFA models, without custom silicon. A more informative KPI would be Pass@k on a downstream reasoning task - for many such benchmarks, increasing token throughput by several orders of magnitude does not even move the needle on sample efficiency.
- mlboss 8mo agoInference is crazy fast! I can see lot of potential for this kind of chip for IOT devices and Robotics.
- flux3125 8mo agoI imagine how advantageous it would be to have something like llama.cpp encoded on a chip instead, allowing us to run more than a single model. It would be slower than Jimmy, for sure, but depending on the speed, it could be an acceptable trade-off.
- coppsilgold 8mo agoPerformance like that may open the door to the strategy of brutefocing solutions to problems for which you have a verifier (problems such as decompilation).
- Aerroon 8mo agoImagine this thing for autocomplete. I'm not sure how good llama 3.1 8b is for that, but it should work, right? Autocomplete models don't have to be very big, but they gotta be fast.
- PeterStuer 8mo agoNot sure, but is this just ASICs for a particular model release?
- maelito 8mo agoTalks about ubiquitous AI but can't make a blog post readable for humans :/
- b0rbb 8mo agoThat animated background is terrible. Incredibly distracting. No way to turn it off (at least within what's provided without using something like devtools.)
- petesergeant 8mo agoFuture is these as small, swappable bits of SD-card sized hardware that you stick into your devices.
- arjie 8mo agoThis is incredible. With this speed I can use LLMs in a lot of pre-filtering etc. tasks. As a trivial example, I have a personal OpenClaw-like bot that I use to do a bunch of things. Some of the things just require it to do trivial tool-calling and tell me what's up. Things like skill or tool pre-filtering become a lot more feasible if they're always done. Anyway, I imagine these are incredibly expensive, but if they ever sell them with Linux drivers and slotting into a standard PCIe it would be absolutely sick. At 3 kW that seems unlikely, but for that kind of speed I bet I could find space in my cabinet and just rip it. I just can't justify $300k, you know.
- deleted 8mo ago[deleted]
- garganzol 8mo agoImagine a mass-produced AI chips with all human knowledge packed in chinesium epoxy blobs running from CR2032 batteries in toys for children. Given the progress in density and power consumption, it's not that far away.
- AlexC04 8mo agoIf I could have one of these cards in my own computer do you think it would be possible to replace claude code? 1. Assume It's running a better model, even a dedicated coding model. High scoring but obviously not opus 4.5 2. Instead of the standard send-receive paradigm we set up a pipeline of agents, each of whom parses the output of the previous. At 17k/tps running locally, you could effectively spin up tasks like "you are an agent who adds semicolons to the end of the line in javascript", with some sort of dedicated software in the style of claude code you could load an array of 20 agents each with a role to play in improving outpus. take user input and gather context from codebase -> rewrite what you think the human asked you in the form of an LLM-optimized instructional prompt -> examine the prompt for uncertainties and gaps in your understanding or ability to execute -> <assume more steps as relevant> -> execute the work Could you effectively set up something that is configurable to the individual developer - a folder of system prompts that every request loops through? Do you really need the best model if you can pass your responses through a medium tier model that engages in rapid self improvement 30 times in a row before your claude server has returned its first shot response?
- AmazingTurtle 8mo agoModels can't improve themselves with their own (model) input, they need to be grounded in truth and reality.
- rockostrich 8mo agoBut at one point the model is sufficiently large enough to accomplish any task a human could specify. For software development, I think we're pretty much at that point with the latest Anthropic/Google/OpenAI models. We have no idea where the direction of token pricing is going to go in the future, but the consensus seems to be that it will only get more expensive. If Taalas can offer the same functionality that we have with frontier models today at a 1/10 of the cost and 10x the speed then they're going to take over a large part of the market.
- dalenw 8mo agoI think so. The last few months have shown us that it isn't necessarily the models themselves that provide good results, but the tooling / harness around it. Codex, Opus, GLM 5, Kimi 2.5, etc. all each have their quirks. Use a harness like opencode and give the model the right amount of context, they'll all perform well and you'll get a correct answer every time. So in my opinion, in a scenario like this where the token output is near instant but you're running a lower tier model, good tooling can overcome the differences between a frontier cloud model.
- heliumtera 8mo agoYep, this is the most exciting demo for me yet. Holy cow this is unbelievably fast. The most impressive demo since gpt 3, honestly. Since we already have open source models that are plenty good, like the new kimi k2.5, all I need is the ability to run it at moderate speed. Honestly I am not bullish on capabilities that models do not yet have, seems we have seen it all and the only advancement have been context size. And honestly I would claim this is the market sentiment aswell, anthropic showed opus 4.6 first and the big release was actually sonnet, the model people would use routinely. Nobody gave a shit about Gemini 3.1 pro, 3.0 flash was very successful... Given all the recent developments in the last 12 months, no new use cases have opened for me. Given this insane speed, even on a limited model/context size, we would approach IA very differently.
- DeathArrow 8mo agoIs amazingly fast but since the model is quantized and pretty limited, I don't know what it is useful for.
- luyu_wu 8mo agoI think this is quite interesting for local AI applications. As this technology basically scales with parameter size, if there could be some ASIC for a QWen 0.5B or Google 0.3B model thrown onto a laptop motherboard it'd be very interesting. Obviously not for any hard applications, but for significantly better autocorrect, local next word predictions, file indexing (tagging I suppose). The efficiency of such a small model should theoretically be great!
- Tehnix 8mo agoBunch of negative sentiment in here, but I think this is pretty huge. There are quite a lot of applications where latency is a bigger requirement than the complexity of needing the latest model out there. Anywhere you'd wanna turn something qualitative into something quantitative but not make it painfully obvious to a user that you're running an LLM to do this transformation. As an example, we've been experimenting with letting users search free form text, and using LLMs to turn that into a structured search fitting our setup. The latency on the response from any existing model simply kills this, its too high to be used for something where users are at most used to the delay of a network request + very little. There are plenty of other usecases like this where.
- llsf 8mo agoThat is what self-driving car should eventually use, whenever they (or the authorities) deem their model good enough. Burn it on a dedicated chip. It would be cheaper (energy) to run, and faster to make decisions.
- wmf 8mo agoIt's more expensive in COGS and self driving doesn't need to run at 900 Hz.
- brcmthrowaway 8mo agoWhat happened to Beff Jezos AI Chip?
- gen220 8mo agoThis is genuinely an incredible proof-of-concept; the business implications of this demo to the AI labs and all the companies that derive a ton of profit from inference is difficult to understate, really. I think this is how I'm going to get my dream of Opus 3.7 running locally, quickly and cheaply on my mid-tier MacBook in 2030. Amazing. Anthropic et al will be able to make marginal revenue from licensing the weights of their frontier-minus-minus models to these folks.
- g-mork 8mo agoI do like the idea of an aftermarket of ancient LLM chips that still have tons of useful life on text processing tasks etc. They don't talk about their architecture much, I wonder how well power can scale down. 200W for such a small model is not something I see happening in a laptop any time soon. Pretty hilarious implications for moat-building of the big providers too.
- gen220 8mo agoYea I mean this is the first publishable draft of a startup cooking on this. I'm confident there are at least 1-2 OOMs of improvement to come here in terms of the (intelligence : wattage) ratio. I really thought we were going to need to see a couple of dramatic OOM-improvement changes to the model composition / software layer, in order to get models of Opus 3.7's capability running on our laptops. This release tells me that eventual breakthrough won't even be strictly necessary, imo.
- g-mork 8mo agoThe way I imagine it in 2-4 years we're going to be hit with a triple glut of better architecture, massive oversupply of hardware and potentially one or two hardware efforts like this really taking off. It's pretty crazy we're already 4 years in and outside of very niche / low availability solutions, it's still either GPU or bust
- gen220 8mo ago
- max8539 8mo agoThis is crazy! These chips could make high-reasoning models run so fast that they could generate lots of solution variants and automatically choose the best. Or you could have a smart chip in your home lab and run local models - fast, without needing a lot of expensive hardware or electricity
- mbh159 8mo agoSo cool, what's underappreciated imo: 17k tokens/sec doesn't just change deployment economics. It changes what evaluation means, static MMLU-style tests were designed around human-paced interaction. At this throughput you can run tens of thousands of adversarial agent interactions in the time a standard benchmark takes. Speed doesn't make static evals better it makes them even more obviously inadequate.
- jameslk 8mo agoThe implications for RLM is really interesting. RLM is expensive because of token economics. But when tokens are so cheap and fast to generate, context size of the model matters a lot less Also interesting implications for optimization-driven frameworks like DSPy. If you have an eval loop and useful reward function, you can iterate to the best possible response every time and ignore the cost of each attempt
- CrzyLngPwd 8mo agoThe demo is dogshit: https://chatjimmy.ai/ https://chatjimmy.ai/ I asked it some basic questions and it fudged it like it was chatgpt 1.0
- TheServitor 8mo agoI don't know the use of this yet but I'm certain there will be one.
- Animats 8mo agoEmbedding the model at chip fab time ought to be useful for robotics, driving, vision, and audio applications, at least. The training sets are good for years. So they use 3 bit values. Is that current thinking? LLMs started at 32-bit floats, and have gradually shrunk. 8-bit floats seem to work. Is 3 bits pushing it?
- amelius 8mo agoIf you're making your own chip, you might as well explore analog computation.
- konaraddi 8mo ago> Taalas’ silicon Llama achieves 17K tokens/sec per user, nearly 10X faster than the current state of the art, while costing 20X less to build, and consuming 10X less power. Insane gains, makes me excited for the future. Imagine Opus-like responses in <1 second. I suspect power efficiency will be nearly entirely offset by increased usage but it’s more bang for watt.
- HenryOsborn 8mo agoToken velocity is great, but the industry is hyper-fixated on speed while completely ignoring the blast radius. If we push to 17k tokens/sec for autonomous agents, we are just accelerating how fast an agent can hit an infinite loop and drain an API budget. Before we make AI ubiquitous, we need deterministic, network-level circuit breakers. Speed without governance is just a faster way to burn capital.
- tgtweak 8mo agoI think frontier models can do more with fewer tokens (and do the wrong thing far less often) than a "really fast" small model. There are use cases for fast/ultrafast inferrence models - classifying text, scoring things, extracting information - but for coding and other knowledge tasks - you're not going to get to your solution faster at 16,000 tokens/s if the solution never comes (or is the wrong one).
- sheepscreek 8mo agoWe need one of these things running an OSS vision model. Having super-fast agentic computer access would be so worthwhile!
- mncharity 8mo agoThere's an old idea of adaptive media. Imagine a video drama that's composed of a graph of clips, like an old "choose your own adventure" book ("Do you X? If yes, goto page 45"). With gaze tracking, one can "hmm, the viewer is more focused on character A than B... so we'll give clips and subplots with more A". Now, when reading, the eye moves in little jumps - saccades. They last 10's of ms, the eye is blind during them, and with high-quality tracking, you know quite early just where that foveal peephole is going to land. So handwave a budget of a few ms for trajectory analysis, a few for 200 Hz rendering latency, and you still have 10-ish ms to play with. At 20k tok/s, that's 200 tok. So perhaps one might JIT the next sentence, or the topic of the next paragraph, or the entire nature of the document, based on the user's attention. Imagine a universal document - you start reading, and you find the document is about, whatever you wanted it to be about?
- awwaiid 8mo agoGenerative TikTok for words
- mncharity 8mo agoHmm... TikTok has apparently long had "text enhanced with background" genres, and TIL, text posts since 2023. So text is ok. But non-independent items? For generative storytelling, "here is a next paragraph for the story", swipe left/right might work? Want to avoid "I don't much like this new paragraph, but I'm afraid to lose it and be stuck with something worse". Swipe left/right and up for continue? Swipe down to revisit old choices? Maybe present new text bolded, appended to old text, for context. Or a "next page of a picture book" idiom. A text field for direct creative or editorial intervention - speech to text. Maybe a side channel input for "story and background should now be soporific". Generative bedtime stories, but incrementally collaboratively created... Thanks for the brainstorming prompt.
- trollbridge 8mo agoThis is, so far, utterly charming. I made a simple prompt of "make an adventure game in the style of cia.bas from pc-sig". It ended up being wildly different than that, but 30 minutes later and I'm still busy trying to play this "game" it fabricated out of thin air. One interesting thing is it keeps randomly emitting "ประก" (meaning "Announcement") and chartInstance. This is recalling the early days of GPT-2 when the light bulb went on that "hey, there's something groundbreaking here".
- oofbey 8mo agoI have a hard time reading beyond factual lies like: > On the cost front, deploying modern models demands massive engineering and capital: room-sized supercomputers consuming hundreds of kilowatts… This is just wrong. The largest models are probably 1-2 trillion parameters. Say 2 trillion and let’s pretend it’s only quantized to 8bit (even though it could easily be half that.) So we need 2TB of VRAM. Not even using the latest hardware, lets say H100 chips with 80GB of vram each, with 8 of them in say an 8U. (Although you can certainly fit these in 6U still air cooled or even 4U water cooled.) Three of these server would almost do, but let’s call it four to include plenty of room for context. The largest physical size would be 32U - most of a single rack. Which is hardly the size of a room, even in Manhattan. Total power maybe 40kW. And you could easily drop these numbers to a half or quarter of that with reasonable modifications or upgrades. If you want to sell your hardware, start by being honest about the problem you’re addressing.
- rajbiswas125 8mo agohe numbers being presented are deliberately misleading. On this model, Groq delivers around 1,300 tokens per second, whereas Cerebras achieves roughly 2,500 tokens per second. With the next generation of Cerebras chips expected to be 5–7× faster, peak throughput could reach the ~17,500 tokens-per-second range. For smaller models like this, that level of performance is entirely realistic. So no, a general-purpose accelerator will likely continue to outperform a fixed-function ASIC with a specific model etched into it. Moreover, we’re only looking at results from a two-year-old, relatively small model. We still don’t know how this architecture will scale with a large MoE model, especially given constraints like limited on-chip KV cache and more complex attention mechanisms. The real test isn’t performance on a small benchmark model, it’s how the system handles large-scale, production-grade workloads under architectural constraints.
- iriisatremotely 8mo ago[dead]
- trippyballs 8mo agoholy fuck this is really gud. imagine this with sota models. we are cooked af.damn
- runeks 8mo ago> Taalas’ silicon Llama achieves 17K tokens/sec per user, nearly 10X faster than the current state of the art, while costing 20X less to build, and consuming 10X less power. Am I reading this right: 10x faster and 10x less power, ie. 100x more power efficient?
- noisy_boy 8mo agoI'm curious how much of "hardcoding" is in the chip? Can it have parts that don't need changing much and "offload" the rest into some sort of high-speed/bandwidth interconnect? Will we reach a state where we have chips on which models can be "flashed" like CPU firmware? Or eventually will we reach a state where none of these tricks will be needed because like run-of-the mill Intel/AMD commodity CPUs, we will have full-power AI chips which will be part of an bigger/integrated mother-chip? Then what will happen to companies that do LLMs-as-a-service? Will they be forced to join and adapt becoming hybrid model+hardware shops? I'm not knowledgeable enough about hardware but throwing these random ideas out in hopes of thought-provoking responses to learn from.
- rockostrich 8mo agoThese chips are large by fab standards and even with state of the art processes we likely won't see any kind of integration on consumer tech any time soon, but I imagine they will absolutely see instant demand if they can deliver on what they laid out in the post.
- Simboo 8mo agoDeep Differentiable Logic Gate Networks?
- doix 8mo agoI was wondering if/when this would happen. My friends and I would discuss this at the pub all the time, "LLM2RTL" or take it a step further and do the the whole process "LLM2GDS". I couldn't find much info here, but I'm guessing they've built tooling to automatically convert model weights to RTL and the reason it's such an old model is that it takes a long time tape a chip out (especially the first one). Would be interesting to know how much is automated and how much is handcrafted. I think the "next big thing" with AI hardware will be when they switch from "digital" implementations of LLMs to "analogue". We already know that we can lose some bits of precision and still have a "workable" model. If/when folks figure the fine-tuning out, I'm guessing it'll be another order of magnitude improvement.
- ttul 8mo agoThe NextPlatform article hints at their approach: “We have got this scheme for the mask ROM recall fabric – the hard-wired part – where we can store four bits away and do the multiply related to it – everything – with a SINGLE TRANSISTOR. So the density is basically insane. And this is not nuclear physics – it is fully digital. It is just a clever trick that we don’t want to broadcast. But once you hardwire everything, you get this opportunity to stuff very differently than if you have to deal with changing things. The important thing is that we can put a weight and do the multiply associated with it all in one transistor. And you know the multipliers are kind of the big boy piece of the computer.“ One transistor doing 4-bit multiplication? A plausible way to get “4-bit weight plus multiply in one transistor” in a 6 nm FinFET mask-ROM fabric is to make the ROM cell a single device whose drive strength is the stored value. At tapeout you pick one of about 16 discrete strengths (for example by choosing fin count and possibly Vt), so that transistor itself encodes a 4-bit weight. Then you do the multiply in the charge/time domain by encoding the input activation as a discrete pulse width or pulse count and letting the cell source or sink a weight-proportional current onto a precharged bitline for that duration. The resulting bitline voltage change (or time-to-threshold) is proportional to current times time, so it behaves like weight times input and can be accumulated along a column before a simple comparator or time-to-digital readout. It’s “digital” in the sense that both weight and input are quantized, but it relies on device physics; the hard parts are keeping 16 levels separable across PVT, mismatch, and aging, plus managing bitline noise and coupling and ensuring the device stays in a predictable operating region. VLSI design produces digital outputs, but in the quantum silicon domain, it’s all about the analog…
- Barbing 8mo agoTIL your salary (Kiddin’, my silly way to say thanks for a deeply technical look, helps me understand the kind of knowledge work that might be useful n years from now!)
- ttul 8mo agoI stand corrected here. The design is "fully digital". They are not using analog trickery here. They are most likely just using a clever collection of very tiny transistors hooked up to two wire layers such that one _digital_ transistor can trigger the correct 4-bit output. This can be accomplished through clever logic design.
- creativeSlumber 8mo agoAre the model weights burned into the silicon / part of the architecture? Or can you update the model weights on these chips? If they cannot be updated, these chips will be outdated the moment they are made given the breakneck speed at which new and improved models are introduced.
- TheServitor 8mo agoApplied for access. I hope to test a parallel fast-inference problem solver with a hybrid MCTS approach.
- snowhale 8mo ago[dead]
- bigcat12345678 8mo agoI am thinking if this can be a low-level substrate for composing dumb LLMs into smart swarm, theoretically: 1. A whole with disparate parts (smart and dumb components) are almost always more cost-effective to reach a given target of performance 2. With that, a whole with disparate parts, are almost always more performant with the same cost A few inspiration: 1. Human body is intelligent composed of so diverse parts 2. Swarm intelligence of insects and small animals are certainly beyond current understanding The cost and speed of this thing is on point to make such a whole composed diverse parts possible.