4 ms·
When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X
by jpadkins 2mo ago
When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon and we will see another 10X improvement in price/performance frontier.
The dynamic GPU clusters will be used for the 5% of tasks, and pushing out the frontier. Also there will be a set of knowledge tasks that are not done today (because they are too difficult for most knowledge workers), that will start being done in the future.
- jrflo 2mo agoBurning the weights into silicon would be many orders of magnitude increase, not just 10x. It's kind of crazy that this hockey stick the AI hype bros talk about seems more and more every day like it might be real
- jaggederest 2mo agohttps://taalas.com/ https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second. https://chatjimmy.ai/ https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
- iamjackg 2mo agoHoly crap, I was not prepared for how fast it responded. I just wrote "Just wanted to see how fast you are! Can you write me a quick story about a tiger who lives inside a block of cheese the size of a house?" I pressed Enter, and the response was instant. > Generated in 0.037s • 14,205 tok/s This is unbelievable.
- HDThoreaun 2mo agoFor what its worth the frontier lab models can surely be a lot faster if they wanted them to be but theyre supply constrained so theyre doing stuff like multi tenancy. Since you cant self host them no one outside the labs really knows speed as a solo tenant
- ericd 2mo agoYou can kind of get a sense by running these things at home - I'm currently running Laguna. One interesting thing is that per stream doesn't actually slow down that much with multiple concurrents, because the bottleneck remains the memory bandwidth until pretty significant request depths, and then eventually you hit the GPU's limits. It's one of the big forces that pushes for centralization in this stuff, the fixed costs to run one are huge, the marginal costs of additional tenants, relatively small.
- jcul 2mo agoIt's crazy. Are they doing any precomputing as you type, I wonder if you paste a block of text is it the same speed.
- nerdsniper 2mo agoI pasted and instantly hit enter on this prompt: "I generated a filter set using REW v5.31.3 using real-world sweep tone measurements from the room I'm listening in . How can I use it as my MacOS output equalizer so that my spotify music is adjusted for this room and speakers" and it gave a very reasonable answer in non-perceptible time.
- jaggederest 2mo agoThat's a hell of a prompt, you should make a post about what you're doing maybe, because I want that for myself.
- weiliddat 2mo agoidk if you're still actually looking for a solution (I got nerdsniped heh), but I found a seamless way to have it locked into the speakers + room, instead of a specific source/software, is using an iLoud subwoofer that routes to any speaker setup and has room correction + EQ that lives on the sub.
- throwuxiytayq 2mo agoNo, it really does take ~0.03s to generate the answer. Try your browser's developer tools and watch the requests.
- 8cvor6j844qw_d6 2mo agoI'd like to imagine the things that can be done with this speed and the current frontier models.
- Dig1t 2mo agoSeriously, if Fable or even Opus was this fast that would be a real game changer.
- reducesuffering 2mo agoRSI will be models better than Fable running faster than this, you won't even need a human in the loop to figure out what to do. The high level goal will be accomplished better than the human in an instant
- ElijahLynn 2mo agoTruth, it feels like we're in the dial-up age of LLMs right now. And this Jimmy AI is fiber.
- throwup238 2mo agoFully interactive games where you can talk to every NPC by text or voice and have an LLM drive the story (with your own meta prompts to guide it, if you so wish). Maybe even have them generate assets on the fly too. I’m still trying to figure out coding agents. I can’t even begin to imagine the things it would enable. Even the most mundane ideas like LLMs-in-HiFreq-trading have huge implications.
- nly 2mo ago0.03 seconds is an eternity in high frequency trading You need to be 4 orders of magnitude faster at least
- baal80spam 2mo agoJust wow, I made a similar request. Result: Generated in 0.042s • 14,201 tok/s This is crazy.
- kooi 2mo ago"Stochastic gradient descent algorithm in Haskel" "LMS algorithm in bash" Just barfed it up lol. Amazing.
- ElijahLynn 2mo agoWow! You weren't kidding, I just tried it too and 14,098 tokens in .05 seconds, I barely blinked and it was done. There was no typing at all appearing on the screen. It just showed up. https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37cea https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37...
- WalterGR 2mo agoThat link isn't bringing up your chat, FYI. It just shows the default new chat state.
- ElijahLynn 2mo agoSure enough. I just tested it in incognito, and it is blank. I thought it would work because I wasn't logged in to Jimmy AI. It still shows for me though in my regular session so they must just have a cookie that makes it stick. Here's a copy and paste prompt if somebody wants to just test it real quick to see what I saw: Write a story about the fastest monkey who ever lived, his name is Jimmy and he is an AI superbot monkey that is part cyborg primate. He can travel through time and is psychic.
- christophilus 2mo agoWow. This is absolutely wild. I didn't expect that. If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.
- apitman 2mo agoSpeed is the metric I'm currently most interested in. The models are smart enough. Once speed significantly increases I think we're going to see some interesting downstream effects. The three things I currently spend the most time waiting on are LLM API requests, Rust compile times, and nix derivations. As AI latency approaches zero I think we're going to start taking a hard look at whether slow-compiling languages are adding enough value over Golang, Typescript, or even dynamic languages to be worth the slowdown.
- blovescoffee 2mo agowhat do you mean "if"? of course we will, and the models will be smarter as well
- jaggederest 2mo agoYeah. A model smarter than Fable by, say, 50 human IQ point equivalents, running in some kind of autoregressive process continously on whatever goals you put into the loop, on baked silicon, would be maybe my low end for the potential in the relatively near future.
- apitman 2mo agoSee Cerebras and Groq as well.
- zahrevsky 2mo agohugged?
- antman 2mo agoAnswers like 8Bish model it appears, but instant response
- herzigma 2mo agoWow! Responded essentially instantaneously to my prompt: "I need a short, 4000 word essay on the the difference between star wars and Star Trek universes from the perspective of graduate level scientific work."
- Yopolo 2mo agoAnd don't underestimate how much money Google, Microsoft, Amazon and Meta still have to spend on this tech. Blocking Fable for sure made it very politicl a lot sooner than i expected it to happen. and because China already has massive problems of getting access, they are pushing it on hardware too like what Huawai did without EUV. It seems China is already able to do DUV a lot sooner than others expected.
- re-thc 2mo ago> It seems China is already able to do DUV a lot sooner than others expected. That's the media and in particular US KOLs of all sorts driving the wrong impression of China and other places. China and many other places for example have fast public transport that the US doesn't and can't even imagine today. They're not behind. China's DUV still isn't that production grade (mass produce-able) so don't get that hyped up the wrong way (in a different direction). The whole China-is-behind with tech and in particular semi wasn't that they can't. The truth is they spent decades in internal politics and corruption. That all got solved with the bans, so thank the bans! Jensen even said the bans were bad.
- coffeebeqn 2mo agoWhat does that mean though? Like some kind of a ROM memory ?
- bob1029 2mo agoStacked ROM can, in theory, be a lot denser than anything that depends on a capacitor and refresh cycle. I don't think it would be that difficult to manufacture compared to other process tech. HBM is really hard to do compared to other memory types.
- FuriouslyAdrift 2mo agoYep it will be ASICs and DSPs all over again. Orders of magnitude changes.
- monkeydust 2mo agoSo which shovels companies are the ones to watch for burnt in silicon models ?
- FuriouslyAdrift 2mo agoImagine the price of a $9 million NVL72 dropped to about $100, used 4 orders of magnitude less power, was the size of ARM cpu, be bundled with pretty much any electronic device, and ran as fast a frontier AI is today. That's about how disrupting DSPs were to the industries they arose out of (over a very long time frame). How would that disrupt the industry?
- HDBaseT 2mo ago$CBRS - Cerebras Systems There is other in the space, Groq and Sambanova are both private companies attempting to develop their own technology.
- FuriouslyAdrift 2mo agoCerebras and AMD collabrated on some Helios system stuff.
- vjvjvjvjghv 2mo agoNot an expert on this but wouldn’t this be possible with something similar to an FPGA?
- kooi 2mo agoWeights can be baked into silicone or programmed into hardware ala FPGA, but the context will always be dynamic. High speed SRAM is where the $$$ is
- jaggederest 2mo agoMy understanding is that for FPGA the issue is either it eats all your gates on internal memory if you interleave, or it takes forever to load everything between the SRAM on the board and the actual FPGA component over a bus, last time I looked into it.
- 5555watch 2mo agoI'm curious, how hard/expensive it is to burn a really large model into silicon, and why aren't we doing this already? Or, when we will start doing this, who's going to be able to do that in scale? I'm seeing the TAALAS example, but it's only an 8B model, suggesting some real limitations parameter wise. And for 2.5kW?
- adrianvi 2mo agoI can't speak for all cases, but the AI space is seeing improvements month by month, so it is beneficial to wait until it settles (a model becomes the standard in intelligence/price) before designing and mass producing an "LLM ASIC" of said model. The big AI labs won't do that unless they are forced to, as they want you to spend more money on the big, expensive, frontier models (so they can live up to their valuation), so it's more likely that you will see this on smaller open weights models.
- re-thc 2mo ago> When model intelligence reliably hits 90%-95% of current day knowledge worker tasks, they are going to burn those weight to silicon Google is already working on a similar idea but more "flexible".
- kridsdale1 2mo agoExplain.
- jmb99 2mo agoI'm not who you responded to and I don't have any info on Google. Nor can I explain in detail due to NDAs. But multiple major players are working on something along the lines of what the parent is alluding to. The "edge" AI landscape (in particular, what you can do with ~5W) is going to be nuts in about 18 months.
- jacekm 2mo agoHow will this affect the newly build data centers? What effect do you think it will have on memory prices?
- jmb99 2mo agoMy uneducated guess says, not much. For running massive models you still need a ton of high-bandwidth interconnects between many individual chips/GPUs/etc since you need to do math across a few TB worth of weights. That's simply going to require more power (and more die area in I/O, and therefore more cost). Being able to run small models in tiny power envelopes is incredibly useful to people, but I believe it will be covering a different niche than what datacenters can provide. Likewise, you'll still need crazy amounts of high-end memory to populate whatever goes in these datacenters. The only thing that will crash prices is reduced demand (duh) or, more interestingly, increased production. In particular, if CXMT is able to get their DDR5 fabs up to a reasonably high yield, that could add some downward price pressure (as could government subsidies). As well, if Micron/Kingston/Hynix think that CXMT is going to start cutting into their market share, they might be willing to either increases supply or drop prices. Unfortunately CXMT looks to be taking quite a while to get their new fab up to max capacity so that may take a year+ before anything manifests. If you're interested in following the (publicly available) info on these sorts of things, check out what companies like Axelera, DeepX, and MemoryX are doing today and have on their roadmaps, as well as the sorts of chips/SoCs Qualcomm, Kinara (now NXP), and Ambarella currently have announced (or have on the market). And remember, that pretty much all of these chips on the market today were in initial development more or less when ChatGPT first launched. If you knew what you knew today (or a year ago) about what requirements current- and next-generation models would have (from a silicon perspective), what might you do differently? Think for instance, host system interconnects, amount and speed of on-package or on-die memory, image/video decode capabilities, int8 vs fp8 vs fp16 vs bf16 compute units, etc. And, consider that most "AI" stuff in development a few years ago was all 15nm or 12nm - because who was gonna pay big money to get fab capacity at 3nm to run some object detection models? So most of the stuff on the market today is on very old nodes and therefore not super power efficient.