3 ms·
8800 GTX in 2006. Cutting-edge, an insanely powered consumer card for the time. Theoretically around 0.3456 TFLOPS. 1080 GTX in 2016. Cutting-edge, an insanely
by apimade 2mo ago
8800 GTX in 2006. Cutting-edge, an insanely powered consumer card for the time. Theoretically around 0.3456 TFLOPS.
1080 GTX in 2016. Cutting-edge, an insanely powerful consumer card for the time. Theoretically around 8.87 to 8.9 TFLOPS.
5090 RTX in 2026. Cutting-edge, an insanely powerful consumer card for today.
Theoretically around 104.8 TFLOPS.
In the same timeframe mobile processor CPU's went from 0.001 TFLOPS, to today's Apple's A19 Pro chip which delivers 2.074 TFLOPS.
That's _without_ getting into ASIC's, or purpose-built hardware like Taalas's model on silicon HC1, or generic AI dies like what they're planning with HC2 or Cerebras, which will massively compress the timeline.
- cududa 2mo agoJust a note that I think the direction most people are paying attention to is memory bandwidth; thats the real bottleneck and “number go up” but also constraint people are designing around
- xbmcuser 2mo agoWe will see such power and price now only when AI market crashes or China reaches node parity and goes after market share as currently the way they are buying out most of the latest node production the consumer prices will only be palatable to the very rich or we will need to be happy with older slower nodes
- formerly_proven 2mo agoExcept the 499$ of a 1080 GTX inflation-adjusted only buys you a 5070 or 5070 Ti even by MSRP.
- zmmmmm 2mo agoSadly while the FLOPS are increasing nicely, total graphics memory is stalled in consumer cards by comparison.
- jack_pp 2mo agoIsn't there such a thing as low hanging fruit? Aren't we already approaching theoretical physical limits? We're at 2nm
- kaashif 2mo ago(1) Yes. (2) Are you saying that you think we're at the limits of computing in general, or that specific technology? We know, for example, that a human brain level intelligence is possible to run on a human brain. We are nowhere near that. And actually that's not even a physical limit necessarily. But that is...not a low hanging fruit.
- cvak 2mo agowe are not at 2nm, we just call it that.
- root_axis 2mo agoOk, now do memory capacity and bandwidth - the things that actually constraint local LLMs.
- apimade 2mo ago8800 GTX in 2006: 768 MB of GDDR3, with 86.4 GB/s of theoretical memory bandwidth. GTX 1080 in 2016: 8 GB of GDDR5X, with 320 GB/s. RTX 5090 in 2026: 32 GB of GDDR7, with 1.792 TB/s. This is fun, what's next?! PCI 8.0 is breaking 1TB/s, GDDR7 is 1TB/s. With just the _current_ timeline, things are looking like they'll compress once we get over this initial lump.
- flaburgan 2mo agoYeah but here you describing the opposite phenomenon. You're saying that the hardware is going to become cheaper and more powerful with the years, to the point a current State of the Art model from today will run on a normal consumer hardware in ten years. What people are trying to do now is the opposite, optimize the software as much as possible so that it does not need the best hardware but the normal one we currently have. As if we were trying to make a current AAA game to run smoothly on the 1080 GTX of your example.
- dtj1123 2mo agoNo, both of these things can happen in parallel. The suggestion is that a 1T model could be made to run on cheap consumer hardware of the future.
- foxrider 2mo agoSpeaking of ASICs - how likely is it that as models get better we'll see someone baking a whole model directly into the silicon? It's like having l0 cache.
- SJC_Hacker 2mo agoYou could do it but there would be no point, The only advantage over would be power consumption. And it would be quite expensive. At the rate models are improving, it would be obsolete in six months.
- HPsquared 2mo agoPower consumption and latency are very important on mobile
- naasking 2mo agoThey're important everywhere of course, but especially on mobile. If AI researchers figure out how to offload knowledge and expertise from reasoning weights, then a core reasoning ASIC linked to the knowledge would totally rock.
- foxrider 2mo agoYes, right now it would be obsolete in six months, but I also must add that this never stopped crypto miners from making new ASICs. However, with how useful Kimi is right now - at some point if someone makes a dedicated hardware board with "good enough" model for daily tasks - that would be a very sought after commodity.
- kaelwd 2mo agoOnly 8B currently but it's been done: https://taalas.com/products/ https://taalas.com/products/
- apimade 2mo agoThe only _public_ example we know. This is definitely being done with private models by HFT/quant firms, data processing agencies/orgs (large intelligence agencies, _every_ data analytics org, etc).
- Azantys 2mo agoHere you have shown yourself that progress slows down and doesnt speed up. 8.9/0.35 = ~25x more performance in 10 years from 2006 to 2016. 104.8/8.9 = ~12x more performance in 10 years from 2016 to 2026. Growth has dropped 50%.
- dghlsakjg 2mo agoThat isn't deceleration, you've just chosen a very selective way to compare. If you use time as a denominator, which is kind of intrinsic when talking about rates of acceleration, you get a very different result. If you graphed .3, 8.9, and 104.4 on the y axis, with years on the x axis, it would be pretty clear that there was in increase in the rate of progress. We went from adding 8 teraflops in a decade, to adding almost 100 the next decade. If we add "only" 400 more teraflops in the next decade the graph will make that initial growth look flat in comparison, even though your math would show that we are basically stalled out. It’s like claiming that a company that goes from making $1 to $1k to $100k to $1mm in a 4 year period has decelerating growth.
- ksec 2mo ago>It’s like claiming that a company that goes from making $1 to $1k to $100k to $1mm in a 4 year period has decelerating growth. Because it is decelerating growth. There is a reason why we use YoY percentage in annual and financial reporting.