3 ms·
https://mikeveerman.github.io/tokenspeed/?rate=750&mode=think https://mikeveerman.github.io/tokenspeed/?rate=750&mode=thin... This is what 750tps looks like, I
by qznc 4mo ago
https://mikeveerman.github.io/tokenspeed/?rate=750&mode=think https://mikeveerman.github.io/tokenspeed/?rate=750&mode=thin...
This is what 750tps looks like, I guess.
- buddhistdude 4mo agoJust to think what this will look like in a couple of years.
- OGWhales 4mo agoHopefully like this (but smarter): https://chatjimmy.ai/ https://chatjimmy.ai/
- nomel 4mo agoThis is genuinely confusing to my senses. The future is going to be so strange/neat/me unemployed.
- falcor84 4mo ago> strange/neat/me unemployed I'm not sure if that's what you were going for, but I read it as if it were written by The Board in the game Control, and found myself with the appropriate level of existential dread.
- matheusmoreira 4mo agoThe future is totally illegible to me. I love these AI models, but I feel like I'm going to be jobless within 10 years. Anomie is at an all time high right now.
- the_af 4mo ago10 years? An optimist, I see.
- razodactyl 4mo agoYeah. It keeps catching me off guard that it answered me already.
- niyazpk 4mo agoWow.. what?! How is this so fast?! Where can I read more?
- dmd 4mo agohttps://taalas.com/ https://taalas.com/
- fcsp 4mo agoFunnily enough, pasting your comment straight into Jimmy leads to a... Funnily suboptimal answer that does not answer the question. As someone else already contributed, this is driven by a Canadian startup taalas that basically makes chips that are llms, so everything is very fast but also, baked into the chip. Once this kind of stuff is a commodity in like 10 years, our world will be very, very different.
- hajile 4mo agoTaalas HC1 AI uses Llama 3.1 8B, but takes up a massive 53B transistors and 815mm2 on TSMC N6 (nearly at the reticle limit of 858mm2). N2 is a little less than 3x as dense (110MTr/mm2 vs 313MTr/mm2). This chip would still be 272mm2 on N2 which is an eye-watering $30k/wafer and bigger than a 9950x or Nvidia 5070. This just isn't feasible. Some of the latest-gen LLMs seem to have 5-10T parameters or about 1000x more. I don't know that taping out just one chip makes economic sense let alone the 300-1000 chips required for a cutting-edge model. Things like continuing education so your model knows about the latest NPM packages or world news is super important, but seems like it would require new chips. There are a TON of uses for an 8B parameter models on the edge, but this is WAY too big to put on the edge of anything. Something like a 10mm2 100m parameter voice model might be feasible on the edge, but only for expensive devices, but most of those are TSMC 28nm (up to 29MTr/mm2) or GF FDX22 (up to 40MTR/mm2) which would increase the AI chip to the point where it would absolutely dominate the BOM.
- HaloZero 4mo agothe flash models have fallen in size at least between deep seek models. Is there a limit to the shrinking capacity of the models?
- plaguuuuuu 4mo ago[dead]
- kkotak 4mo agoWhy is the insane speed of 13KTPS of this site is not more on the the top of the AI conversations?
- Ey7NFZ3P0nzAe 4mo agoIt's pretty well known by now.
- chromadon 4mo agoI asked it for a block of C++ code and it hit 14,189 tok/s. I assume it cached someone else's session?
- fcsp 4mo agoNo - it's custom silicon https://news.ycombinator.com/item?id=48693490 https://news.ycombinator.com/item?id=48693490
- mlrtime 4mo agoBecause I just tested it and it took 3-4 clarifications before it actually gave a correct response vs gemini/google search. It's not great, but good. I'd rather wait 3x as long.
- mike_hearn 4mo agoBecause there's been nothing to discuss since their announcement. Their API access immediately closed due to overwhelming demand and they didn't fab newer models than Llama3 yet. Probably they will make bank selling to HFT for a while.
- vitorgrs 4mo agoNot opening here... HN killed?
- Bombthecat 4mo agoWhat How? Which model is behind it?
- victorbjorklund 4mo agoDamn that is crazy.
- archon810 4mo agoThis is the reaction every time it's posted, and deservedly so.
- dirasieb 4mo agohugged to death?
- jeingham 4mo agoThis caused me to have some sense what blistering fast AI actually is. What it means for the future is a question that remains.
- refulgentis 4mo agoImagine a Beowulf cluster of these…
- noisy_boy 4mo agoThat's a name I haven't heard in a while.
- mlrtime 4mo agoFirst post?
- cactusplant7374 4mo agoThe user has many comments and updoots if you look at their profile.
- refulgentis 4mo agoWe’re being silly and spamming Slashdot spam comments :p (“imagine a Beowulf cluster” and “first post?”)
- addaon 4mo agoMe too!
- chromadon 4mo agoI always think of Furbies because of that geocities (memories!) site.
- senectus1 4mo agoprobably something like this https://sb0xw.csb.app/ https://sb0xw.csb.app/
- alienbaby 4mo agoI started with a 2400baud modem, I've seen how this goes
- accrual 4mo agoSometimes I visualize a setup like this [0], based on 2D art by Simon Stålenhag. Someone has their home robot sitting on a desk connected to their old PC with thick cabling, dumping endless lines of each subsystem's <think> logs to diagnosis why it did something weird earlier in the day. Systems pushing 750+ tokens per second per subsystem might even be considered on the slow side for realtime tasks by then. [0] https://www.therookies.co/entries/39513 https://www.therookies.co/entries/39513
- bredren 4mo agoProbably will not be looking at text like this in a few years.
- cactusplant7374 4mo agoProbably not. Everyone will still need a lot of reasoning tokens and tool calls. Running the tests for every round is tiring but must be done.
- amluto 4mo agoThat’s an awful visualization. I can skim code quite quickly, but not when it shows up one character at a time in a small window, modem style. At least that site should draw out a full page then start replacing that page with the next, starting from the top and working downwards, repeating each time it hits the bottom.
- HSO 4mo ago> I can skim code quite quickly are you by any chance hyperlexic? interested to hear more about this, like how fast is considered fast
- jaapz 4mo agoThis is how tools like claude code and chat prompts output their tokens, so I'd say it's actually a pretty good visualisation.
- amluto 4mo agoNot for prefill. I suppose if you just want to imagine what generation speed looks like in the current generation of TUIs, it’s an okay visualization.
- elxr 4mo agoThat's exactly what it looks like in the tools I use most (opencode and codex), so for that purpose it's a pretty good visualization.
- me-vs-cat 4mo agoYou get used to it. I don't even see the code. All I see is blonde.. brunette.. redhead.