3 ms·
Reminds me of Deep Thought from Hitchhiker's Guide to the Galaxy
by lukeduff 27d ago
Reminds me of Deep Thought from Hitchhiker's Guide to the Galaxy
- schmorptron 27d agoIt's kind of insane how having this tech at this speed 5 years ago would have probably still been seen as insanely useful and revolutionary. If LLMs were more capable but dramatically slower, I wonder how it would impact how we use it? Dramatically more thought being put into prompts, much more preparation probably
- IgorPartola 27d agoMy parents learned to program on punch cards. They told me it was a day of preparing the program, an hour of running it, just to get a syntax error.
- deleted 26d ago[deleted]
- bluedino 26d agoWrite the program, punch the cards, send the cards to another building to be loaded, program runs, printout comes out in another building, somehow this takes 2-3 days
- lurker919 26d agoCoding is the new punch card slots now. My children will listen in awe about how typing and testing used to take hours or even (gasp!) days.
- redox99 27d agoAt 1t/s it's still faster than humans for a lot of tasks, basically doing overnight what could take humans half a week. Plus you can always parallelize.
- Izmaki 26d agoThis is what people forget when they see slow performance: at 1 t/s it's still roughly the equivalent of having another person work for you at no extra cost besides the initial purchase/sign-on-bonus. Frontier models are amazing, but what will really be useful for us is having models and hardware so efficient that you can run useful LLMs locally. One of my favourite LLMs to this day is still my jail-broken gemma4 12b because it's small enough to run on my computer, but also 100% local and free as in liberty.
- mhaberl 26d agoI would agree, but I want to add that I have real issues with combination of opencode plus slow inference (4-5tok/s). I get weird interruptions. I can only guess its related to some kind of timeouts in the harness or something. Its not a problem of the model of course, but it seems impractical atm. I wonder if anyone else had this kind of thing happening.
- batperson 26d agoThe future of inference is likely in ASICs, so we'll get the inverse, a bit less capable than frontier but super fast models. Like this 14k tok/s beast https://chatjimmy.ai/ https://chatjimmy.ai/ from Taalas (who got acquired by AMD recently). GPT-6-astra runs at like ~40 tok/s, I have a hard time imagining what could be accomplished with that type of model at 10k+ tok/s when in the hands of the public. Will certainly make cybersecurity a challenge for older systems.
- domhudson 26d agoThis is incredible! Are there other big players in this space (freezing models to silicon)?
- HeWhoLurksLate 26d agotake a look at Cerebras, who are doing wafer-scale compute
- timcobb 26d agoI imagine Astra is/will soon will be on Cerebras?
- SPascareli13 26d agoLike how crypto used ASICS but then didn't because the scaling of consumer hardware made it obsolete?
- KumaBear 26d agorunning my own locally. I just set the tasks to start when systems go idle over x. Read and copy only to external drive projects, codes, ect for review. I review the reports the changes and apply them myself or correct them. Is it slower than say throwing it into fable yes. But I don't have to be monitoring it 24/7
- mitxela 26d agoThat's because 5 years ago it was still brand new for a computer to be able to speak English. 5 years later, we have accepted that LLMs can speak English and we expect them to do useful things.