4 ms·
> If it can do 90% of the tasks the big boys do, at 50% speed I want to live in this world too, but these numbers, as of today, are very aspirational and far r
by overfeed 4mo ago
> If it can do 90% of the tasks the big boys do, at 50% speed
I want to live in this world too, but these numbers, as of today, are very aspirational and far removed from reality.
I'm no tokenmaxxer; I find my modest local setup useful, I also know the limitations, it's slow and it sucks (relatively) at high-level and/or long-context planning, compared to frontier models. Only a minority of my prompts are max-effort - its not all I do, but, it also means frontier labs aren't dying any time soon
- mikestorrent 4mo agoConsider also that right now LLMs run slowly enough you can watch them think. I've seen a demo of an LLM running at an absurdly high speed and it reminds me of when I moved from a 2400 baud modem to a 14.4 - BBS screens that I could watch draw were all of a sudden nigh-interactive. Faster-than-realtime video generation is also coming, and will also continue to require huge hardware for a long while yet. I love local models - I have a machine at home that runs a few for me and it's a lot of fun - but for the time being they are not super trustworthy on tool calls and staying on script. Another year or so might change all that!
- ChickeNES 4mo agoWhat does your local setup look like?
- deleted 4mo ago[deleted]
- mikestorrent 4mo ago1650 watts of liquid cooled Risc-V
- KoolKat23 4mo agoIf anyone wishes to see the future. A fast LLM is quite eye-opening. I think chatjimmy uses Talaas' chips where models are hardcoded into the silicon. https://chatjimmy.ai/ https://chatjimmy.ai/
- flyingjoe 4mo agoThanks, I didn't know that one! Very impressive speed although quality seems very bad
- msdz 4mo ago> although quality seems very bad The weights they “etched” into the FPGA card that’s used for the ChatJimmy demo are that of a Llama 3-something 8b model. The actually impressive and novel thing is that Taalas’ve managed to automate that process (clearly – nobody transforms 8 billion numbers into a physical representation by hand). So now, they can work on scaling this process up, and with low enough lead times (I’ll be convinced they have inside connections to TSMC if they can actually deliver on the promised mere 3-4 months delay), will be able to offer 30-100b+ parameter models under half a year after they’re released, at thousands of tokens per second while probably drawing less wattage (per token, not sure about overall). Exciting times ahead, folks.
- mikestorrent 4mo agoAt some point we're going to have these chips for a couple bucks in kids toys
- calgoo 4mo agoYea, is almost "scary fast" in a sense... the amount of compute you can do in parallel one one of those chips is amazing. hopefully they get their next chip completed as that will be a lot more useful for general workloads. I think the current one is based on llama3. *corrected llama version to 3
- kennywinker 4mo agoI’m sure you’re right, for the things you are asking of an llm, just as I am right about the things I am asking of an llm. The real question is, what are 90% of people going to ask llms to do. I’d argue mostly it’s going to be stuff that works-now or almost-works on local models, but that’s just an opinion. It also depends on the frontier models hitting a wall of steeply diminishing returns, since they set the expectations for all of this stuff - my gut says that’s happened already they just won’t admit it for a while - but we’ll see.