3 ms·
I wonder if there is way local small LLMs can complement each other in away that the sum-total yields a much more performant LLM
by pizzao 4mo ago
I wonder if there is way local small LLMs can complement each other in away that the sum-total yields a much more performant LLM
- killerstorm 4mo agoPerhaps some radical MoE where you download _exactly_ the components you need as you need them. Currently MoE is switched usually on per-token per-layer basis, so you need all weights locally. But e.g. Apple made one which pre-selects all experts based on prompt embedding. That might be further scaled up - e.g. predict exactly what's needed
- salter2 4mo agoPerhaps something similar to speculative decoding. Speculating Experts Accelerates Inference for Mixture-of-Experts: https://arxiv.org/abs/2603.19289 https://arxiv.org/abs/2603.19289
- eblanshey 4mo agoI don't understand why no labs create dedicated models per industry/expert. E.g. physics, electronics, chemistry, etc. Each model would be much smaller and better suitable for running locally. Everyone is trying to cram everything into a single model.
- Flere-Imsaho 4mo agoSort of like how ants in a colony produce a working "society" that no individual ant could muster.