3 ms·
With these patterns emerging, does anyone know how local LLMs are faring? It seems to me that by combining MCP and "skills", we are adopting LLMs to be more us
by e12e 1y ago
With these patterns emerging, does anyone know how local LLMs are faring?
It seems to me that by combining MCP and "skills", we are adopting LLMs to be more useful tools; with MCP we restrict input and output when dealing with APIs so that the LLM can do what it is good at; translate between languages - in this case from English to various json subsets - and back.
And with skills we're serializing and formalizing prompts/context - narrowing the search space.
So that "summarize q1 numbers" gets reduced to "pick between these tools/MCP calls and parameterize on q1" - rather than the open ended task of "locate the data" and "try to match a sequence of tokens representing numbers - and generate tokens that look like a summary".
Given that - can we get away with much stupider LLMs for these types of use cases now - vs before we had these patterns?
- simonw 1y agoThis is definitely a problem. Skills require a very strong model - one with a longer context (32,000 tokens minimum at a guess) that can reliably drive Unix CLI tools over a multiple step conversation. I haven't yet run a local model that feels strong enough at these things for skills to make sense. Really I think the unlock for skills was o3/Claude 4/GPT-5 - prior to those the models weren't capable enough for something like skills to work well. That said, the rate of improvement of local models has been impressive over the past 18 months. It's possible we have a 70B local model that's capable enough to run skills now and I've not yet used it with the right harness.
- saltwounds 1y agoI connect local models to MCPs with LM Studio and I'm blown away at how good they are. But the issues creep up when you hit longer context like you said.
- saltwounds 1y agoOpenAI and Anthropic's real moat is hardware. For local LLMs, context length and hardware performance are the limiting factors. Qwen3 4B with a 32,768 context window is great. Until it begins filling up and performance drops quickly. I use local models when possible. MCPs work well, but their large context injection makes switching to an online provider the no-brainer.