3 ms·
How small can be a LLM transformer in order to be able to understand basic human language and search for answers on the internet? It should not contain all the
by novaRom 3y ago
How small can be a LLM transformer in order to be able to understand basic human language and search for answers on the internet? It should not contain all the facts and knowledge, but must be quick (so, it's a small model), understand at least one language, and know how and where to look for answers.
Would it be sufficient to have 1B, 3B or 7B parameters to achieve this? Or is it doable with 100M or even fewer parameters? I mean vocabulary size might be quite small, max context size could also be limited to 256 or 512 tokens. Is there any paper on that maybe?
- minimaxir 3y agoYou don't need a 1B+ parameter model for this workflow, but it helps, particularly for the quality of the final output. > know how and where to look for answers The point of calling them Agents is that they don't know how. All the examples, including AutoGPT that makes AI influencers go OMG AGI, are operating in a discrete action space with user-specified hints to select which action (or none at all).
- quickthrower2 3y agoI like the idea that a small model could fine tune itself for each task more affordably, so that could be an advantage.
- lhl 3y agoA team at Microsoft Research asked the same question and just published a paper about part of that at least: TinyStories: How Small Can Language Models Be and Still Speak Coherent English? https://arxiv.org/abs/2305.07759 https://arxiv.org/abs/2305.07759 "We show that TinyStories can be used to train and evaluate LMs that are much smaller than the state-of-the-art models (below 10 million total parameters), or have much simpler architectures (with only one transformer block), yet still produce fluent and consistent stories with several paragraphs that are diverse and have almost perfect grammar, and demonstrate reasoning capabilities." They also trained a TinyStories-Instruct instruction following variant. It looks like the 28M parameter 8 layer model had a 9/10 on grammar and consistency. You could probably combine that with something like Jsonformer or parserLLM to enforce valid formatting.
- lt 3y agoThe contents of whatever it found online is fed back in the context, plus the response it generates based on that also counts to the limit. So 512 is really inadequate unless you just want to make search queries using natural language.