4 ms·
Indeed, it begs the question why we have "everything" models where instead we could have very efficient "something" models. Typical LLMs out there can generate
by glitchc 2mo ago
Indeed, it begs the question why we have "everything" models where instead we could have very efficient "something" models. Typical LLMs out there can generate code and translate between 60 different languages. Sometimes I only need the first part, sometimes the second. Two distinct models would be a lot smaller and run much faster (token-wise).
- panarky 2mo agoBut they'd be stupider. The results for English and Python are much better because the model is also trained on Mandarin and Greek and Lisp even if you never make a request or receive a response in Mandarin, Greek or Lisp.
- glitchc 2mo agoThat's news to me since it's unlikely most of those weights are activated when responding to a coding prompt. Can you point to a source/paper that validates this claim?
- panarky 2mo agoIt's well known and well documented in AI research. Instead of taking syntactic shortcuts, the richer abstractions learned from multiple languages, and code, and math, and images and audio, ultimately make English comprehension and reasoning far stronger. If you want citations, ask your favorite LLM how linguistic diversity prevents "surface memorization", overfitting on surface-level English patterns instead of representing the deeper concepts in latent vector space.