2 ms·
> Somehow I expected inference engines are generic LLM runtimes that can execute any weight. In fact, this was close to be true until last year: almost every o
by stymaar 23d ago
> Somehow I expected inference engines are generic LLM runtimes that can execute any weight.
In fact, this was close to be true until last year: almost every open model except DeepSeek had a very similar architecture that was pretty close to the GPT-2 one with very few variations on top (and sometimes an MoE architecture, which itself was a few year old at that point).
But a year ago there's been a cambrian explosion, first in attention mechanism but also in a bunch of other directions, mostly coming from China, and now there's a very massive diversity today's space.
- anuj0456 23d agoyes, that is correct. after chinca came into picture the advancement in this field sky rockted