2 ms·
Seems like what Apple's going for with afm3. Their latest model that will be embedded in macOS 27 is a quantized dense 20B that only select between 1 to 4B at i
by marci 2mo ago
Seems like what Apple's going for with afm3. Their latest model that will be embedded in macOS 27 is a quantized dense 20B that only select between 1 to 4B at inference, based on the prompt, not token by token. If only they could make a 100B or 400B dense that selects ~5 to 15B...
- josu 2mo agoI don't understand, if they are only using a subset of the tokens then it's a sparse model. What do you mean by dense?
- l33tman 2mo agoCould it be some sort of permanently routed MoE where they detect and switch for the whole prompt instead of token by token?
- marci 2mo agoNothing to understand. Straight up hallucination. I could have sworn I read that they used a novel architecture where the model is dense but you could select specific layers or something at inference. reread the announcement: just said MoE. Corrected my brain's weights so thanks. https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models https://machinelearning.apple.com/research/introducing-third...
- josu 2mo agoThanks for responding.