8 ms·
At 8x86B, looks like the largest open model yet by far. Would be interesting to hear how many tokens it's been trained on. Especially important for higher param
by extheat 3y ago
At 8x86B, looks like the largest open model yet by far. Would be interesting to hear how many tokens it's been trained on. Especially important for higher param models in order to efficiently utilize all those parameters.
- zone411 3y agoIt's actually not the largest. https://huggingface.co/google/switch-c-2048 https://huggingface.co/google/switch-c-2048 is 1.6T parameters.
- WeMoveOn 3y agobut is switch c even usable? iirc the training set was nowhere near enough for a model of that size to be coherent in a conversation
- p1esk 3y agoIt’s not 8x86B. Total number of parameters is 314B. Perhaps it’s 8x39B to fit on a single 8xA100 (40GB) server?
- dheera 3y agoThey all do this marketing bull. Mixtral has an 8x7B model but it's actually 46.7B, not 56B params. Kinda similar to how 4K displays are 3840 pixels wide, not true 4K which would be 4096. Marketing people called it 4K, not engineers.
- guitarlimeo 3y agoI've always thought of 4K as "4x FullHD". In that way it makes sense.
- mavhc 3y agoTV and Digital Cinema have different standards, because of course they do
- deleted 3y ago[deleted]
- dheera 3y agoBleh no, K means thousand. For a long time we specified displays by their vertical dimension -- 480p, 720p, 1080p. Then the marketing guys came along and decided that the horizontal dimension sounds bigger. If we stuck with the less-bullshitty way of doing things and kept comparisons 1:1, we'd call 3840x2160 displays 2160p or "2K" displays, but instead, the marketing people decided that we're going to change things to horizontal and called 3840x2160 "4K".
- throwaway11460 3y agoIt's 2x Full HD though
- _kuvn 3y ago2x in a single direction, 4x the number of pixels
- throwaway11460 3y agoOh yeah... What I meant is, 1920x2 = 3840 ~~ 4000
- moffkalast 3y agoMost likely it's a MoE of Grok-0 which would be 8x33B + 50B for the router.
- cma 3y agoActive parameters is 86B, so wouldn't that be the size of the largest two experts (where they may all be the same) + the weights of the selector?
- swalsh 3y agoConsidering how poor it is compared to other models, it really emphasises how important fine tuning is. Models with MUCH smaller parameter counts are outperforming it in many metrics.
- gordian-mind 3y agoCurrent metrics are a poor way to measure the usefulness of LLMs.
- make3 3y agono it empathizes the importance of training smaller models for longer, like the Mistral "overtrained" models
- gdiamos 3y agoShow the proof? Does it include IFT?
- lukan 3y ago"it really emphasises how important fine tuning is" Or rather the quality of the training data?
- jakderrida 3y agoAren't they usually built on most of the same training data?
- fragmede 3y agothat's a subtle dig at the fact that they have all of Twitter as a training corpus to use, but we don't know how they weight tweets. which, we know they're not gonna be weighted evenly.
- rezonant 3y agoI'm sure just like in X's algorithms, @elon tweets are weighted heavily.