3 ms·
Yes, and also waiting for the next iteration of Gemma. Muse or Qwen are optimized for coding, while IMO Gemma is still better for non-coding tasks. https://x.c
by karimf 2mo ago
Yes, and also waiting for the next iteration of Gemma. Muse or Qwen are optimized for coding, while IMO Gemma is still better for non-coding tasks.
https://x.com/osanseviero/status/2086107547535122767 https://x.com/osanseviero/status/2086107547535122767
- dannyw 2mo agoYou can partially tell by the tokeniser; which gives you some hint into the training corpus mix. </div> is four Gemma4 tokens, but one Qwen3.6 token.
- venusenvy47 2mo agoWhere do you find this information for each model?
- ComputerGuru 2mo agoThe tokenizers are included in the open s̶o̶u̶r̶c̶e̶ weights releases; you wouldn’t be able to use the weights without the corresponding encoder/decoder, in fact.
- adrian_b 2mo agoWhen you look on HuggingFace.co at the files of a model, for each model you will see a file "tokenizer.json". In that file you can see all tokens and their corresponding numeric codes.
- deleted 2mo ago[deleted]
- stymaar 2mo agoLooks like we have a /r/localllama dweller here.
- malshe 2mo agoI am working on a project where we have to classify customer calls into more than 10 categories. As the client wants everything locally I tried a few local LLMs. Gemma turned out to be the best model for this task. The classification accuracy is impressive, and the client is happy that I am using an American model.