6 ms·
Gemma 3 270M re-implemented in pure PyTorch for local tinkering
- vi0g0d 1y ago[dead]
- canyon289 1y agoHey all, I created this model with a top notch team. I answered many questions last week when this hit the front page, and happy to answer more here as well. https://news.ycombinator.com/item?id=44902148 https://news.ycombinator.com/item?id=44902148 Personally I'm excited that you all have access to this model now and hope you all get value out of using them.
- deleted 1y ago[deleted]
- GaggiX 1y agoI imagine you and your team have finetuned the model on different tasks, can you share some results? (I have only seen the alien NPC finetuning)
- canyon289 1y agoThe Unsloth folks have finetuning numbers. Linking their post here https://www.reddit.com/r/unsloth/comments/1mq5hbb/google_gemma_3_270m_out_now/ https://www.reddit.com/r/unsloth/comments/1mq5hbb/google_gem...
- WithinReason 1y agoI would like to know your thoughts on using 2/3 of such a small the model's size for embeddings. What would be different if you used a byte-level vocabulary and spent the parameter budget on transformer parameters instead? I think you would lose performance (tok/s) but might gain accuracy.
- canyon289 1y agoAt this small scale the embeddings indeed were a big focus. Consider this thought process. The tokens themselves are a form of compression. Lets say we have the word "WaffleHouse", character level this would be 11 tokens, but with an embedder this would be perhaps 2 or 3 tokens (I didn't actually run through the tokenizer but we could verify precisely). This matters a lot for on device processing especially. So while we could get more intelligence out of the model by bumping up the "knowledge" parameters, the device would need to process more input and output tokens. Another advantage on small devices is the embeddings are just a lookup table which requires little to no computation. Its the rest of the parameters that have the expensive matrix multplications, so if we increased those we'd also be increasing the number of FLOPs needed for a forward pass. This blog post explains it well. https://www.adamcasson.com/posts/transformer-flops https://www.adamcasson.com/posts/transformer-flops So all this to say is there are definite tradeoffs between model size, performance on evals, and compute cost. We ran many internal experiments with different choices to see could work well, and then picked what we believed work will best for the open community.
- Scene_Cast2 1y agoHow would this matrix get trained with PyTorch? I currently have a toy Transformer network - I ended up marking the matrix as sparse and using SparseAdam - gives a bit of a performance boost, but at the same time I can't use torch.compile() on the fetch from this matrix.
- WithinReason 1y agoMakes sense, thank you.
- PoignardAzur 1y agoDoes Gemma use any specific scheme to compress embeddings? Which have you considered? For instance, it's well-known that transformer embeddings tend to form clusters. Have you considered splitting the embedding table into "cluster centroid" and "offset from centroid" tables, where the later would presumably have a smaller range and precision?
- 1y ago
- tarruda 1y agoThanks for your work, it is really an amazing small LM. Can you share what kind of hardware is necessary to train it, and how long it took?
- canyon289 1y agoThank you! The Gemma3 technical report contains many details on training setup https://arxiv.org/pdf/2503.19786 https://arxiv.org/pdf/2503.19786 This was released with the initial batch of Gemma3 so it doesn't contain the 270m details, nonetheless you'll get a good idea of what it takes to build these models.
- owebmaster 1y agoDoes it have function calls? Can we use it with MCP?
- canyon289 1y agoIt can possibly perform basic prompted FC but I wouldn't get your hopes up. It should be to be a solild FC model if trained on specific tools and format. I would not expect great MCP performance because the context window is 32k and most MCP servers I've see implicitly assume massive context windows.
- riedel 1y agoVery stupid question: why does the tflite model output only '[multimodal][multimodal]' when executed on GPU in the AI edge gallery app, while fully working on the CPU.
- bigyabai 1y agoThanks for making this! One of my favorite projects was having a Discord chatbot powered by the original BERT model - these 270M weights are a fine upgrade.
- dcreater 1y agoAs a non MLE, what are the pros/cons of OP's PyTorch re-implementation?
- mdaniel 1y agoI'm not a ML engineer, so I can speak to the "non MLE" bit from my perspective (literal tl;dr: learning and experimentation opportunity) 1. Since it's just PyTorch, that means one can run it locally upon whatever accelerator you have that PyTorch supports. For quite a few people that includes Metal Performance Shaders: https://docs.pytorch.org/docs/stable/mps.html https://docs.pytorch.org/docs/stable/mps.html I can attest that building PyTorch from git is achievable in about 15 minutes on my M1 Pro, if you really want to chase the rabbithole. Cloning PyTorch is its own special 'please. wait.', but building it is fine 2. Since it's (of the ones that I've looked at) approximately 500 lines long, it's much, much, much more digestable than a lot of the vomit that comes out of so-called production systems. Those systems usually have only heard about typed Python in passing, and they believe it is a fad that will blow over. The ones in this repo aren't stellar about it, but at 500 lines it's easily achievable to type hint the code yourself, which can serve as an excellent learning opportunity 3. PyTorch offers some fun conversion tools, also, allowing one to compare-and-contrast how it executes under Torch versus ONNX <https://docs.pytorch.org/docs/stable/onnx.html https://docs.pytorch.org/docs/stable/onnx.html>, TorchScript <https://docs.pytorch.org/docs/stable/generated/torch.jit.save.html https://docs.pytorch.org/docs/stable/generated/torch.jit.sav...>, CoreML <https://apple.github.io/coremltools/docs-guides/source/convert-pytorch.html https://apple.github.io/coremltools/docs-guides/source/conve...>, or a bazillion other competing frameworks 4. Related, one can play around with quantization and other "inference related" concerns (e.g. https://github.com/pytorch/ao#pytorch-native-training-to-serving-model-optimization https://github.com/pytorch/ao#pytorch-native-training-to-ser... ) 5. Further related, one can play around with the fine-tuning mentioned elsewhere, to better understand what is and isn't possible to achieve using that process. Because the code is digestable, and the models are reasonably sized (Qwen 0.6B weighs only 1.4GB and is Apache 2), it brings FAFO opportunities in ways that gpt-oss-20b (or bigger!) won't I do appreciate that some of what I said may skate close to "ML engineer" concerns, so obviously your situation will be different, but for me having a better grip on how these things work enables me to have better conversations with my colleagues and also helps trip my bullshit detector when someone claims they're the second coming and are going to cure cancer or whatever
- khalic 1y agoI'm going to have so much fun tinkering with it, thank you!!!
- n0vella 1y agoDo you think these very small models have some utility in the real world? Apart from learning and academic purposes of course.
- colechristensen 1y agoSure, interacting with natural language without expectation that the model contains knowledge. Good for things like tool use and embeddings where the information is all retrieved.
- throw310822 1y agoAre these small models are trained to privilege "raw intelligence" over factual knowledge? Is there any indication of how much of current model is dedicated to the knowledge of multiple languages and tons of facts rather than pure understanding and reasoning?
- canyon289 1y agoThe evaluations provide this indication. You'll see MMLU, GPQA, Big Bench etc in reports for many models. Those numbers provide the indication you're looking for. To answer a question you didn't ask. With small models especially we need to make choices as to which to focus on. For this model we focused on text summarization and instruction following, with the idea that users would finetune to gain performance on the task set that is relevant to them
- canyon289 1y agoYes! To me the primary value is not just as a teaching or toy model. I see a lot o value in repeatable tasks if we think about enterprise and a local fast developer model for individual usage. Here's some examples that are inspired by previous roles I had outside of Google, where a business I was working in needed real time text processing. This tutorials were made with Gemma versions from a year ago, but could now be recreated with Gemma 270m https://developers.googleblog.com/en/gemma-for-streaming-ml-with-dataflow/ https://developers.googleblog.com/en/gemma-for-streaming-ml-... https://www.youtube.com/watch?v=YxhzozLH1Dk https://www.youtube.com/watch?v=YxhzozLH1Dk
- lsb 1y agoThat’s wild that with a KV cache and compilation on the Mac CPU you are faster than on an A100 GPU.
- Weryj 1y agoThis would be because the GPU can’t fill its waveform and hide memory latency, no? I’m curious for a reason why
- punnerud 1y agoBecause on Mac the CPU and GPU share memory, but A100 need to transfer to RAM/CPU on the parts that’s not supported by GPU? (My first guess)
- ladberg 1y agoGiven that the compiled version is slower than then eager version on A100, there's definitely something suboptimal happening there
- ModelForge 1y agoNo the compiled version is actually faster. From that table, the A100 tok/sec (larger is faster) numbers are: - Eager: 28 - Compiled: 128 And - KV cache eager: 26 - KV cache compiled: 99 The reason that the KV cache is slower is likely because it's not GPU-optimized code. On CPU the KV cache is faster. To make it faster on GPU, you would pre-allocate the tensors on the device for example instead of `torch.cat`ting them on the fly
- ladberg 1y agoAh yep read the labels backwards and meant that - ty for catching and for the explanation
- ModelForge 1y agoCould be an artifact of the small size not fully taking advantage of the GPU. For example, for the slightly larger Qwen3 0.6B model the A100 is faster (you can see it when scrolling to the bottom here: https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/11_qwen3 https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/11...)
- shekhar101 1y agoCan someone (or OP) point me to a recipe to fine tune a model like this for natural language tasks like complicated NER or similar workflows? I tried finetuning Gemma3 270M when it came out last week without any success. A lot of tutorials are geared towards chat applications and role playing but I feel this model could be great for usecases like mine where I am trying to extract clean up and extract data from PDFs with entity identification and such.
- hmottestad 1y agoHave you tried this one here by any chance? https://huggingface.co/dslim/bert-base-NER https://huggingface.co/dslim/bert-base-NER Just wondering if it’s worth testing and what it would be most useful for.
- nolist_policy 1y agoThis is using the gemma-llm python library which uses JAX in the background: https://gemma-llm.readthedocs.io/en/latest/colab_finetuning.html https://gemma-llm.readthedocs.io/en/latest/colab_finetuning....
- lgessler 1y agoIf you're really just doing traditional NER (identifying non-overlapping spans of tokens which refer to named entities) then you're probably better off using encoder-only (e.g. https://huggingface.co/dslim/bert-large-NER https://huggingface.co/dslim/bert-large-NER) or encoder-decoder (e.g. https://huggingface.co/dbmdz/t5-base-conll03-english https://huggingface.co/dbmdz/t5-base-conll03-english) models. These models aren't making headlines anymore because they're not decoder-only, but for established NLP tasks like this which don't involve generation, I think there's still a place for them, and I'd assume that at equal parameter counts they quite significantly outperform decoder-only models at NER, depending on the nature of the dataset.
- keeeba 1y agoWhat use-cases do you see for the 270M’s embeddings, and should we be sticking to token embeddings or can we meaningfully pool for sentence/document embeddings? Do we need to fine-tune for the embeddings to be meaningful at the sentence/document level?
- eachro 1y agoIf you wanted to train it from scratch, how long would it take on a reasonable GPU setup?
- canyon289 1y agoThe world reasonable is vague but assuming you mean something that could be run in a residential unit it would long a very long time if training from pure scratch. This is part of the rationale for releasing this model. Now you don't have to start from scratch and finetuning is reasonable on a wide variety of hardware, including reasonable GPU setups (and smaller)
- rck 1y agoFor the sake of comparison, you can train a 124M model on a 3090 (see nanoGPT). In that case, each batch ends up having about 500,000 tokens and takes maybe around 10ish seconds to run forward and backward. Then the 6 trillion tokens that this model was trained on would take about 4 years, approximately. Or just "too long" for a shorter answer.
- torben-friis 1y agoThis might be a very basic question, but as a dev whose only interaction with models is using the main commercial ones (sonnet, ChatGPT and the like), what are some usecases for these smaller local models? What usages can be reasonable to expect from them? Are there uses out of the box or does one have to go through some custom post-training to get useful behavior? I feel like there is a huge gap between understanding models as a user of commercial tools and the kind of discussions happening in these threads, but I’m not sure what are the in-between steps.
- barrkel 1y agoSummarization, very basic tool use, without needing to go across the internet and back, and zero cost because of edge compute.
- _giorgio_ 1y agoMaybe also secrecy and privacy.
- canyon289 1y agoIts a crucial question. I wrote up a long answer here. Let me know it helps https://news.ycombinator.com/item?id=44913558 https://news.ycombinator.com/item?id=44913558
- torben-friis 1y agoThanks for the reply! It does help to figure out where in the space this model fits. I'm still a bit confused about this part: >since it needs to be shaped to match specific tasks, we did our best to design it to be a flexible starting point for LLM-style tasks and worked with partners to put it into the right frameworks and places for you all to be able to shape it to what you need it to be. What does shaping mean in this case? What tools are used, what requirements are there, both in terms of hardware and knowledge? I would like to go beyond being spoonfed by large companies' high usability products, both to improve my knowledge and not be a victim of potential future rug pulls. In the classic software world, I guess the equivalent would be someone who runs open source software navigating the extra complexity, and ocassionally collaborates with the projects. But I don't know what that looks like in the AI world. I've gone through some courses on machine learning but learning the basics about hessian matrices and gradient descent seems as detached from the practical point I'm searching as taking a compilers class is from learning React, so I think I've been looking in the wrong places (?).
- _giorgio_ 1y agowhat a legend
- quesne 1y agoThought it was a new 3270 interface, bummed.
- mattfrommars 1y agoIs this the same thing as people did the past with '<model> inference written in vanilla Go, Python, Java, etc" ?
- dboon 1y agoFirst, thanks for doing everything you do! I, and I’m sure countless others, genuinely benefit from you. How would you recommend someone with a strong background in undergraduate level traditional ML get into deep learning? I use that as a broad term to encompass all the knowledge needed to understand how these models work, starting from the deep learning models of a decade ago, plus the practical ability to collect data or build RL gyms and fine tune them. I understand ML math well enough that I’m confident I could follow a modern white paper after a lot of effort and research. But there are so many pieces — quantizations, flash attention, Mode, batch sizes, layer sizes, model sparsity. I feel very overwhelmed trying to piece together how all of the pieces arose, and even more overwhelmed trying to figure out how one even goes about fine tuning one. I (like most people here) am extremely technical, and it’s not often I feel this way about a field. Thanks again! Best of luck on your work
- Quarrel 1y ago> I’m confident I could follow a modern white paper after a lot of effort and research. Without having done it for deep learning, I'm sure it is like any other area of computer science. You get to exactly the level you're at now, and then you put in that effort following modern papers, and each one gets easier and easier. A year later you've done the literature review for your Phd. :)
- hodgehog11 1y agoAs someone who has students that work in deep learning, I can say that it is unwise to approach deep learning in the same way as traditional ML. Most classical methods are strongly mathematically motivated and have excellent theory to accompany them. Deep learning is still alchemy; it is a matter of experience, trying things out and getting a feel for how the pieces fit together in a modular format. Once you are experienced with the common building blocks, you can develop an intuition for how they might be improved. I would start with training a basic MLP on tabular data. Then switch to CNNs: LeNet, VGG, then ResNet. Understand each of the new blocks that are incorporated into each architecture and how they improve stability and training efficiency. There are good PyTorch tutorials for these. Use these as a playground to understand what each of the training knobs do. Look at how their implicit biases induce double descent; this should give you confidence that overfitting is rarely an issue anymore. Give finetuning a try by taking a pretrained ResNet on ImageNet, adding layers to the start and end, and training only these to adapt the model to another image dataset. This should demonstrate the power of finetuning and why pretrained models are so powerful. Next, briefly consider a tutorial on LSTMs, recognizing the exploding and vanishing gradient problems and the traditional challenges with sequential data. Then move to transformers. Work with language first, starting from Andrej Karpathy's excellent YouTube tutorials. Train the model in full for a bit, then see about using an existing GPT2 checkpoint. Try adapting NanoGPT to a mathematical dataset as an exercise. Then take a look at llm.c to see how to really improve performance. Finally, take a look at ViT and DETR. Use pretrained models and finetune them on smaller datasets again. By this point, you should have a good grounding to start reading much of the surrounding literature and understand them. You should also understand that models are never built from scratch anymore, and every model is a collection of individual pieces built elsewhere for a particular purpose.
- PunchTornado 1y agoanybody can help me with some tutorials on how to use this for mechanistic interpretability?
- Xplan 1y ago[dead]