12 ms·
Hello on behalf of the Gemma team! We are really excited to answer any questions you may have about our models. Opinions are our own and not of Google DeepMind
by alekandreev 3y ago
Hello on behalf of the Gemma team! We are really excited to answer any questions you may have about our models.
Opinions are our own and not of Google DeepMind.
- brucethemoose2 3y agoWill there be "extended context" releases like 01.ai did for Yi? Also, is the model GQA?
- hustwindmaple1 3y agoIt's MQA, documented in the tech report
- nuclearjam 3y agoMay I ask what is the ram requirement for running the 2B model on CPU on an average consumer windows laptop? I have 16 gb RAM but I am seeing CPU/memory traceback. I’m using the transformer implementation.
- zitterbewegung 3y agoDo you have a plan of releasing higher parameter models?
- alekandreev 3y agoWe have many great things in research and development phases, so stay tuned. I’m hopeful we can share more in the coming weeks and month!
- brucethemoose2 3y agoThat is awesome! I hope y'all consider longer context models as well. Also, are ya'll looking alternative architectures like Mamba? Being "first" with a large Mamba model would cement your architectural choices/framework support like llama did for Meta.
- efilife 3y agoThis doesn't answer the question at all
- deleted 3y ago[deleted]
- declaredapple 3y agoCongrats on the launch and thanks for the contribution! This looks like it's on-par or better compared to mistral 7B 0.1 or is that 0.2? Are there plans for MoE or 70B models?
- kathleenfromgdm 3y agoGreat question - we compare to the Mistral 7B 0.1 pretrained models (since there were no pretrained checkpoint updates in 0.2) and the Mistral 7B 0.2 instruction-tuned models in the technical report here: https://goo.gle/GemmaReport https://goo.gle/GemmaReport
- neximo64 3y agoHow are these performing so well compared to Llama 2, are there any documents on the architecture and differences, is it MoE? Also note some of the links on the blog post don't work, e.g debugging tool.
- kathleenfromgdm 3y agoWe've documented the architecture (including key differences) in our technical report here (https://goo.gle/GemmaReport https://goo.gle/GemmaReport), and you can see the architecture implementation in our Git Repo (https://github.com/google-deepmind/gemma https://github.com/google-deepmind/gemma).
- h1t35h 3y agoIt seems you have exposed the internal debugging tool link in the blog post. You may want to do something about it.
- trisfromgoogle 3y agoAh, I see -- the link is wrong, thank you for flagging! Fixing now.
- neximo64 3y agoThe link to the debugging tool is an internal one, no one outside Google can access it
- h1t35h 3y agoThe blog post shares the link for debugging tool as https://*.*.corp.google.com/codelabs/responsible-ai/lit-gemma https://*.*.corp.google.com/codelabs/responsible-ai/lit-gemm... .corp and the login redirect makes me believe it was supposed to be an internal link
- littlestymaar 3y agoSame for the “safety classifier”
- barrkel 3y agohttps://codelabs.developers.google.com/codelabs/responsible-ai/lit-gemma https://codelabs.developers.google.com/codelabs/responsible-...
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- 3y ago
- pama 3y agoWill these soon be available on lmsys for human comparison against other models? Can they run with llama.cpp?
- ErneX 3y agoYes to llama.cpp https://twitter.com/ggerganov/status/1760293079313973408 https://twitter.com/ggerganov/status/1760293079313973408
- sbarre 3y agoI came here wondering if these models are "open" in the sense that they'll show up on sites like Ollama where you can download and run them locally. Am I correct to conclude that this means they eventually will? It's unclear to me from Google's docs exactly what "open" means for Gemma
- benpacker 3y agoYes - they are open weights and open inference code, which means they can be integrated into Ollama. They are not “open training” (either in the training code or training data sense), so they are not reproducible, which some have suggested ought to be a component of the definition of open models.
- OJFord 3y agoIt really should shouldn't it? I'm quite ML-naïve, but surely providing the model without 'training code or training data' is just like providing a self-hostable binary without the source code? Nobody calls that open source, it's not even source available.
- sunnybeetroot 3y agoThat’s why they’re called open as in free to use how you wish, not open source where the source of the training is also provided.
- artninja1988 3y agoI find the snyde remarks around open source in the paper and announcement rather off putting. As the ecosystem evolves, we urge the corporate AI community to move beyond demanding to be taken seriously as a player in open source for models that are not actually open, and avoid preaching with a PR statement that can be interpreted as uniformed at best or malicious at worst.
- silentsanctuary 3y agoWhich remarks are you referring to?
- artninja1988 3y agoThe synde remarks at metas llama license that doesn't allow companies with 700 million monthly active users to use it, while this model also doesn't have a really 'open' license itself and also this paragraph: >As the ecosystem evolves, we urge the wider AI community to move beyond simplistic ’open vs. closed’ debates, and avoid either exaggerating or minimising potential harms, as we believe a nuanced, collaborative approach to risks and benefits is essential. At Google DeepMind we’re committed to developing high-quality evaluations and invite the community to join us in this effort for a deeper understanding of AI systems.
- tomComb 3y agoWell, given that that restriction added to the meta-llama license is aimed at Google, is petty, and goes against open source norms, I think it’s reasonable that they should feel this way about it.
- lordswork 3y agoHow is this a snide remark? It's factual and prevented their team from benchmarking against Llama 2.
- trisfromgoogle 3y agoQuick question -- can you tell me where you got that quote? It's not in the main blog or any of the launch communications that I can see.
- tosh 3y agoAre there any plans for releasing the datasets used?
- alekandreev 3y agoThis would be really interesting in my opinion, but we are not releasing datasets at this time. See the C4 dataset for an earlier open dataset from Google.
- sbarre 3y agoCan the Gemma models be downloaded to run locally, like open-source models Llama2, Mistral, etc ? Or is your definition of "open" different?
- kathleenfromgdm 3y agoYes, you can get started downloading the model and running inference on Kaggle: https://www.kaggle.com/models/google/gemma https://www.kaggle.com/models/google/gemma ; for a full list of ways to interact with the model, you can check out https://ai.google.dev/gemma https://ai.google.dev/gemma.
- dartharva 3y agoCan we have llamafile releases as well? https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile
- syntaxing 3y agoA small typo in your model link that breaks it. There’s an extra ; on the end.
- kathleenfromgdm 3y agoCorrected - thanks :)
- Kostic 3y agoIt should be possible to run it via llama.cpp[0] now. [0] https://github.com/ggerganov/llama.cpp/pull/5631 https://github.com/ggerganov/llama.cpp/pull/5631
- nerdix 3y agoAmazing how quickly this happened.
- tomp 3y agoTheir definition of "open" is "not open", i.e. you're only allowed to use Gemma in "non-harmful" way. We all know that Google thinks that saying that 1800s English kings were white is "harmful".
- vorticalbox 3y agoare there plans to release an official GGUF version to use with llama.ccp?
- espadrine 3y agoIt is already part of the release on Huggingface: https://huggingface.co/google/gemma-7b/blob/main/gemma-7b.gguf https://huggingface.co/google/gemma-7b/blob/main/gemma-7b.gg... It is a pretty clean release! I had some 500 issues with Kaggle validating my license approval, so you might too, but after a few attempts I could access the model.
- vorticalbox 3y agoI didn't see this when searching thanks
- sqreept 3y agoWhat are the supported languages of these models?
- alekandreev 3y agoThis v1 model is focused on English support, but you may find some multilingual capabilities.
- lnyan 3y agoWill there be Gemma-vision models or multimodal Gemma models?
- Jayakumark 3y agoHave the same question.
- alekandreev 3y agoWe have many exciting things planned that we can't reveal just yet :)
- CuriouslyC 3y agoIt's cool that you guys are able to release open stuff, that must be a nice change from the modus operandi at goog. I'll have to double check but it looks like phi-2 beats your performance in some cases while being smaller, I'm guessing the value proposition of these models is being small and good while also having more knowledge baked in?
- alekandreev 3y agoWe deeply respect the Phi team and all other teams in the open model space. You’ll find that different models have different strengths and not all can be quantified with existing public evals. Take them for a spin and see what works for you.
- turnsout 3y agoWhat is the license? I couldn’t find it on the 1P site or Kaggle.
- trisfromgoogle 3y agoYou can find the terms on our website, ai.google.dev/gemma: https://ai.google.dev/gemma/terms https://ai.google.dev/gemma/terms
- deleted 3y ago[deleted]
- spiantino 3y agoout of curiosity, why is this a "terms" and not a license? I'm used to reading and understanding the software as coming with a license to use it. Do the terms give us license to use this explicitly?
- turnsout 3y agoThey do, but unlike a known license, these terms are custom and non-standard. Which means I would guide my commercial clients away from this particular model.
- deleted 3y ago[deleted]
- audessuscest 3y agoDoes this model also thinks german were black 200 years ago ? Or is afraid to answer basic stuff ? because if this is the case no one will care about that model.
- freedomben 3y agoI don't know anything about these twitter accounts so I don't know how credible they are, but here are some examples for your downvoters that I'm guessing just think you're just trolling or grossly exaggerating: https://twitter.com/aginnt/status/1760159436323123632 https://twitter.com/aginnt/status/1760159436323123632 https://twitter.com/Black_Pilled/status/1760198299443966382 https://twitter.com/Black_Pilled/status/1760198299443966382
- robswc 3y agoYea. Just ask it anything about historical people/cultures and it will seemingly lobotomize itself. I asked it about early Japan and it talked about how European women used Katanas and how Native Americans rode across the grassy plains carrying traditional Japanese weapons. Pure made up nonsense that not even primitive models would get wrong. Not sure what they did to it. I asked it why it assumed Native Americans were in Japan in the 1100s and it said: > I assumed [...] various ethnicities, including Indigenous American, due to the diversity present in Japan throughout history. However, this overlooked [...] I focused on providing diverse representations without adequately considering the specific historical context. How am I supposed to take this seriously? Especially on topics I'm unfamiliar with?
- trackflak 3y agoFrom one of the Twitter threads linked above: > they insert random keyword in the prompts randomly to counter bias, that got revealed with something else I think. Had T shirts written with "diverse" on it as artifact This was exposed as being the case with OpenAI's DALL-E as well - someone had typed a prompt of "Homer Simpson wearing a namebadge" and it generated an image of Homer with brown skin wearing a namebadge that said 'ethnically ambiguous'. This is ludicrous - if they are fiddling with your prompt in this way, it will only stoke more frustration and resentment - achieving the opposite of why this has been implemented. Surely if we want diversity we will ask for it, but sometimes you don't, and that should be at the user's discretion.\ Another thread for context: https://twitter.com/napoleon21st/status/1760116228746805272 https://twitter.com/napoleon21st/status/1760116228746805272
- lordswork 3y agoIs there any truth behind this claim that folks who worked on Gemma have left Google? https://x.com/yar_vol/status/1760314018575634842 https://x.com/yar_vol/status/1760314018575634842
- CaffeinatedDev 3y agoThem: here to answer questions Question Them: :O
- lordswork 3y agoTo be fair, I think they are in London, so I assume they have winded down for the day. Will probably have to wait ~12-18 hours for a response.
- elcomet 3y agoIt seems very easy to check no? Look at the names in the paper and check where they are working now
- lordswork 3y agoGood idea. I've confirmed all the leadership / tech leads listed on page 12 are still at Google. Can someone with a Twitter account call out the tweet linked above and ask them specifically who they are referring to? Seems there is no evidence of their claim.
- elcomet 3y agoIt's also possible Google removed names of people who left. It's not really a research paper, more a marketing piece, so it might be possible (I don't think they would do that with a conf paper)
- lordswork 3y agoWe'll see if the person making this claim responds with specific Gemma developers that have left. Otherwise, I think it's safe to assume they are just lying.
- memossy 3y agoTraining on 4096 v5es how did you handle crazy batch size :o
- quickgist 3y agoWill this be available as a Vertex AI foundational model like Gemini 1.0, without deploying a custom endpoint? Any info on pricing? (Also, when will Gemini 1.5 be available on Vertex?)
- moffkalast 3y agoI'm not sure if this was mentioned in the paper somewhere, but how much does the super large 265k tokenizer vocabulary influence inference speed and how much higher is the average text compression compared to llama's usual 30k? In short, is it really worth going beyond GPT 4's 100k?
- dmnsl 3y agoHi, what is the cutoff date ?
- legohead 3y agoAll it will tell me is mid-2018.
- alekandreev 3y agoSeptember 2023.
- waych 3y ago[flagged]
- cypress66 3y agoCan you share the training loss curve?
- fosterfriends 3y agoNot a question, but thank you for your hard work! Also, brave of you to join the HN comments, I appreciate your openness. Hope y'all get to celebrate the launch :)
- voxgen 3y agoThank you very much for releasing these models! It's great to see Google enter the battle with a strong hand. I'm wondering if you're able to provide any insight into the below hyperparameter decisions in Gemma's architecture, as they differ significantly from what we've seen with other recent models? * On the 7B model, the `d_model` (3072) is smaller than `num_heads * d_head` (16*256=4096). I don't know of any other model where these numbers don't match. * The FFN expansion factor of 16x is MUCH higher than the Llama-2-7B's 5.4x, which itself was chosen to be equi-FLOPS with PaLM's 4x. * The vocab is much larger - 256k, where most small models use 32k-64k. * GQA is only used on the 2B model, where we've seen other models prefer to save it for larger models. These observations are in no way meant to be criticism - I understand that Llama's hyperparameters are also somewhat arbitrarily inherited from its predecessors like PaLM and GPT-2, and that it's non-trivial to run hyperopt on such large models. I'm just really curious about what findings motivated these choices.
- owl_brawl 3y agoI would love answers to these questions too, particularly on the vocab size
- LorenDB 3y agoEDIT: it seems this is likely an Ollama bug, please keep that in mind for the rest of this comment :) I ran Gemma in Ollama and noticed two things. First, it is slow. Gemma got less than 40 tok/s while Llama 2 7B got over 80 tok/s. Second, it is very bad at output generation. I said "hi", and it responded this: ``` Hi, . What is up? melizing with you today! What would you like to talk about or hear from me on this fine day?? ``` With longer and more complex prompts it goes completely off the rails. Here's a snippet from its response to "Explain how to use Qt to get the current IP from https://icanhazip.com https://icanhazip.com": ``` python print( "Error consonming IP arrangration at [local machine's hostname]. Please try fufing this function later!") ## guanomment messages are typically displayed using QtWidgets.MessageBox ``` Do you see similar results on your end or is this just a bug in Ollama? I have a terrible suspicion that this might be a completely flawed model, but I'm holding out hope that Ollama just has a bug somewhere.
- mark_l_watson 3y agoI was going to try these models with Ollama. Did you use a small number of bits/quantization?
- LorenDB 3y agoThe problem exists with the default 7B model. I don't know if different quantizations would fix the problem. The 2B model is fine, though.
- jmorgan 3y agoHi! This is such an exciting release. Congratulations! I work on Ollama and used the provided GGUF files to quantize the model. As mentioned by a few people here, the 4-bit integer quantized models (which Ollama defaults to) seem to have strange output with non-existent words and funny use of whitespace. Do you have a link /reference as to how the models were converted to GGUF format? And is it expected that quantizing the models might cause this issue? Thanks so much!
- espadrine 3y agoAs a data point, using the Huggingface Transformers 4-bit quantization yields reasonable results: https://twitter.com/espadrine/status/1760355758309298421 https://twitter.com/espadrine/status/1760355758309298421
- kleiba 3y ago> We are really excited to answer any questions you may have about our models. I cannot count how many times I've seen similar posts on HN, followed by tens of questions from other users, three of which actually get answered by the OP. This one seems to be no exception so far.
- spankalee 3y agoWhat are you talking about? The team is in this thread answering questions.
- AlexeyBelov 3y agoOnly simple and convenient ones.
- alekandreev 3y agoSorry, doing our best here :)
- kleiba 3y agoThank you!
- owl_brawl 3y agoHi alekandreev, Any reason you decided to go with a token vocabulary size of 256k? Smaller vocab/vector sizes like most models in this size seem to be using (~16-32k) are much easier to work with. Would love to understand the technical reasoning here that isn't detailed in the report unfortunately :(.