7 ms·
Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions! Opinions are our own and not of Google Deep
by alekandreev 2y ago
Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions!
Opinions are our own and not of Google DeepMind.
- zerojames 2y agoHow is Gemma-2 licensed?
- alekandreev 2y agoThe terms of use remain the same as Gemma 1 - https://ai.google.dev/gemma/terms https://ai.google.dev/gemma/terms.
- coreypreston 2y agoNo question. Thanks for thinking of 27B.
- moffkalast 2y agoThe 4k sliding window context seems like a controversial choice after Mistral 7B mostly failed at showing any benefits from it. What was the rationale behind that instead of just going for full 8k or 16k?
- alekandreev 2y agoThis is mostly about inference speed, while maintaining long context performance.
- luke-stanley 2y agoAny gemma-2-9b or 27b 4 bit GGUF's on HuggingFace yet? Thanks!
- XzAeRosho 2y agoIt's on HuggingFace already: https://huggingface.co/google/gemma-2-9b https://huggingface.co/google/gemma-2-9b
- luke-stanley 2y agoI know the safe tensors are there, but I said GGUF 4-bit quantised, which is kinda the standard for useful local applications, a typical balanced sweet spot of performance and quality. It's makes it much easier to use, works in more places, be it personal devices or a server etc.
- chown 2y agoIf you are still looking for it, I just made it available on an app[1] that I am working on with Gemma2 support. https://msty.app https://msty.app
- luke-stanley 2y agoAre you saying you put a 4-bit GGUF on HuggingFace?
- luke-stanley 2y agoActually for the 9B model, this has 4-bit quantised weights (and others): https://huggingface.co/bartowski/gemma-2-9b-it-GGUF https://huggingface.co/bartowski/gemma-2-9b-it-GGUF Still no 27B 4-bit GGUF quants on HF yet! I'm monitoring this search: https://huggingface.co/models?library=gguf&sort=trending&search=gemma-2+27b https://huggingface.co/models?library=gguf&sort=trending&sea...
- SubiculumCode 2y agohttps://huggingface.co/bartowski/gemma-2-27b-it-GGUF https://huggingface.co/bartowski/gemma-2-27b-it-GGUF
- thot_experiment 2y agoI'm curious about the quantization quality claims in the table there. Is this a Gemma 2 specific thing (more subtlety in the weights somehow?). In my testing and testing I've seen elsewhere at least for llama3 8B (and some less rigorous testing with other models) q_8 -> q4_K_M are basically indistinguishable from one another?
- janwas 2y agoYes, PPL and certain benchmarks do not detect differences from quantization. But recent work gives cause for concern, e.g., https://arxiv.org/pdf/2310.01382 https://arxiv.org/pdf/2310.01382, https://arxiv.org/pdf/2405.18137 https://arxiv.org/pdf/2405.18137.
- luke-stanley 2y agoThe first paper is good to critique the performance of quantised models, it points out that 40-50% 'compression' typically results in only slight loss for RAG tasks relying on in-context learning, but for factual tasks replying on stored knowledge, performance very quickly dropped off. They looked at Vicuna, one of the earlier models, so I wonder how applicable it is to recent models like the Phi 3 range. I don't think deliberate clever adversarial attacks like those of the 2nd paper are a sensible worry for most, but it is fun. Thanks for the links @janwas.
- jpcapdevila 2y agoWill gemma2 be available through gemma.cpp? https://github.com/google/gemma.cpp https://github.com/google/gemma.cpp
- austinvhuang 2y agoThis is in the works in the dev branch (thanks pchx :) https://github.com/google/gemma.cpp/pull/274 https://github.com/google/gemma.cpp/pull/274
- janwas 2y ago:) Confirmed working. We've just pushed the dev branch to main.
- jpcapdevila 2y agoAwesome, I love this .cpp trend! Thanks for your work!!
- luke-stanley 2y agoGiven the goal of mitigating self-proliferation risks, have you observed a decrease in the model's ability to do things like help a user setup a local LLM with local or cloud software? How much is pre-training dataset changes, how much is tuning? How do you think about this problem, how do you solve it? Seems tricky to me.
- luke-stanley 2y ago[flagged]
- alekandreev 2y agoTo quote Ludovic Peran, our amazing safety lead: Literature has identified self-proliferation as dangerous capability of models, and details about how to define it and example of form it can take have been openly discussed by GDM (https://arxiv.org/pdf/2403.13793 https://arxiv.org/pdf/2403.13793). Current Gemma 2 models' success rate to end-to-end challenges is null (0 out 10), so the capabilities to perform such tasks are currently limited.
- moffkalast 2y agoTurns out LLM alignment is super easy, barely an inconvenience.
- dinosaurdynasty 2y agoOne should not confuse alignment and current incapability.
- josh-sematic 2y agoAlignment is tight!
- mdrzn 2y agoWow wow wow.... wow.
- luke-stanley 2y ago
- canyon289 2y agoI also work at Google and on Gemma (so same disclaimers) You can try 27b at www.aistudio,google.com. Send in your favorite prompts, and we hope you like the responses.
- dandanua 2y agoWhy is AIStudio not available in Ukraine? I have no problem with using Gemini web UI or other LLM providers from Ukraine, but this Google API constrain is strange.
- luke-stanley 2y agoIt's fairly easy to pay OpenAI or Mistral money to use their API's. Figuring out how Google Cloud Vertex works and how it's billed is more complicated. Azure and AWS are similar in how complex they are to use for this. Could Google Cloud please provide an OpenAI compatible API and service? I know it's a different department. But it'd make using your models way easier. It often feels like Google Cloud has no UX or end-user testing done on it at all (not true for aistudio.google.com - that is better than before, for sure!).
- alekandreev 2y agoHappy to pass on any feedback to our Google Cloud friends. :)
- luke-stanley 2y agoThank you!
- anxman 2y agoI also hate the billing. It feels like configuring AWS more than calling APIs.
- hnuser123456 2y agoI plan on downloading a Q5 or Q6 version of the 27b for my 3090 once someone puts quants on HF, loading it in LM studio and starting the API server to call it from my scripts based on openai api. Hopefully it's better at code gen than llama 3 8b.
- bapcon 2y agoI have to agree with all of this. I tried switching to Gemini, but the lack of clear billing/quotas, horrible documentation, and even poor implementation of status codes on failed requests have led me to stick with OpenAI. I don't know who writes Google's documentation or does the copyediting for their console, but it is hard to adapt. I have spent hours troubleshooting, only to find out it's because the documentation is referring to the same thing by two different names. It's 2024 also, I shouldn't be seeing print statements without parentheses.
- WhitneyLand 2y agoThe paper suggests on one hand Gemma is on the same Pareto curve as Llama3, while on the other hand seems to suggest it’s exceeded its efficiency. Is this a contradiction or am I misunderstanding something? Btw overall very impressive work great job.
- alekandreev 2y agoI think it makes sense to compare models trained with the same recipe on token count - usually more tokens will give you a better model. However, I wouldn't draw conclusions about different model families, like Llama and Gemma, based on their token count alone. There are many other variables at play - the quality of those tokens, number of epochs, model architecture, hyperparameters, distillation, etc. that will have an influence on training efficiency.
- causal 2y agoThanks for your work on this; excited to try it out! The Google API models support 1M+ tokens, but these are just 8K. Is there a fundamental architecture difference, training set, something else?
- np_space 2y agoAre Gemma-2 models available via API yet? Looks to me like it's not yet on vertexai
- zone411 2y ago"Soon" https://x.com/LechMazur/status/1806366744706998732 https://x.com/LechMazur/status/1806366744706998732
- kristianpaul 2y agoDo run gemma2 on your Google phone?