Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alekandreev
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
alekandreev
2y ago
Yes we have measured the tradeoff. We don't see a drop of perplexity in English when introducing multilingual, and there is a slight drop in some English language-specific evals (~1%).
2.
▲
by
alekandreev
2y ago
We never train at 128k, only 32k, changing the scaling factor at the end. We wanted the long context recipe to be friendly for finetuning, and training at 128k is a bit of a pain we don't do it. For inference, we see inference at 128k
3.
▲
by
alekandreev
2y ago
We recommend using <start_of_turn>user for the system prompt as well.
4.
▲
by
alekandreev
2y ago
That's a very interesting area, but nothing we can announce today.
5.
▲
by
alekandreev
2y ago
Hey, Gemma engineer here. Can you please share reports on the type of prompts and the implementation you used?
6.
▲
by
alekandreev
2y ago
Thank you for the report! We are working with the Ollama team directly and will look into it.
7.
▲
by
alekandreev
2y ago
Thank you for the feedback! This is why we are so excited to push more and more on small models for both low end and high end smartphones!
8.
▲
by
alekandreev
2y ago
That's an idea we've thought about. However, we think the open source community has already created a very impressive set of language or region-specific finetunes [1] [2]. Also there is a lot of cultural and nuance context in ever
9.
▲
by
alekandreev
2y ago
Picking model sizes is not an exact science. We look for sizes that will fit quantized on different categories on devices (e.g., low-end and high-end smartphone, laptops and 16GB GPUs, and bigger GPUs/TPUs). We also want the ratio of m
10.
▲
by
alekandreev
2y ago
Greetings from the Gemma team! We just got Gemma 3 out of the oven and are super excited to show it to you! Please drop any questions here and we'll answer ASAP. (Opinions our own and not of Google DeepMind.) PS we are hiring: https:&
11.
▲
by
alekandreev
2y ago
I think it makes sense to compare models trained with the same recipe on token count - usually more tokens will give you a better model. However, I wouldn't draw conclusions about different model families, like Llama and Gemma, based o
12.
▲
by
alekandreev
2y ago
To quote Ludovic Peran, our amazing safety lead: Literature has identified self-proliferation as dangerous capability of models, and details about how to define it and example of form it can take have been openly discussed by GDM ( https:&#
13.
▲
by
alekandreev
2y ago
Happy to pass on any feedback to our Google Cloud friends. :)
14.
▲
by
alekandreev
2y ago
This is mostly about inference speed, while maintaining long context performance.
15.
▲
by
alekandreev
2y ago
In addition to the HF links shared by sibling comments, the 2B will be released soon.
16.
▲
by
alekandreev
2y ago
The terms of use remain the same as Gemma 1 - https://ai.google.dev/gemma/terms .
17.
▲
by
alekandreev
2y ago
Your training input has the shape of (sequence length x batch size). If a lot of your samples are shorter than sequence length, as is usually the case, you will have a lot of padding tokens in the input, which is wasted compute. To compensa
18.
▲
by
alekandreev
2y ago
Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions! Opinions are our own and not of Google DeepMind.
19.
▲
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models [pdf]
(storage.googleapis.com)
6 points
by
alekandreev
2y ago
|
1 comments
20.
▲
MediaPipe: LLM Inference on iOS and Android for Gemma, Phi-2 and Others
(developers.googleblog.com)
1 points
by
alekandreev
3y ago
|
0 comments
21.
▲
by
alekandreev
3y ago
As a fellow Bulgarian from the 80s and 90s myself, and now a part of the Gemma team, I’d say Austin, Jan, and team very much live up to the ethos of hackers I'd meet on BBSes back then. :) They are driven entirely by their own curiosit
22.
▲
by
alekandreev
3y ago
Sorry, doing our best here :)
23.
▲
by
alekandreev
3y ago
We have many exciting things planned that we can't reveal just yet :)
24.
▲
by
alekandreev
3y ago
September 2023.
25.
▲
by
alekandreev
3y ago
We deeply respect the Phi team and all other teams in the open model space. You’ll find that different models have different strengths and not all can be quantified with existing public evals. Take them for a spin and see what works for yo
26.
▲
by
alekandreev
3y ago
We have implementations in different ML frameworks, so I am not quite sure which one you are referring to. Would you like to file a bug at the relevant GitHub repo?
27.
▲
by
alekandreev
3y ago
This would be really interesting in my opinion, but we are not releasing datasets at this time. See the C4 dataset for an earlier open dataset from Google.
28.
▲
by
alekandreev
3y ago
This v1 model is focused on English support, but you may find some multilingual capabilities.
29.
▲
by
alekandreev
3y ago
We have many great things in research and development phases, so stay tuned. I’m hopeful we can share more in the coming weeks and month!
30.
▲
by
alekandreev
3y ago
Hello on behalf of the Gemma team! We are really excited to answer any questions you may have about our models. Opinions are our own and not of Google DeepMind.
More ›