Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kathleenfromgdm
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
kathleenfromgdm
3y ago
The context length for these models is 8192 tokens.
2.
▲
by
kathleenfromgdm
3y ago
Good catch - just corrected. Thanks!
3.
▲
by
kathleenfromgdm
3y ago
The context length for these models is 8192 tokens.
4.
▲
by
kathleenfromgdm
3y ago
Great question - we compare to the Mistral 7B 0.1 pretrained models (since there were no pretrained checkpoint updates in 0.2) and the Mistral 7B 0.2 instruction-tuned models in the technical report here: https://goo.gle/Gem
5.
▲
by
kathleenfromgdm
3y ago
We release our non-aligned models (marked as pretrained or PT models across platforms) alongside our fine-tuned checkpoints; for example, here is our pretrained 7B checkpoint for download: https://www.kaggle.com/models/
6.
▲
by
kathleenfromgdm
3y ago
Corrected - thanks :)
7.
▲
by
kathleenfromgdm
3y ago
Yes, you can get started downloading the model and running inference on Kaggle: https://www.kaggle.com/models/google/gemma ; for a full list of ways to interact with the model, you can check out https://
8.
▲
by
kathleenfromgdm
3y ago
Thank you! You can get started downloading the model and running inference on Kaggle: https://www.kaggle.com/models/google/gemma ; for a full list of ways to interact with the model, you can check out https:/
9.
▲
by
kathleenfromgdm
3y ago
We've documented the architecture (including key differences) in our technical report here ( https://goo.gle/GemmaReport ), and you can see the architecture implementation in our Git Repo ( https://github.com&#