4 ms·
If anyone is interested in evaling Gemma locally, this can be done pretty easily using ollama[0] and promptfoo[1] with the following config: prompts: - '
by typpo 2y ago
If anyone is interested in evaling Gemma locally, this can be done pretty easily using ollama[0] and promptfoo[1] with the following config:
prompts:
- 'Answer this coding problem in Python: {{ask}}'
providers:
- ollama:chat:gemma2:9b
- ollama:chat:llama3:8b
tests:
- vars:
ask: function to find the nth fibonacci number
- vars:
ask: calculate pi to the nth digit
- # ...
One small thing I've always appreciated about Gemma is that it doesn't include a "Sure, I can help you" preamble. It just gets right into the code, and follows it with an explanation. The training seems to emphasize response structure and ease of comprehension.
Also, best to run evals that don't rely on rote memorization of public code... so please substitute with your personal tests :)
[0] https://ollama.com/library/gemma2 https://ollama.com/library/gemma2
[1] https://github.com/promptfoo/promptfoo https://github.com/promptfoo/promptfoo
- deleted 2y ago[deleted]
- roywiggins 2y agoIn Ollama, Gemma:9b works fine, but 27b seems to be producing a lot of nonsense for me. Asking for a bit of python or JavaScript code rapidly devolves into producing code-like gobbledegook, extending for hundreds of lines.
- bugglebeetle 2y agoThe tokenizer in llama.cpp probably needs fixing then or it has some other bug.
- 0x7cfe 2y agoDefinitely. I tried gemma2:27B model with phrases like "translate the following sentence to language X" and it even failed to understand the task and spat out completely irrelevant things, like math formulas. OTOH, smaller model did it perfectly.
- brandall10 2y ago27b is working fine for me, hosted on ollama w/ continue.dev in VSCode.
- thot_experiment 2y agoHad a chance to do some testing and it seems quite good on oneshot tasks with a small context window but as you approach context saturation it starts to go way off the rails. Maybe this is an implementation issue? I'm using Q6_K quants of both sizes in ollama. I'll report back if I figure it out. A larger context window really helps on RAG tasks, it's frustrating that a lot of the foundational models have such small windows.
- jmorgan 2y agoSorry about this – working on fixing the issue with hitting the context limit. Gemma 2 supports a 8192 context limit – which can be selected if you provide the `num_ctx` parameter in the API or via `ollama run` with `/set parameter num_ctx 8192`
- thot_experiment 2y agoThanks! If you have a moment can you give me a quick explainer on what happens when you hit the context limit in ollama? I had assumed that ollama would just trunc the context to whatever is set in the model, but I guess this isn't the case?
- jmorgan 2y agoCurrently when the context limit is hit, there's a halving of the context window (or a "context shift") to allow inference to continue – this is helpful for smaller (e.g. 1-2k) context windows. However, not all models (especially newer ones) respond well to this, which makes sense. We're working on changing the behavior in Ollama's API to be more similar to OpenAI, Anthropic and similar APIs so that when the context limit is hit, the API returns a "limit" finish/done reason. Hope this is helpful!