Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
JegernOUTT
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
JegernOUTT
3y ago
It can be checked if the model predicts canonical solutions from humaneval. I understand it is not ideal, but at least you can check it yourself There are a bunch of other benchmarks too, check out the page https://huggingface.co
2.
▲
by
JegernOUTT
3y ago
Hi, thank you for your attention! > They compare the performance of this model to the worst 7B code llama model. The base code llama 7B python model scores 38.4% on humaneval, versus the non-python model, which only scores 33%. We are co