Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kouroshh
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
kouroshh
3y ago
Hey, I am one of the co-authors of the post. So the training data for ViGGO has about 5.1k rows which we trained with a block size of 512 (you can lower the block size if you want but we didn't do so because it was just easier to not c
2.
▲
by
kouroshh
3y ago
Would be good to see a rigorous analysis of these PEFT methods on quality. There still seems to be a debate on whether these methods sacrifice quality or not.
3.
▲
by
kouroshh
3y ago
Llama-2-chat models have been overly fine-tuned to be like this. You can give a few-shot prompting a try, but they still don't gurantee a desired output. The best way to guarantee is to fine-tune on small (~1k) data points and go from
4.
▲
by
kouroshh
3y ago
Training times for GSM8k are mentioned here: https://github.com/ray-project/ray/tree/master/doc/source/te...