Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bddppq
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
bddppq
2y ago
yep with <= 8bit (int8/fp8) quantization
2.
▲
by
bddppq
3y ago
lepton is at a different layer comparing to llama.cpp, in fact for LLM model files that are of GGUF format, it's using llama.cpp (ctransformers to be precise) as the execution engine
3.
▲
by
bddppq
3y ago
Are there some examples(prompts) that Falcon 180B is performing better than Llama 70B?
4.
▲
by
bddppq
3y ago
Thanks for pointing out. We have now updated the prompt template to follow the format.