Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ssheng
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
ssheng
2y ago
To get a full A10, one must also get 36 vCPU and 440 GB of memory. I must be missing something.
2.
▲
What is Azure thinking with this VM size NVadsA10 v5-series
(learn.microsoft.com)
1 points
by
ssheng
2y ago
|
1 comments
3.
▲
by
ssheng
2y ago
Quality loss with quantization is expected. It seems like with GPTQ the loss is within acceptable range based on the perplexity score shown.
4.
▲
by
ssheng
2y ago
How does Exllama rank among these? Heard good things about it.
5.
▲
by
ssheng
2y ago
Creative idea. Any data on how much it takes to load the LoRAs and how much latency it adds to the generation speed?
6.
▲
by
ssheng
2y ago
Hello! We're the authors of this blog post. Please let us know if there are other models and inference backends you'd like us to benchmark next.
7.
▲
by
ssheng
4y ago
The user feedback seems really encouraging. https://www.linkedin.com/posts/eric-riddoch_its-early-to-say...
8.
▲
by
ssheng
4y ago
FastAPI is great building block but can't expect it to work for model serving out of box.