Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mcharytoniuk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Show HN: I rewrote most of Llama-server in Rust, and made it scalable
(old.reddit.com)
3 points
by
mcharytoniuk
1y ago
|
0 comments
2.
▲
Ask HN: Async Python Adoption?
1 points
by
mcharytoniuk
2y ago
|
0 comments
3.
▲
by
mcharytoniuk
2y ago
Yes, exactly. You can split the available context into "slots" (chunks) so it can handle multipe requests concurrently. The number of them is configurable.
4.
▲
by
mcharytoniuk
2y ago
It divides the context into smaller "slots", so it can process requests concurrently with continuous batching. See also: https://github.com/ggerganov/llama.cpp/tree/master/examples/...
5.
▲
by
mcharytoniuk
2y ago
Just open an issue if you need anything. I want to make it as good and helpful as possible. Every kind of feedback is appreciated.
6.
▲
by
mcharytoniuk
2y ago
Currently, it is a single instance in memory, so it doesn't transfer state. HA is on the roadmap; only then will it need some kind of distributed state store. Local states are reported by the agents installed alongside llama.cpp to the
7.
▲
by
mcharytoniuk
2y ago
In progress. I added that to the readme; I need the feature myself. :)
8.
▲
Show HN: Open-source load balancer for llama.cpp
(github.com)
151 points
by
mcharytoniuk
2y ago
|
20 comments
9.
▲
Show HN: Language-agnostic structured data extractor
(github.com)
3 points
by
mcharytoniuk
2y ago
|
0 comments
10.
▲
Show HN: Resonance – PHP Framework That Solves Real-Life Issues
(github.com)
2 points
by
mcharytoniuk
3y ago
|
0 comments