Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mich5632
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
mich5632
1y ago
I think this the difference between compute bound pre-fill (a cpu has a high bandwidth/compute ratio), vs decode. The time to first token is below 0.5s - even for a 10k context.
2.
▲
by
mich5632
1y ago
We wrote a rust py03 client for OpenAI embeddings compatible servers (openai.com, or infinity, TEI, vllm, sglang). Most server-side ML infrastructure auto-scales based on the workload. On embedding workloads, this is no longer the bottlene
3.
▲
High performance client for Baseten.co
(github.com)
7 points
by
mich5632
1y ago
|
1 comments
4.
▲
by
mich5632
2y ago
Looking forward to a serving system that can actually use this!
5.
▲
by
mich5632
3y ago
Looking great, thanks for building this