Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
george_123
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Loading Llama-2 70B 20x faster with Anyscale Endpoints
(anyscale.com)
3 points
by
george_123
3y ago
|
0 comments
2.
▲
by
george_123
3y ago
1) when the blog was released, TGI didn’t support paged attention, 2) many people don’t even know about TGI to reduce inference costs.
3.
▲
by
george_123
3y ago
A lot of AWS services (especially SageMaker) aren’t very good in customer experience. People buy them for nominal capabilities and AWS core bread and butter — short-term and long-term reliability. Most of these startups (AI and others) have
4.
▲
How continuous batching improves LLM inference throughput 23x
(twitter.com)
1 points
by
george_123
3y ago
|
0 comments
5.
▲
by
george_123
3y ago
this approach to managing KV cache can work with 4bit. imagine the speedup of pagedattention with quantization..
6.
▲
Ant Group – scaling to 1.37M QPS on Ray
(anyscale.com)
3 points
by
george_123
4y ago
|
1 comments
7.
▲
by
george_123
4y ago
This is a guest engineering blog post from Ray contributors at Ant Group, discussing how Ant Group implemented scalable Ray Serving architecture atop Ray, deploying 240,000 cores for model serving, scaling by 3.5x from previous year, and re
8.
▲
by
george_123
4y ago
This is a guest engineering blog post from Ray contributors at Ant Group, discussing how Ant Group implemented scalable Ray Serving architecture atop Ray, deploying 240,000 cores for model serving, scaling by 3.5x from previous year, and re