Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
philipkiely
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
philipkiely
19d ago
Hey all Philip from Baseten here. Posting this on behalf of our security team. I wanted to confirm that we collaborated with Strix on the remediation of the reported vulnerability. We thank Strix for their responsible disclosure. We took im
2.
▲
by
philipkiely
1mo ago
I draw my diagrams on notecards and send them to our designer who brings them to life. The images start out looking like this: https://philipkiely.com/images/blogs/how-to-write-a-book/des...
3.
▲
by
philipkiely
1mo ago
I also wrote this as somewhat of a defense of the techniques that don't move the frontier -- there is a lot of value in being able to pick a point on the curve.
4.
▲
by
philipkiely
1mo ago
These are both good points that I attempted to cover, quotes: > In practice, the efficient frontier is very jagged. Rather than a smooth, continuous line between outcomes, small changes can have big impacts. These cutoff points are often
5.
▲
by
philipkiely
1mo ago
I think the biggest net new recent technique is P/D disaggregation. And that spec dec is very different now especially post DSpark/DFlash. But overall yes the fundamentals of LLM performance optimization have been remarkably stabl
6.
▲
The efficient frontier of LLM inference
(baseten.co)
155 points
by
philipkiely
1mo ago
|
46 comments
7.
▲
We built a day-0 API for Kimi K3
(baseten.co)
2 points
by
philipkiely
2mo ago
|
0 comments
8.
▲
We built the new fastest API for GLM-5.2
(baseten.co)
2 points
by
philipkiely
2mo ago
|
0 comments
9.
▲
by
philipkiely
3mo ago
Good thing we have GLM-5.2
10.
▲
We built the fastest API for GLM-5.2 (280 TPS)
(baseten.co)
6 points
by
philipkiely
3mo ago
|
0 comments
11.
▲
by
philipkiely
6mo ago
https://github.com/AliesTaha/polar_quant
12.
▲
The Math Behind TurboQuant
(baseten.co)
8 points
by
philipkiely
6mo ago
|
3 comments
13.
▲
Show HN: Inference Engineering
(baseten.com)
2 points
by
philipkiely
7mo ago
|
0 comments
14.
▲
How We Built the Fastest Kimi K2.5 on Artificial Analysis
(baseten.co)
3 points
by
philipkiely
8mo ago
|
0 comments
15.
▲
Nvidia Invests $150M in AI Inference Startup Baseten
(wsj.com)
1 points
by
philipkiely
8mo ago
|
1 comments
16.
▲
by
philipkiely
10mo ago
GLM 4.6 has been very popular from my perspective as an inference provider with a surprising number of people using it as a daily driver for coding. Excited to see the improvements 4.7 delivers, this model has great PMF so to speak.
17.
▲
by
philipkiely
10mo ago
The Information link, for those with a subscription: https://www.theinformation.com/articles/inference-provider-b...
18.
▲
by
philipkiely
11mo ago
You give it a text prompt and optional image. What you get is a 3D room based on the prompt/image. It rewrites your prompt to a specific format. Overall the rooms tend to be detailed and imaginative. Then you can fly around the room li
19.
▲
by
philipkiely
11mo ago
I played with Marble yesterday, Fei-Fei/World Labs' new product. It is the most impressed I've been with an AI experience since the first time I saw a model one-shot material code. Sure, its an early product. The visual outpu
20.
▲
Baseten raises $150M Series D at $2.15B
(fortune.com)
2 points
by
philipkiely
1y ago
|
1 comments
21.
▲
by
philipkiely
1y ago
We have built a ton of tooling on top of TRT-LLM and use it not just for LLMs but also for TTS models (Orpheus), STT models (Whisper), and embedding models.
22.
▲
by
philipkiely
1y ago
Yeah the custom hardware providers are super good at TPS. Kudos to their teams for sure, and the demos of instant reasoning are incredibly impressive. That said, we are serving the model at its full 131K context window, and they are serving
23.
▲
by
philipkiely
1y ago
Yeah we have tried to build calculators before it just depends so much. Your equation is roughly correct, but I tend to multiply by a factor of 2 not 1.2 to allow for highly concurrent traffic.
24.
▲
by
philipkiely
1y ago
TRT-LLM has its challenges from a DX perspective and yeah for Multi-modal we still use vLLM pretty often. But for the kind of traffic we are trying to serve -- high volume and latency sensitive -- it consistently wins head-to-head in our be
25.
▲
by
philipkiely
1y ago
This comment made my day ty! Yeah definitely speaking from a datacenter perspective -- fastest piece of hardware I have in the parts drawer is probably my old iPhone 8.
26.
▲
by
philipkiely
1y ago
Went to bed with 2 votes, woke up to this. Thank you so much HN!
27.
▲
Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
(baseten.co)
247 points
by
philipkiely
1y ago
|
175 comments
28.
▲
by
philipkiely
1y ago
For prod inference, 1xH100 is working well.
29.
▲
by
philipkiely
1y ago
WhisperX! https://github.com/basetenlabs/truss-examples/tree/main/whis...
30.
▲
by
philipkiely
1y ago
Example implementation with sample inference code + voice cloning example: https://github.com/basetenlabs/truss-examples/tree/main/chat... Still working on streaming
More ›