Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
smarterclayton
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
smarterclayton
1y ago
A good rule of thumb is that a prefill token is about 1/6th the compute cost of decode token, and that you can get about 15k prefill tokens a second on Llama3 8B on a single H100. Bigger models will require more compute per token, and
2.
▲
by
smarterclayton
1y ago
That's a good example - I can at least answer about why it's a difference: different target user. As I understand the Dynamo SDK it is about simplifying and helping someone get started with Dynamo on Kubernetes. From the user se
3.
▲
by
smarterclayton
1y ago
llm-d is intended to be three clean layers: 1. Balance / schedule incoming requests to the right backend 2. Model server replicas that can run on multiple hardware topologies 3. Prefix caching hierarchy with well-tested variants for di
4.
▲
by
smarterclayton
1y ago
We inherit any multi-host support from vLLM, so https://docs.vllm.ai/en/latest/serving/distributed_serving.h... would be the expected path. We plan to publish examples of multi-host inference that leverages L
5.
▲
by
smarterclayton
1y ago
Inference is the process of evaluating a model ("inferring" a response to the inputs). LLMs are uniquely difficult to serve because they push the limits on the hardware. The models we support come from the model server vLLM https
6.
▲
by
smarterclayton
1y ago
llm-d would make sense if you are running a very large production LLM serving setup - say 5+ full H100 hosts. The aim is to be much more focused than kserve is on exactly the needs of serving LLMs. It would of course be possible to run alo
7.
▲
LLM-D: Kubernetes-Native Distributed Inference
(llm-d.ai)
120 points
by
smarterclayton
1y ago
|
15 comments
8.
▲
by
smarterclayton
2y ago
CPU tracking is provided by the metrics API, which either reads kubelet metrics directly (the original, old, but simplest way), or a metrics adapter that reads the metrics from a third party collector and implements the API. The behavior is
9.
▲
by
smarterclayton
2y ago
Pretty well. Anthropic runs some of Claude inference on GKE and TPU v5e - this talk at the last GCP conference has some details: https://youtu.be/b87I1plPeMg?si=T4XSFUzXG8BwpphR Ecosystem support for GPU is very strong and
10.
▲
by
smarterclayton
3y ago
There is a lot of work to make the actual infrastructure and lower level management of lots and lots of GPUs/TPUs open as well - my team focuses on making the infrastructure bit at least a bit more approachable on GKE and Kubernetes.
11.
▲
by
smarterclayton
3y ago
As an aside: This principle (always shutdown uncleanly) was a significant point of design discussion in Kubernetes, another one of the projects that adapted lessons learned inside Google on the outside (and had to change as a result). All o
12.
▲
Reference Architecture for ML Training and Batch on GKE with Kueue
(github.com)
2 points
by
smarterclayton
3y ago
|
0 comments
13.
▲
by
smarterclayton
3y ago
From the report: "As part of the evaluation process, on a popular benchmark, HellaSwag (Zellers et al., 2019), we find that an additional hundred finetuning steps on specific website extracts corresponding to the HellaSwag training set
14.
▲
by
smarterclayton
3y ago
True. The pod’s monotonic and atomic lifecycle across containers is a significant difference, but you can broadly accomplish similar behaviors with an alloc for sharing resources.
15.
▲
by
smarterclayton
3y ago
I will note that a trend I have observed with recent ML - as we increasingly use accelerators and models correspondingly grow in size, we are returning to a "one machine, one workload" paradigm for the biggest training and inferen
16.
▲
by
smarterclayton
3y ago
Also, I should point out that a set of machines hosting TPUs is referred to as a "pod", which is not the same thing as a Kubernetes pod (also referenced in this doc). The term "pod" originated in early data center design
17.
▲
by
smarterclayton
3y ago
Disclaimer: work associated with this team, didn't write or review the blog post Article stated that it was throughput scheduling the pods on the clusters (from unrelated benchmarks that's usually ~300 pods/sec throughput for
18.
▲
by
smarterclayton
3y ago
Our bottleneck was serialization of objects to bytes to send to etcd. Etcd cpu should be about 0.01-0.001 control plane CPU, and control plane apiserver CPU has been dominated by serialization since day 1 (except for brief periods around w
19.
▲
by
smarterclayton
3y ago
That’s fair - as the perpetrator of much of that abstraction I believe that highlighting the type system aspect of this code obscures the real problem we were solving for, which Go is spectacularly good at: letting you get within inches of
20.
▲
by
smarterclayton
3y ago
I have a lot of respect for Kris but in this context, as the person who approved the PR adding the “Object” interface to the code base (and participated in most of the subsequent discussions about how to expand it), it was not done because
21.
▲
by
smarterclayton
3y ago
> kubernetes went ahead and implemented an oop system in golang I don’t think this was ever an objective, can you clarify what you mean by “oop system” and where we implemented it? We aggressively used composition of interfaces up to a p
22.
▲
by
smarterclayton
3y ago
Agreed that is a feature and not a bug. But! The one thing that custom orchestrators can’t do is easily get the benefit of kubelet isolation of containers and resource management. Part of slowly moving down this path is to allow those orc
23.
▲
by
smarterclayton
3y ago
In general, the intent here is to leave open room for just that. dependsOn was proposed during the kep review but deferred. But because init containers and regular containers share the same behavior and shape, and differ only on container
24.
▲
by
smarterclayton
3y ago
The challenge with a separate attribute is that it is not forward compatible with new features we might add to pods around ordering and lifecycle. If we used a simple boolean, eventually we’d have to have it interact with other fields and
25.
▲
by
smarterclayton
3y ago
We did that to leave open more complex ordering of both init containers and sidecars (regular containers do not have a restart order). For instance, you might have a service mesh that needs a vault secret - those both might be sidecars, an
26.
▲
by
smarterclayton
4y ago
In that scenario it looks like members must coordinate to identify the highest committed transaction (identifying the list of valid members) and then bootstrap from that member? Stateful Sets were designed to standardize two hard problems:
27.
▲
by
smarterclayton
4y ago
They are pretty well tested as of today (now that multiple vendors respect it during node upgrade), but now they’re relatively under featured for the next set of problems: 1. No way to signal that the workload is ready to accept traffic but
28.
▲
by
smarterclayton
4y ago
Since I helped design them, I take some issue with that :). Certainly we never expected they would completely solve problems for the database, but they were definitely intended to provide guarantees that simplify normal consensus operatio
29.
▲
by
smarterclayton
4y ago
Re out of order: Is https://kubernetes.io/docs/concepts/workloads/controllers/st... unsuitable for that?
30.
▲
by
smarterclayton
4y ago
You might be thinking about kcp? https://github.com/kcp-dev/kcp https://www.youtube.com/watch?v=oaPBYUfdFE8 The team is developing the idea and has made a lot of progress. Fair warning, some of the re
More ›